Command-Line Interface
The polymon CLI provides three main commands for training, prediction, and active learning:
Train Command
The train command is used to train machine learning or deep learning models for polymer property prediction.
Usage:
polymon train [OPTIONS]
Required Arguments:
--labelsTarget property/properties to predict. Choices:
Tg,FFV,Density,Rg,Tc. Multiple labels can be specified to train multiple models.
Common Optional Arguments:
--raw-csvPATHPath to the raw CSV file containing polymer data (default:
database/database.csv)--sourcesSOURCE [SOURCE ...]Data sources to filter from the dataset (default:
['Kaggle']) Common sources:Kaggle,PI1070,PolyMetriX,MAFA-exp--modelNAMEModel type to train:
- Tabular models (use with ``–feature-names``):
rf: Random Forestxgb: XGBoostlgbm: LightGBMcatboost: CatBoosttabpfn: TabPFN
- Graph Neural Networks:
gatv2: Graph Attention Network v2gin: Graph Isomorphism Networkpna: Principal Neighbourhood Aggregationattentivefp: Attentive Fingerprintingdimenetpp: DimeNet++gps: Graph with Positional Encodingfastkan_gatv2: KAN + GATv2kan_gps: KAN + GPS
--feature-namesFEATURE [FEATURE ...]Feature names for tabular models (default:
['rdkit2d']) Choices:rdkit2d,ecfp4,mordred,maccs,xenonpy_desc--n-foldINTNumber of folds for cross-validation (default:
1) Use5or10for reliable performance estimates--n-trialsINTNumber of trials for hyperparameter optimization (default:
None) When specified, enables Optuna-based optimization--out-dirPATHDirectory to save training results (default:
./results)--tagNAMEIdentifier for organizing this training run (default:
debug)
Deep Learning Specific Arguments:
--descriptorsFEATURE [FEATURE ...]Additional descriptors to concatenate with graph features
--hidden-dimINTHidden dimension for neural networks (default:
32)--num-layersINTNumber of layers in the neural network (default:
3)--num-epochsINTMaximum number of training epochs (default:
2500)--lrFLOATLearning rate (default:
1e-3)--batch-sizeINTBatch size for training (default:
128)--early-stopping-patienceINTPatience for early stopping (default:
250)--deviceNAMEDevice for training:
cudaorcpu(default:cuda)
Advanced Training Arguments:
--hparams-fromPATHPath to hyperparameters file (
.json,.pt, or.pkl) to reuse from a previous run--run-productionEnable production mode (95:5 train:val split, no test set)
--finetuneEnable fine-tuning of a pretrained model
--pretrained-modelPATHPath to pretrained model for fine-tuning
--finetune-csv-pathPATHPath to CSV file with fine-tuning data
--train-residualTrain on residuals from a base estimator or low-fidelity model
--estimator-nameNAMEName of empirical estimator for delta-learning (e.g.,
Density-IBM,Rg-monomer)--low-fidelity-modelPATHPath to low-fidelity model for residual learning
--emb-modelPATHPath to embedding model for property knowledge transfer
--n-estimatorINTNumber of estimators for ensemble learning (default:
1,>1enables ensemble)--ensemble-typeTYPEType of ensemble:
voting,bagging,gradient_boosting,snapshot,soft_gradient_boosting--split-modeMODEData splitting strategy:
random,source,scaffold(default:random)--normalizer-typeTYPELabel normalization:
normalizer,log_normalizer,none(default:normalizer)--augmentationEnable data augmentation (oligomer building)
--remove-hydrogensRemove hydrogens from molecular graphs
--seedINTRandom seed for reproducibility (default:
42)
Examples:
Train a Random Forest model with 5-fold cross-validation:
polymon train --labels Tg --model rf --feature-names rdkit2d --n-fold 5
Train a GNN with hyperparameter optimization:
polymon train --labels Rg --model gatv2 --n-trials 15 --n-fold 5 --num-epochs 2500
Fine-tune a pretrained model:
polymon train --labels Density --model gatv2 --finetune \
--pretrained-model ./results/gatv2/Density/model.pt \
--finetune-csv-path ./experimental.csv
Recommend Command
The rec command recommends molecules for active learning based on acquisition functions.
Usage:
polymon rec [OPTIONS]
Required Arguments:
--pool-csvPATHPath to CSV file with candidate molecules (must contain a
SMILEScolumn)--trained-modelPATHPath to trained model file (
.ptor.pkl)
Optional Arguments:
--acquisitionFUNCTIONAcquisition function:
epig(expected improvement),uncertainty, orrandom(default:uncertainty)--model-typeTYPEType of model:
kfoldorensemble(default:kfold)--sample-sizeINTNumber of molecules to recommend (default:
100)--save-pathPATHPath to save recommended molecules as CSV (default:
None)
Examples:
Select 20 most uncertain samples:
polymon rec --pool-csv pool.csv --trained-model model.pt \
--acquisition uncertainty --sample-size 20 --save-path recommended.csv
Predict Command
The predict command makes predictions on new polymer data using a trained model.
Usage:
polymon predict [OPTIONS]
Required Arguments:
--model-pathPATHPath to trained model file (
.ptor.pkl)--csv-pathPATHPath to CSV file with molecules to predict
--smiles-columnNAMEName of column containing SMILES strings
Examples:
Predict properties for new polymers:
polymon predict --model-path ./results/gatv2/Tg/model.pt \
--csv-path ./new_data.csv --smiles-column SMILES
Python API
For more advanced usage, you can use the Python API directly:
from polymon.model.base import ModelWrapper
# Load a trained model
model = ModelWrapper.from_file('results/gatv2/Tg/model.pt')
# Make predictions
predictions = model.predict(['*C*', '*CC*', '*CCC*'])
print(predictions)