polymon.data
dataset
- class polymon.data.dataset.PolymerDataset(raw_csv_path: str, feature_names: List[str], label_column: str, sources: List[str], smiles_column: str = 'SMILES', identifier_column: str = 'id', save_processed: bool = True, force_reload: bool = False, add_hydrogens: bool = False, pre_transform: Callable | None = None, estimator: Callable | None = None, must_keep: List[str] = None)[source]
Bases:
DatasetA dataset that contains the polymers. During the initialization, the SMILES strings will be featurized to the corresponding features and converted to the
Polymerobject. Ifpre_transformis provided, it will be used to update thePolymerobject. Ifestimatoris provided, it will be used to estimate the labels of thePolymerobject, and make the prediction task be the residual between the ground truth and the estimated labels.- Parameters:
raw_csv_path (str) – The path to the raw csv file. Typically, it should have at least three columns:
SMILES,label(e.g.,Rg), andSource(e.g.,PI1070).feature_names (List[str]) – The names of the features to use. The available features are in
AVAIL_FEATURESfrompolymon.data.featurizer.label_column (str) – Label column name.
sources (List[str]) – The names of the sources.
smiles_column (str) – SMILES column name.
identifier_column (str) – Identifier column name.
save_processed (bool) – Whether to save the processed dataset.
force_reload (bool) – Whether to force reload the processed dataset.
add_hydrogens (bool) – Whether to add hydrogens to the molecules.
pre_transform (Callable) – The pre-transform to apply to the data. It should be a function that takes a
Polymerobject and returns aPolymerobject.estimator (Callable) – The estimator to apply to the data. It is used to provide the estimated labels for the
Polymerobject. It should be the object ofpolymon.estimator.BaseEstimator.
- get_loaders(batch_size: int, n_train: int | float, n_val: int | float, mode: Literal['random', 'scaffold'] = 'random', num_workers: int = 0, augmentation: bool = False) Tuple[DataLoader, DataLoader, DataLoader][source]
Get the data loaders for the training, validation, and test sets.
- Parameters:
batch_size (int) – The batch size.
n_train (Union[int, float]) – The number of training samples. If it is a float, it will be converted to an integer by multiplying the length of the dataset.
n_val (Union[int, float]) – The number of validation samples. If it is a float, it will be converted to an integer by multiplying the length of the dataset.
mode (Literal['random', 'scaffold']) – The split mode.
random: Split the dataset randomly.scaffold: Split the dataset by scaffold.num_workers (int) – The number of workers for the data loaders.
augmentation (bool) – Whether to augment the training set. Currently, it only supports the augmentation of oligomers.
dedup
- class polymon.data.dedup.Dedup(df: DataFrame, label: str, must_keep: List[str] = None, rtol: float = 0.05)[source]
Bases:
object- add_df(file_path: str, smiles_col: str, label_col: str, source: str)[source]
Add a dataframe to the
self.df.- Parameters:
file_path (str) – The path to the dataframe.
smiles_col (str) – The column name of the SMILES.
label_col (str) – The column name of the label.
source (str) – The source name.
- compare(source1: str, source2: str, fitting: bool = False)[source]
Compare the label of the two sources with the same smiles.
- Parameters:
source1 (str) – The first source name.
source2 (str) – The second source name.
fitting (bool) – Whether to fit the data. If True, the function will fit the data with a linear regression and add the fitted points as a new source.
- run(sources: List[str] | None = None, save: bool = False)[source]
Run the deduplication. The deduplication is performed by the following steps:
Remove rows with higher relative difference.
Keep indices if the source is in
self.must_keep.
- Parameters:
sources (Optional[List[str]]) – The sources to deduplicate.
save (bool) – Whether to save the deduplicated dataframe.
- show_distribution(sources: List[str], bins: int = 50, alpha: float = 0.5, density: bool = False)[source]
Plot the distribution of the label for each source.
- Parameters:
sources (List[str]) – The sources to plot.
bins (int) – The number of bins.
alpha (float) – The alpha of the histogram.
density (bool) – Whether to plot the density.
featurizer
Name |
Featurizer |
Attributes in |
|---|---|---|
x |
|
|
edge |
|
|
pos |
|
|
z |
|
|
relative_position |
|
|
seq |
|
|
desc |
|
|
monomer |
|
|
Feature Name |
Available Features |
Description |
|---|---|---|
x |
Node features (default set, |
|
xenonpy_atom |
XenonPy atom features |
|
cgcnn |
CGCNN atom features |
|
source |
Source features |
|
z |
Atomic numbers as integers |
|
edge |
Default edge features ( |
|
bond |
Chemical bond features |
|
fully_connected_edges |
Fully connected edge indices |
|
periodic_bond |
Add bonds between attachment points to the chemical bonds |
|
virtual_bond |
Add virtual bonds between virtual node and each other node |
|
pos |
3D coordinates |
|
relative_position |
The relative position of the atom to the nearest attachment |
|
seq |
SMILES sequence features |
|
desc |
Default descriptor features ( |
|
rdkit2d |
RDKit 2D descriptors |
|
ecfp4 |
ECFP4 fingerprints |
|
rdkit3d |
RDKit 3D descriptors |
|
mordred |
Mordred descriptors |
|
maccs |
MACCS keys |
|
oligomer_rdkit2d |
RDKit 2D descriptors of the oligomer |
|
oligomer_mordred |
Mordred descriptors of the oligomer |
|
oligomer_ecfp4 |
ECFP4 fingerprints of the oligomer |
|
xenonpy_desc |
XenonPy composition descriptors |
|
mordred3d |
Mordred 3D descriptors |
|
fedors_density |
Estimated density from Fedors method |
|
monomer |
Preprocess molecule as monomer (remove attachment points) |
|
polycl |
Pretrained Polycl embeddings |
|
polybert |
Pretrained PolyBERT embeddings |
|
gaff2_mod |
Pretrained GAFF2 descriptors |
- class polymon.data.featurizer.AtomFeaturizer(feature_names: List[str] = None, unique_atom_nums: List[int] = None, unique_sources: List[str] = None)[source]
Featurize atoms in a molecule. Default features can be found in
DEFAULT_ATOM_FEATURES. We provide the following features:degree: The degree of the atom.is_aromatic: Whether the atom is aromatic.chiral_tag: The chiral tag of the atom.num_hydrogens: The number of implicit hydrogens.hybridization: The hybridization of the atom.mass: The mass of the atom.formal_charge: The formal charge of the atom.is_attachment: Whether the atom is an attachment point.xenonpy_atom: The xenonpy atom features.cgcnn: The cgcnn atom features.source: The source of the molecule.
- Parameters:
feature_names (List[str]) – The features to use.
unique_atom_nums (List[int]) – The unique atomic numbers. If
xinfeature_names, this parameter is required.
The features are assigned to
xin the output dictionary.- __call__(rdmol: Mol) Dict[str, Tensor][source]
Featurize the atoms in a molecule.
- Parameters:
rdmol (Chem.Mol) – The molecule to featurize.
- Returns:
The dictionary containing the features. The key is
xand the value is the concatenated features with the shape[num_atoms, num_features].- Return type:
Dict[str, torch.Tensor]
- class polymon.data.featurizer.BondFeaturizer(feature_names: List[str] = None)[source]
Featurize bonds in a molecule. This featurizer is used to obtain the edge index and edge attributes. We provide the following features:
fully_connected_edges: The fully connected edges, which means all atoms are connected to each other.bond: The chemical bond and their features.periodic_bond: This feature is used to add bonds between attachment points to the chemical bonds.virtual_bond: This feature is used to add virtual bonds between virtual node and each other node.
- Parameters:
feature_names (List[str]) – The features to use.
- __call__(rdmol: Mol) Dict[str, Tensor][source]
Featurize the bonds in a molecule.
- Parameters:
rdmol (Chem.Mol) – The molecule to featurize.
- Returns:
The dictionary containing the features. The key is
edge_indexandedge_attrwith the shape[2, num_edges]and[num_edges, num_features]respectively.- Return type:
Dict[str, torch.Tensor]
- class polymon.data.featurizer.PosFeaturizer(feature_names: List[str] = None)[source]
Featurize the positions of the atoms in a molecule. We first transform the polymer SMILES to a monomer SMILES and then embed the monomer SMILES with RDKit and conduct geometry optimization with MMFF94s. The positions are then normalized to the center of mass of the molecule.
Note
This featurizer may fail to embed the molecule.
- class polymon.data.featurizer.AtomNumFeaturizer(feature_names: List[str] = None)[source]
Featurize the atomic numbers of the atoms in a molecule. Although the atomic numbers are already included in the
xfeature, this featurizer is used to obtain the integer atomic numbers.
- class polymon.data.featurizer.RelativePositionFeaturizer(feature_names: List[str] = None)[source]
Featurize the relative positions of the atoms in a molecule. The relative position indicates the distance between the atom and the closest attachment point. The value is 200 if there is no attachment point.
- __call__(rdmol: Mol) Dict[str, Tensor][source]
Featurize the relative positions of the atoms in a molecule.
- Parameters:
rdmol (Chem.Mol) – The molecule to featurize.
- Returns:
The dictionary containing the features. The key is
relative_positionwith the shape[num_atoms].- Return type:
Dict[str, torch.Tensor]
- class polymon.data.featurizer.SeqFeaturizer(feature_names: List[str] = None)[source]
Featurize the sequence of a molecule. We first replace the double letters with a single token. The sequence is then encoded with the
SMILES_VOCAB. The sequence is then padded to the length ofMAX_SEQ_LEN. The SOS, EOS, and PAD tokens are added to the sequence.- __call__(rdmol: Mol) Dict[str, Tensor][source]
Featurize the sequence of a molecule.
- Parameters:
rdmol (Chem.Mol) – The molecule to featurize.
- Returns:
- The dictionary containing the features.
The key is
seqandseq_lenwith the shape[1, seq_len]and[1]respectively.
- Return type:
Dict[str, torch.Tensor]
- class polymon.data.featurizer.DescFeaturizer(feature_names: List[str] = None)[source]
Featurize descriptors of a molecule. Features shape should be
[1, num_features]. We provide the following features:rdkit2d: The RDKit 2D descriptors.ecfp4: The ECFP4 fingerprints.rdkit3d: The RDKit 3D descriptors.mordred: The Mordred descriptors.maccs: The MACCS keys.oligomer_rdkit2d: The RDKit 2D descriptors of the oligomer.oligomer_mordred: The Mordred descriptors of the oligomer.oligomer_ecfp4: The ECFP4 fingerprints of the oligomer.xenonpy_desc: The xenonpy descriptors.mordred3d: The Mordred 3D descriptors.fedors_density: The Fedors density.
- Parameters:
feature_names (List[str]) – The features to use.
polymer
- class polymon.data.polymer.OligomerBuilder[source]
Bases:
objectBuilder for oligomers. We use the following reactions to build the oligomers:
[:1][Au].[:2][Cu]>>[:1][:2]
[:1]=[Au].[:2]=[Cu]>>[:1]=[:2]
Note
The builder only supports for the polymer with two attachment points.
- Returns:
The oligomer molecule
- Return type:
Chem.Mol
- class polymon.data.polymer.Polymer(x: Tensor | None = None, edge_index: Tensor | None = None, edge_attr: Tensor | None = None, attachments: Tensor | None = None, z: Tensor | None = None, pos: Tensor | None = None, seq: Tensor | None = None, seq_len: Tensor | None = None, descriptors: Tensor | None = None, y: Tensor | None = None, smiles: str | None = None, identifier: str | None = None, source: str | None = None, **kwargs)[source]
Bases:
DataData object for multi-modal representation of polymers.
- Graph (2D/3D) Attributes:
x: Node features (num_nodes, num_node_features)
edge_index: Edge indices (2, num_edges)
edge_attr: Edge features (num_edges, num_edge_features)
attachments: Attachments (num_nodes, ) with 1 for attachment nodes
z: Atomic numbers (num_nodes, )
pos: Positions (num_nodes, 3)
- Sequence Attributes:
seq: Sequence (1, max_seq_len)
seq_len: Sequence length (1, )
- Descriptor Attributes:
descriptors: Descriptors (1, num_descriptors)
- Other Attributes:
y: Target (num_targets, )
smiles: a SMILES string
identifier: a unique identifier for the polymer
- property num_atoms: int
The number of atoms in the polymer.
- property num_bonds: int
The number of bonds in the polymer.
- property num_descriptors: int
The number of descriptors in the polymer.