polymon.data

dataset

class polymon.data.dataset.PolymerDataset(raw_csv_path: str, feature_names: List[str], label_column: str, sources: List[str], smiles_column: str = 'SMILES', identifier_column: str = 'id', save_processed: bool = True, force_reload: bool = False, add_hydrogens: bool = False, pre_transform: Callable | None = None, estimator: Callable | None = None, must_keep: List[str] = None)[source]

Bases: Dataset

A dataset that contains the polymers. During the initialization, the SMILES strings will be featurized to the corresponding features and converted to the Polymer object. If pre_transform is provided, it will be used to update the Polymer object. If estimator is provided, it will be used to estimate the labels of the Polymer object, and make the prediction task be the residual between the ground truth and the estimated labels.

Parameters:
  • raw_csv_path (str) – The path to the raw csv file. Typically, it should have at least three columns: SMILES, label (e.g., Rg), and Source (e.g., PI1070).

  • feature_names (List[str]) – The names of the features to use. The available features are in AVAIL_FEATURES from polymon.data.featurizer.

  • label_column (str) – Label column name.

  • sources (List[str]) – The names of the sources.

  • smiles_column (str) – SMILES column name.

  • identifier_column (str) – Identifier column name.

  • save_processed (bool) – Whether to save the processed dataset.

  • force_reload (bool) – Whether to force reload the processed dataset.

  • add_hydrogens (bool) – Whether to add hydrogens to the molecules.

  • pre_transform (Callable) – The pre-transform to apply to the data. It should be a function that takes a Polymer object and returns a Polymer object.

  • estimator (Callable) – The estimator to apply to the data. It is used to provide the estimated labels for the Polymer object. It should be the object of polymon.estimator.BaseEstimator.

get(idx: int) → Polymer[source]

Get the Polymer object at the given index.

get_loaders(batch_size: int, n_train: int | float, n_val: int | float, mode: Literal['random', 'scaffold'] = 'random', num_workers: int = 0, augmentation: bool = False) → Tuple[DataLoader, DataLoader, DataLoader][source]

Get the data loaders for the training, validation, and test sets.

Parameters:
  • batch_size (int) – The batch size.

  • n_train (Union[int, float]) – The number of training samples. If it is a float, it will be converted to an integer by multiplying the length of the dataset.

  • n_val (Union[int, float]) – The number of validation samples. If it is a float, it will be converted to an integer by multiplying the length of the dataset.

  • mode (Literal['random', 'scaffold']) – The split mode. random: Split the dataset randomly. scaffold: Split the dataset by scaffold.

  • num_workers (int) – The number of workers for the data loaders.

  • augmentation (bool) – Whether to augment the training set. Currently, it only supports the augmentation of oligomers.

len() → int[source]

Get the number of Polymer objects in the dataset.

sample_batch(batch_size: int = 64) → Batch[source]

Sample a batch of Polymer objects.

dedup

class polymon.data.dedup.Dedup(df: DataFrame, label: str, must_keep: List[str] = None, rtol: float = 0.05)[source]

Bases: object

add_df(file_path: str, smiles_col: str, label_col: str, source: str)[source]

Add a dataframe to the self.df.

Parameters:
  • file_path (str) – The path to the dataframe.

  • smiles_col (str) – The column name of the SMILES.

  • label_col (str) – The column name of the label.

  • source (str) – The source name.

compare(source1: str, source2: str, fitting: bool = False)[source]

Compare the label of the two sources with the same smiles.

Parameters:
  • source1 (str) – The first source name.

  • source2 (str) – The second source name.

  • fitting (bool) – Whether to fit the data. If True, the function will fit the data with a linear regression and add the fitted points as a new source.

run(sources: List[str] | None = None, save: bool = False)[source]

Run the deduplication. The deduplication is performed by the following steps:

  1. Remove rows with higher relative difference.

  2. Keep indices if the source is in self.must_keep.

Parameters:
  • sources (Optional[List[str]]) – The sources to deduplicate.

  • save (bool) – Whether to save the deduplicated dataframe.

show_distribution(sources: List[str], bins: int = 50, alpha: float = 0.5, density: bool = False)[source]

Plot the distribution of the label for each source.

Parameters:
  • sources (List[str]) – The sources to plot.

  • bins (int) – The number of bins.

  • alpha (float) – The alpha of the histogram.

  • density (bool) – Whether to plot the density.

featurizer

Summary of Supported Featurizers

Name

Featurizer

Attributes in Polymer

x

AtomFeaturizer

x

edge

BondFeaturizer

edge_index and edge_attr

pos

PosFeaturizer

pos

z

AtomNumFeaturizer

z

relative_position

RelativePositionFeaturizer

relative_position

seq

SeqFeaturizer

seq and seq_len

desc

DescFeaturizer

descriptors

monomer

RDMolPreprocessor

monomer

Available Feature Names

Feature Name

Available Features

Description

x

Node features (default set, polymon.setting.DEFAULT_ATOM_FEATURES).

xenonpy_atom

XenonPy atom features

cgcnn

CGCNN atom features

source

Source features

z

Atomic numbers as integers

edge

Default edge features (bond)

bond

Chemical bond features

fully_connected_edges

Fully connected edge indices

periodic_bond

Add bonds between attachment points to the chemical bonds

virtual_bond

Add virtual bonds between virtual node and each other node

pos

3D coordinates

relative_position

The relative position of the atom to the nearest attachment

seq

SMILES sequence features

desc

Default descriptor features (rdkit2d)

rdkit2d

RDKit 2D descriptors

ecfp4

ECFP4 fingerprints

rdkit3d

RDKit 3D descriptors

mordred

Mordred descriptors

maccs

MACCS keys

oligomer_rdkit2d

RDKit 2D descriptors of the oligomer

oligomer_mordred

Mordred descriptors of the oligomer

oligomer_ecfp4

ECFP4 fingerprints of the oligomer

xenonpy_desc

XenonPy composition descriptors

mordred3d

Mordred 3D descriptors

fedors_density

Estimated density from Fedors method

monomer

Preprocess molecule as monomer (remove attachment points)

polycl

Pretrained Polycl embeddings

polybert

Pretrained PolyBERT embeddings

gaff2_mod

Pretrained GAFF2 descriptors

class polymon.data.featurizer.AtomFeaturizer(feature_names: List[str] = None, unique_atom_nums: List[int] = None, unique_sources: List[str] = None)[source]

Featurize atoms in a molecule. Default features can be found in DEFAULT_ATOM_FEATURES. We provide the following features:

  • degree: The degree of the atom.

  • is_aromatic: Whether the atom is aromatic.

  • chiral_tag: The chiral tag of the atom.

  • num_hydrogens: The number of implicit hydrogens.

  • hybridization: The hybridization of the atom.

  • mass: The mass of the atom.

  • formal_charge: The formal charge of the atom.

  • is_attachment: Whether the atom is an attachment point.

  • xenonpy_atom: The xenonpy atom features.

  • cgcnn: The cgcnn atom features.

  • source: The source of the molecule.

Parameters:
  • feature_names (List[str]) – The features to use.

  • unique_atom_nums (List[int]) – The unique atomic numbers. If x in feature_names, this parameter is required.

The features are assigned to x in the output dictionary.

__call__(rdmol: Mol) → Dict[str, Tensor][source]

Featurize the atoms in a molecule.

Parameters:

rdmol (Chem.Mol) – The molecule to featurize.

Returns:

The dictionary containing the features. The key is x and the value is the concatenated features with the shape [num_atoms, num_features].

Return type:

Dict[str, torch.Tensor]

class polymon.data.featurizer.BondFeaturizer(feature_names: List[str] = None)[source]

Featurize bonds in a molecule. This featurizer is used to obtain the edge index and edge attributes. We provide the following features:

  • fully_connected_edges: The fully connected edges, which means all atoms are connected to each other.

  • bond: The chemical bond and their features.

  • periodic_bond: This feature is used to add bonds between attachment points to the chemical bonds.

  • virtual_bond: This feature is used to add virtual bonds between virtual node and each other node.

Parameters:

feature_names (List[str]) – The features to use.

__call__(rdmol: Mol) → Dict[str, Tensor][source]

Featurize the bonds in a molecule.

Parameters:

rdmol (Chem.Mol) – The molecule to featurize.

Returns:

The dictionary containing the features. The key is edge_index and edge_attr with the shape [2, num_edges] and [num_edges, num_features] respectively.

Return type:

Dict[str, torch.Tensor]

class polymon.data.featurizer.PosFeaturizer(feature_names: List[str] = None)[source]

Featurize the positions of the atoms in a molecule. We first transform the polymer SMILES to a monomer SMILES and then embed the monomer SMILES with RDKit and conduct geometry optimization with MMFF94s. The positions are then normalized to the center of mass of the molecule.

Note

This featurizer may fail to embed the molecule.

__call__(rdmol: Mol) → Dict[str, Tensor][source]

Featurize the positions of the atoms in a molecule.

Parameters:

rdmol (Chem.Mol) – The molecule to featurize.

Returns:

The dictionary containing the features. The key is pos with the shape [num_atoms, 3].

Return type:

Dict[str, torch.Tensor]

class polymon.data.featurizer.AtomNumFeaturizer(feature_names: List[str] = None)[source]

Featurize the atomic numbers of the atoms in a molecule. Although the atomic numbers are already included in the x feature, this featurizer is used to obtain the integer atomic numbers.

__call__(rdmol: Mol) → Dict[str, Tensor][source]

Featurize the atomic numbers of the atoms in a molecule.

Parameters:

rdmol (Chem.Mol) – The molecule to featurize.

Returns:

The dictionary containing the features. The key is z with the shape [num_atoms].

Return type:

Dict[str, torch.Tensor]

class polymon.data.featurizer.RelativePositionFeaturizer(feature_names: List[str] = None)[source]

Featurize the relative positions of the atoms in a molecule. The relative position indicates the distance between the atom and the closest attachment point. The value is 200 if there is no attachment point.

__call__(rdmol: Mol) → Dict[str, Tensor][source]

Featurize the relative positions of the atoms in a molecule.

Parameters:

rdmol (Chem.Mol) – The molecule to featurize.

Returns:

The dictionary containing the features. The key is relative_position with the shape [num_atoms].

Return type:

Dict[str, torch.Tensor]

class polymon.data.featurizer.SeqFeaturizer(feature_names: List[str] = None)[source]

Featurize the sequence of a molecule. We first replace the double letters with a single token. The sequence is then encoded with the SMILES_VOCAB. The sequence is then padded to the length of MAX_SEQ_LEN. The SOS, EOS, and PAD tokens are added to the sequence.

__call__(rdmol: Mol) → Dict[str, Tensor][source]

Featurize the sequence of a molecule.

Parameters:

rdmol (Chem.Mol) – The molecule to featurize.

Returns:

The dictionary containing the features.

The key is seq and seq_len with the shape [1, seq_len] and [1] respectively.

Return type:

Dict[str, torch.Tensor]

class polymon.data.featurizer.DescFeaturizer(feature_names: List[str] = None)[source]

Featurize descriptors of a molecule. Features shape should be [1, num_features]. We provide the following features:

  • rdkit2d: The RDKit 2D descriptors.

  • ecfp4: The ECFP4 fingerprints.

  • rdkit3d: The RDKit 3D descriptors.

  • mordred: The Mordred descriptors.

  • maccs: The MACCS keys.

  • oligomer_rdkit2d: The RDKit 2D descriptors of the oligomer.

  • oligomer_mordred: The Mordred descriptors of the oligomer.

  • oligomer_ecfp4: The ECFP4 fingerprints of the oligomer.

  • xenonpy_desc: The xenonpy descriptors.

  • mordred3d: The Mordred 3D descriptors.

  • fedors_density: The Fedors density.

Parameters:

feature_names (List[str]) – The features to use.

__call__(rdmol: Mol) → Dict[str, Tensor][source]

Featurize descriptors of a molecule.

Parameters:

rdmol (Chem.Mol) – The molecule to featurize.

Returns:

The dictionary containing the features. The key is descriptors with the shape [1, num_features].

Return type:

Dict[str, torch.Tensor]

class polymon.data.featurizer.RDMolPreprocessor[source]

Preprocess the molecule. We provide the following preprocessors:

  • monomer: Remove the attachment points and set the attachment points

to the neighbors of the attachment points.

static monomer(rdmol: Mol) → Mol[source]

Preprocess the molecule.

Parameters:

rdmol (Chem.Mol) – The molecule to preprocess.

Returns:

The preprocessed molecule.

Return type:

Chem.Mol

polymer

class polymon.data.polymer.OligomerBuilder[source]

Bases: object

Builder for oligomers. We use the following reactions to build the oligomers:

  • [:1][Au].[:2][Cu]>>[:1][:2]

  • [:1]=[Au].[:2]=[Cu]>>[:1]=[:2]

Note

The builder only supports for the polymer with two attachment points.

Returns:

The oligomer molecule

Return type:

Chem.Mol

static get_oligomer(smiles: str, n_oligomer: int) → Mol[source]

Get the oligomer molecule from the smiles string.

Parameters:
  • smiles – The smiles string of the polymer

  • n_oligomer – The number of oligomers

Returns:

The oligomer molecule

Return type:

Chem.Mol

class polymon.data.polymer.Polymer(x: Tensor | None = None, edge_index: Tensor | None = None, edge_attr: Tensor | None = None, attachments: Tensor | None = None, z: Tensor | None = None, pos: Tensor | None = None, seq: Tensor | None = None, seq_len: Tensor | None = None, descriptors: Tensor | None = None, y: Tensor | None = None, smiles: str | None = None, identifier: str | None = None, source: str | None = None, **kwargs)[source]

Bases: Data

Data object for multi-modal representation of polymers.

Graph (2D/3D) Attributes:
  • x: Node features (num_nodes, num_node_features)

  • edge_index: Edge indices (2, num_edges)

  • edge_attr: Edge features (num_edges, num_edge_features)

  • attachments: Attachments (num_nodes, ) with 1 for attachment nodes

  • z: Atomic numbers (num_nodes, )

  • pos: Positions (num_nodes, 3)

Sequence Attributes:
  • seq: Sequence (1, max_seq_len)

  • seq_len: Sequence length (1, )

Descriptor Attributes:
  • descriptors: Descriptors (1, num_descriptors)

Other Attributes:
  • y: Target (num_targets, )

  • smiles: a SMILES string

  • identifier: a unique identifier for the polymer

property num_atoms: int

The number of atoms in the polymer.

property num_bonds: int

The number of bonds in the polymer.

property num_descriptors: int

The number of descriptors in the polymer.