polymon.model

ModelWrapper

class polymon.model.base.ModelWrapper(model: BaseModel, normalizer: Normalizer, featurizer: ComposeFeaturizer, transform_cls: str = None, transform_kwargs: Dict[str, Any] = None, estimator: BaseEstimator = None)[source]

Bases: Module

Model Wrapper. This wrapper is used to wrap the model, normalizer, featurizer, transform, and estimator, and make it easier to do inference.

Parameters:
  • model (BaseModel) – The model.

  • normalizer (Normalizer) – The normalizer.

  • featurizer (ComposeFeaturizer) – The featurizer.

  • transform_cls (str) – The class of the transform.

  • transform_kwargs (Dict[str, Any]) – The initial parameters of the transform.

  • estimator (BaseEstimator) – The estimator.

forward(batch: Batch, loss_fn: Module, device: str = 'cuda') → Tensor[source]

Forward pass.

Parameters:
  • batch (Batch) – The batch of polymers.

  • loss_fn (nn.Module) – The loss function.

  • device (str) – The device to use.

Returns:

The loss.

Return type:

torch.Tensor

classmethod from_dict(model_info: Dict[str, Any]) → ModelWrapper[source]

Build a ModelWrapper from a dictionary.

Parameters:

model_info (Dict[str, Any]) – The information of the model.

Returns:

The ModelWrapper.

Return type:

‘ModelWrapper’

classmethod from_file(path: str, map_location: str = 'cpu', weights_only: bool = False) → ModelWrapper[source]

Build a ModelWrapper from a file.

Parameters:
  • path (str) – The path to the file.

  • map_location (str) – The map location.

  • weights_only (bool) – Whether to load only the weights.

Returns:

The ModelWrapper.

Return type:

‘ModelWrapper’

property info: Dict[str, Any]

Get the information of the model.

Returns:

The information of the model.

Return type:

Dict[str, Any]

predict(smiles_list: List[str], batch_size: int = 128, device: str = 'cpu', backup_model: ModelWrapper = None) → Tensor[source]

Predict the output of the model for a list of polymer SMILES strings.

Parameters:
  • smiles_list (List[str]) – The list of SMILES strings.

  • batch_size (int) – The batch size.

  • device (str) – The device to use.

  • backup_model (ModelWrapper) – The backup model. If the model fails to predict the output for a polymer, the backup model will be used to predict the output.

Returns:

The output of the model.

Return type:

torch.Tensor

predict_batch(batch: Batch, device: str = 'cuda') → Tensor[source]

Predict the output of the model for a batch of polymers.

Parameters:
  • batch (Batch) – The batch of polymers.

  • device (str) – The device to use.

Returns:

The output of the model.

Return type:

torch.Tensor

write(path: str) → str[source]

Write the model to a file.

Parameters:

path (str) – The path to the file.

Returns:

The path to the file.

Return type:

str

KFoldModel

class polymon.model.base.KFoldModel(model_cls: str, model_init_params: Dict[str, Any], n_fold: int = 5)[source]

Bases: BaseModel

K-Fold Model. The output is the average of the predictions of the models trained on the different folds.

Parameters:
  • model_cls (str) – The class of the model.

  • model_init_params (Dict[str, Any]) – The initial parameters of the model.

  • n_fold (int) – The number of folds.

forward(batch: Polymer) → Tensor[source]

Forward pass. The output is the predictions of k-fold models stacked. shape: (n_polymers, n_folds)

Parameters:

batch (Polymer) – The batch of polymers.

Returns:

The output of the model.

Return type:

torch.Tensor

classmethod from_models(models: List[ModelWrapper]) → KFoldModel[source]

Build a K-Fold Model from a list of models.

Parameters:

models (List['ModelWrapper']) – The models.

Returns:

The K-Fold Model.

Return type:

‘KFoldModel’

property init_params: Dict[str, Any]

Get the initial parameters of the model.

Returns:

The initial parameters of the model.

Return type:

Dict[str, Any]

LinearEnsembleRegressor

EnsembleModelWrapper

Models

Available Models in polymon.model

Model Type

Class Name

Description

gatv2

GATv2

Graph Attention Network v2

attentivefp

AttentiveFPWrapper

AttentiveFP

dimenetpp

DimeNetPP

DimeNet++

gatv2vn

GATv2VirtualNode

GATv2 with virtual node

gin

GIN

Graph Isomorphism Network

pna

PNA

Principal Neighbourhood Aggregation

gvp

GVPModel

Geometric Vector Perceptron

gatv2chainreadout

GATv2ChainReadout

GATv2 with chain readout

gt

GraphTransformer

Graph Transformer

kan_gatv2

KAN_GATv2

KAN-augmented GATv2

gps

GraphGPS

Graph GPS

kan_gps

KAN_GPS

KAN-augmented GraphGPS

fastkan

FastKANWrapper

Fast KAN for descriptors

efficientkan

EfficientKANWrapper

Efficient KAN for descriptors

fourierkan

FourierKANWrapper

Fourier KAN for descriptors

fastkan_gatv2

FastKAN_GATv2

FastKAN-augmented GATv2

gatv2_lineevo

GATv2LineEvo

GATv2 with line evolution

gatv2_sage

GATv2SAGE

GATv2 with SAGE aggregation

gatv2_source

GATv2_Source

GATv2 for multi-fidelity/source

gatv2_pe

GATv2_PE

GATv2 with position encoding

gatv2_embed_residual

GATv2EmbedResidual

GATv2 with embedding residuals

kan_gin

KAN_GIN

KAN-augmented GIN

fastkan_gin

FastKAN_GIN

FastKAN-augmented GIN

kan_gcn

KAN_GCN

KAN-augmented GCN

dmpnn

DMPNN

Directed Message Passing Neural Network

kan_dmpnn

KAN_DMPNN

KAN-augmented DMPNN

Note

The model_type string is used as the key in configuration and when calling polymon.model.build_model().

gatv2

class polymon.model.gnn.GATv2(num_atom_features: int, hidden_dim: int, num_layers: int, num_heads: int = 8, pred_hidden_dim: int = 128, pred_dropout: float = 0.2, pred_layers: int = 2, activation: str = 'prelu', num_tasks: int = 1, bias: bool = True, dropout: float = 0.1, edge_dim: int = None, num_descriptors: int = 0)[source]

Bases: BaseModel

GATv2 model.

Parameters:
  • num_atom_features (int) – The number of atom features.

  • hidden_dim (int) – The number of hidden dimensions.

  • num_layers (int) – The number of layers.

  • num_heads (int) – The number of heads. Default to 8.

  • pred_hidden_dim (int) – The number of hidden dimensions in the prediction layer. Default to 128.

  • pred_dropout (float) – The dropout rate in the prediction layer. Default to 0.2.

  • pred_layers (int) – The number of layers in the prediction layer. Default to 2.

  • activation (str) – The activation function. Default to prelu.

  • num_tasks (int) – The number of tasks. Default to 1.

  • bias (bool) – Whether to use bias in the GATv2Conv layers. Default to True.

  • dropout (float) – The dropout rate in the GATv2Conv layers. Default to 0.1.

  • edge_dim (int) – The number of edge features. Default to None.

  • num_descriptors (int) – The number of descriptors. If not zero, the descriptors will be concatenated to the output of the model. Default to 0.

forward(batch: Polymer)[source]

Forward pass.

Parameters:

batch (Polymer) – The batch of data.

Returns:

The output of the model.

Return type:

torch.Tensor

get_embeddings(batch: Polymer)[source]

Get global graph embeddings.

Parameters:

batch (Polymer) – The batch of data.

Returns:

The embeddings of the batch.

Return type:

torch.Tensor

attentivefp

class polymon.model.gnn.AttentiveFPWrapper(in_channels: int, hidden_dim: int, edge_dim: int, num_layers: int, out_channels: int = 1, num_timesteps: int = 2)[source]

Bases: BaseModel

AttentiveFP model wrapper.

Parameters:
  • in_channels (int) – The number of input channels.

  • hidden_dim (int) – The number of hidden dimensions.

  • edge_dim (int) – The number of edge features.

  • num_layers (int) – The number of layers.

  • out_channels (int) – The number of output channels. Default to 1.

  • num_timesteps (int) – The number of timesteps. Default to 2.

forward(batch: Polymer)[source]

Forward pass.

Parameters:

batch (Polymer) – The batch of data.

Returns:

The output of the model.

Return type:

torch.Tensor

dimenetpp

class polymon.model.gnn.DimeNetPP(hidden_dim: int = 128, out_channels: int = 1, num_layers: int = 3, int_emb_size: int = 64, basis_emb_size: int = 8, out_emb_channels: int = 256, num_spherical: int = 7, num_radial: int = 6, cutoff: float = 5.0, max_num_neighbors: int = 32, envelope_exponent: int = 5, num_before_skip: int = 1, num_after_skip: int = 2, num_output_layers: int = 2, act: str = 'swish', output_initializer: str = 'zeros')[source]

Bases: DimeNetPlusPlus, BaseModel

DimeNet++ model wrapper.

No-index:

Parameters:
  • hidden_dim (int) – The number of hidden dimensions. Default to 128.

  • out_channels (int) – The number of output channels. Default to 1.

  • num_layers (int) – The number of layers. Default to 3.

  • int_emb_size (int) – The number of embedding dimensions for the integer features. Default to 64.

  • basis_emb_size (int) – The number of embedding dimensions for the basis features. Default to 8.

  • out_emb_channels (int) – The number of output embedding channels. Default to 256.

  • num_spherical (int) – The number of spherical harmonics. Default to 7.

  • num_radial (int) – The number of radial basis functions. Default to 6.

  • cutoff (float) – The cutoff radius. Default to 5.0.

  • max_num_neighbors (int) – The maximum number of neighbors. Default to 32.

  • envelope_exponent (int) – The exponent of the envelope function. Default to 5.

  • num_before_skip (int) – The number of layers before skip connections. Default to 1.

  • num_after_skip (int) – The number of layers after skip connections. Default to 2.

  • num_output_layers (int) – The number of output layers. Default to 2.

  • act (str) – The activation function. Default to swish.

  • output_initializer (str) – The initializer for the output layer. Default to zeros.

forward(data: Polymer)[source]

Forward pass.

Parameters:

data (Polymer) – The batch of data.

Returns:

The output of the model.

Return type:

torch.Tensor

gatv2vn

class polymon.model.gnn.GATv2VirtualNode(num_atom_features: int, hidden_dim: int, num_layers: int, num_heads: int = 8, pred_hidden_dim: int = 128, pred_dropout: float = 0.2, pred_layers: int = 2, activation: str = 'prelu', num_tasks: int = 1, bias: bool = True, dropout: float = 0.1, edge_dim: int = None, num_descriptors: int = 0)[source]

Bases: BaseModel

GATv2VirtualNode model. Add virtual node as the graph node and use its features as the graph embedding.

Parameters:
  • num_atom_features (int) – The number of atom features.

  • hidden_dim (int) – The number of hidden dimensions.

  • num_layers (int) – The number of layers.

  • num_heads (int) – The number of heads. Default to 8.

  • pred_hidden_dim (int) – The number of hidden dimensions in the prediction layer. Default to 128.

  • pred_dropout (float) – The dropout rate in the prediction layer. Default to 0.2.

  • pred_layers (int) – The number of layers in the prediction layer. Default to 2.

  • activation (str) – The activation function. Default to prelu.

  • num_tasks (int) – The number of tasks. Default to 1.

  • bias (bool) – Whether to use bias in the GATv2Conv layers. Default to True.

  • dropout (float) – The dropout rate in the GATv2Conv layers. Default to 0.1.

  • edge_dim (int) – The number of edge features. Default to None.

  • num_descriptors (int) – The number of descriptors. If not zero, the descriptors will be concatenated to the output of the model. Default to 0.

forward(batch: Polymer)[source]

Forward pass.

Parameters:

batch (Polymer) – The batch of data. It should have descriptors attribute.

Returns:

The output of the model.

Return type:

torch.Tensor

gin

class polymon.model.gnn.GIN(num_atom_features: int, hidden_dim: int, num_layers: int, dropout: float = 0.2, n_mlp_layers: int = 2, pred_hidden_dim: int = 128, pred_dropout: float = 0.2, pred_layers: int = 2)[source]

Bases: BaseModel

GIN model wrapper.

Parameters:
  • num_atom_features (int) – The number of atom features.

  • hidden_dim (int) – The number of hidden dimensions.

  • num_layers (int) – The number of layers.

  • dropout (float) – The dropout rate. Default to 0.2.

  • n_mlp_layers (int) – The number of layers in the MLP. Default to 2.

  • pred_hidden_dim (int) – The number of hidden dimensions in the prediction layer. Default to 128.

  • pred_dropout (float) – The dropout rate in the prediction layer. Default to 0.2.

  • pred_layers (int) – The number of layers in the prediction layer. Default to 2.

forward(batch: Polymer)[source]

Forward pass.

Parameters:

batch (Polymer) – The batch of data.

Returns:

The output of the model.

Return type:

torch.Tensor

pna

class polymon.model.gnn.PNA(in_channels: int, hidden_dim: int, num_layers: int, deg: Tensor, towers: int = 1, edge_dim: int = None, pred_hidden_dim: int = 128, pred_dropout: float = 0.2, pred_layers: int = 2)[source]

Bases: BaseModel

PNA model wrapper.

Parameters:
  • in_channels (int) – The number of input channels.

  • hidden_dim (int) – The number of hidden dimensions.

  • num_layers (int) – The number of layers.

  • deg (torch.Tensor) – The degree tensor.

  • towers (int) – The number of towers. Default to 1.

  • edge_dim (int) – The number of edge features. Default to None.

  • pred_hidden_dim (int) – The number of hidden dimensions in the prediction layer. Default to 128.

  • pred_dropout (float) – The dropout rate in the prediction layer. Default to 0.2.

  • pred_layers (int) – The number of layers in the prediction layer. Default to 2.

forward(batch: Polymer)[source]

Forward pass.

Parameters:

batch (Polymer) – The batch of data. It should have been preprocessed by polymon.data.utils.AddRandomWalkPE.

Returns:

The output of the model.

Return type:

torch.Tensor

gvp

gatv2chainreadout

class polymon.model.gatv2.gat_chain_readout.GATv2ChainReadout(num_atom_features: int, hidden_dim: int, num_layers: int, num_heads: int = 8, pred_hidden_dim: int = 128, pred_dropout: float = 0.2, pred_layers: int = 2, activation: str = 'prelu', num_tasks: int = 1, bias: bool = True, dropout: float = 0.1, edge_dim: int = None, num_descriptors: int = 0, chain_length: int = 10)[source]

Bases: BaseModel

GATv2 with chain readout.

Parameters:
  • num_atom_features (int) – The number of atom features.

  • hidden_dim (int) – The number of hidden dimensions.

  • num_layers (int) – The number of layers.

  • num_heads (int) – The number of heads. Default to 8.

  • pred_hidden_dim (int) – The number of hidden dimensions for the prediction MLP. Default to 128.

  • pred_dropout (float) – The dropout rate for the prediction MLP. Default to 0.2.

  • pred_layers (int) – The number of layers for the prediction MLP. Default to 2.

  • activation (str) – The activation function. Default to 'prelu'.

  • num_tasks (int) – The number of tasks. Default to 1.

  • bias (bool) – Whether to use bias. Default to True.

  • dropout (float) – The dropout rate. Default to 0.1.

  • edge_dim (int) – The number of edge dimensions.

  • num_descriptors (int) – The number of descriptors. Default to 0.

  • chain_length (int) – The length of the chain. Default to 10.

forward(batch: Polymer)[source]

Forward pass.

Parameters:

batch (Polymer) – The batch of data.

Returns:

The output tensor.

Return type:

torch.Tensor

gt

class polymon.model.gnn.GraphTransformer(in_channels: int, hidden_dim: int, num_layers: int, num_heads: int = 8, dropout: float = 0.2, pred_hidden_dim: int = 128, pred_dropout: float = 0.2, pred_layers: int = 2)[source]

Bases: BaseModel

GraphTransformer model wrapper.

Parameters:
  • in_channels (int) – The number of input channels.

  • hidden_dim (int) – The number of hidden dimensions.

  • num_layers (int) – The number of layers.

  • num_heads (int) – The number of heads. Default to 8.

  • dropout (float) – The dropout rate. Default to 0.2.

  • pred_hidden_dim (int) – The number of hidden dimensions in the prediction layer. Default to 128.

  • pred_dropout (float) – The dropout rate in the prediction layer. Default to 0.2.

  • pred_layers (int) – The number of layers in the prediction layer. Default to 2.

forward(batch: Polymer)[source]

Forward pass.

Parameters:

batch (Polymer) – The batch of data.

Returns:

The output of the model.

Return type:

torch.Tensor

kan_gatv2

class polymon.model.gatv2.kan_gatv2.KAN_GATv2(num_node_features: int, hidden_dim: int, num_layers: int, num_heads: int = 8, grid_size: int = 3, dropout: float = 0.1, pred_hidden_dim: int = 128, pred_dropout: float = 0.2, pred_layers: int = 2)[source]

Bases: BaseModel

KAN-augmented GATv2.

Parameters:
  • num_node_features (int) – The number of node features.

  • hidden_dim (int) – The number of hidden dimensions.

  • num_layers (int) – The number of layers.

  • num_heads (int) – The number of heads. Default to 8.

  • grid_size (int) – The size of the grid. Default to 3.

  • dropout (float) – The dropout rate. Default to 0.1.

  • pred_hidden_dim (int) – The number of hidden dimensions for the prediction MLP. Default to 128.

  • pred_dropout (float) – The dropout rate for the prediction MLP. Default to 0.2.

  • pred_layers (int) – The number of layers for the prediction MLP. Default to 2.

forward(batch: Polymer)[source]

Forward pass.

Parameters:

batch (Polymer) – The batch of data.

Returns:

The output tensor.

Return type:

torch.Tensor

gps

class polymon.model.gps.gps.GraphGPS(in_channels: int, edge_dim: int, heads: int = 4, hidden_dim: int = 64, num_layers: int = 6, walk_length: int = 20, pe_dim: int = 8, attn_type: Literal['performer', 'multihead'] = 'multihead', attn_dropout: float = 0.0)[source]

Bases: BaseModel

GraphGPS model.

Parameters:
  • in_channels (int) – The number of input channels.

  • edge_dim (int) – The number of edge dimensions.

  • heads (int) – The number of heads. Default to 4.

  • hidden_dim (int) – The number of hidden dimensions. Default to 64.

  • num_layers (int) – The number of layers. Default to 6.

  • walk_length (int) – The length of the walk. Default to 20.

  • pe_dim (int) – The dimension of the positional encoding. Default to 8.

  • attn_type (Literal['performer', 'multihead']) – The type of attention. Default to 'multihead'.

  • attn_kwargs (Dict[str, Any]) – The keyword arguments for the attention.

  • grid_size (int) – The size of the grid. Default to 3.

forward(batch: Polymer)[source]

Forward pass.

Parameters:

batch (Polymer) – The batch of data. It should have pe attribute.

Returns:

The output tensor.

Return type:

torch.Tensor

kan_gps

class polymon.model.gps.gps.KAN_GPS(in_channels: int, edge_dim: int, heads: int = 4, hidden_dim: int = 64, num_layers: int = 6, walk_length: int = 20, pe_dim: int = 8, attn_type: Literal['performer', 'multihead', 'fastkan'] = 'fastkan', attn_dropout: float = 0.0, grid_size: int = 3)[source]

Bases: BaseModel

KAN-augmented GraphGPS model.

Parameters:
  • in_channels (int) – The number of input channels.

  • edge_dim (int) – The number of edge dimensions.

  • heads (int) – The number of heads. Default to 4.

  • hidden_dim (int) – The number of hidden dimensions. Default to 64.

  • num_layers (int) – The number of layers. Default to 6.

  • walk_length (int) – The length of the walk. Default to 20.

  • pe_dim (int) – The dimension of the positional encoding. Default to 8.

  • attn_type (Literal['performer', 'multihead', 'fastkan']) – The type of attention. Default to 'fastkan'.

  • attn_kwargs (Dict[str, Any]) – The keyword arguments for the attention.

  • grid_size (int) – The size of the grid. Default to 3.

forward(batch: Polymer)[source]

Forward pass.

Parameters:

batch (Polymer) – The batch of data. It should have pe attribute.

Returns:

The output tensor.

Return type:

torch.Tensor

fastkan

class polymon.model.kan.fast_kan.FastKANWrapper(in_channels: int, hidden_dim: int, num_layers: int, grid_min: float = -2.0, grid_max: float = 2.0, num_grids: int = 8)[source]

Bases: BaseModel

Fast KAN wrapper.

Parameters:
  • in_channels (int) – The number of input channels.

  • hidden_dim (int) – The number of hidden dimensions.

  • num_layers (int) – The number of layers.

  • grid_min (float) – The minimum value of the grid. Default to -2.0.

  • grid_max (float) – The maximum value of the grid. Default to 2.0.

  • num_grids (int) – The number of grids. Default to 8.

Note

The implementation is adapted from Fast KAN for descriptors.

forward(batch: Polymer)[source]

Forward pass.

Parameters:

batch (Polymer) – The batch of data. It should have descriptors attribute.

Returns:

The output tensor.

Return type:

torch.Tensor

efficientkan

class polymon.model.kan.efficient_kan.EfficientKANWrapper(in_channels: int, hidden_dim: int, num_layers: int, grid_size: int = 5, spline_order: int = 3, scale_noise: float = 0.1, scale_base: float = 1.0, scale_spline: float = 1.0, base_activation: ~torch.nn.modules.module.Module = <class 'torch.nn.modules.activation.SiLU'>, grid_eps: float = 0.02, grid_range: ~typing.List[float] = [-1, 1])[source]

Bases: BaseModel

Efficient KAN wrapper.

Parameters:
  • in_channels (int) – The number of input channels.

  • hidden_dim (int) – The number of hidden dimensions.

  • num_layers (int) – The number of layers.

  • grid_size (int) – The number of grid points. Default to 5.

  • spline_order (int) – The order of the spline. Default to 3.

  • scale_noise (float) – The scale of the noise. Default to 0.1.

  • scale_base (float) – The scale of the base. Default to 1.0.

  • scale_spline (float) – The scale of the spline. Default to 1.0.

  • base_activation (torch.nn.Module) – The activation function for the base. Default to torch.nn.SiLU.

  • grid_eps (float) – The epsilon for the grid. Default to 0.02.

  • grid_range (List[float]) – The range of the grid. Default to [-1, 1].

Note

The implementation is adapted from Fast KAN for descriptors.

forward(batch: Polymer)[source]

Forward pass.

Parameters:

batch (Polymer) – The batch of data. It should have descriptors attribute.

Returns:

The output tensor.

Return type:

torch.Tensor

fourierkan

class polymon.model.kan.fourier_kan.FourierKANWrapper(in_channels: int, hidden_dim: int, num_layers: int, grid_size: int = 5, add_bias: bool = True, add_act: bool = False)[source]

Bases: BaseModel

Fourier KAN wrapper.

Parameters:
  • in_channels (int) – The number of input channels.

  • hidden_dim (int) – The number of hidden dimensions.

  • num_layers (int) – The number of layers.

  • grid_size (int) – The number of grid points. Default to 5.

  • add_bias (bool) – Whether to add bias. Default to True.

  • add_act (bool) – Whether to add activation. Default to False.

Note

The implementation is adapted from Fourier KAN for descriptors.

forward(batch: Polymer)[source]

Forward pass.

Parameters:

batch (Polymer) – The batch of data. It should have descriptors attribute.

Returns:

The output tensor.

Return type:

torch.Tensor

fastkan_gatv2

class polymon.model.gatv2.fastkan_gatv2.FastKAN_GATv2(num_atom_features: int, hidden_dim: int, num_layers: int, num_heads: int = 8, pred_hidden_dim: int = 128, grid_min: float = -2.0, grid_max: float = 2.0, num_grids: int = 8, num_tasks: int = 1, bias: bool = True, dropout: float = 0.1, edge_dim: int = None, num_descriptors: int = 0)[source]

Bases: BaseModel

Fast KAN-augmented GATv2.

Parameters:
  • num_atom_features (int) – The number of atom features.

  • hidden_dim (int) – The number of hidden dimensions.

  • num_layers (int) – The number of layers.

  • num_heads (int) – The number of heads. Default to 8.

  • pred_hidden_dim (int) – The number of hidden dimensions for the prediction MLP. Default to 128.

  • grid_min (float) – The minimum value of the grid. Default to -2.0.

  • grid_max (float) – The maximum value of the grid. Default to 2.0.

  • num_grids (int) – The number of grids. Default to 8.

  • num_tasks (int) – The number of tasks. Default to 1.

  • bias (bool) – Whether to use bias. Default to True.

  • dropout (float) – The dropout rate. Default to 0.1.

  • edge_dim (int) – The number of edge dimensions.

  • num_descriptors (int) – The number of descriptors. Default to 0.

forward(batch: Polymer)[source]

Forward pass.

Parameters:

batch (Polymer) – The batch of data.

Returns:

The output tensor.

Return type:

torch.Tensor

gatv2_lineevo

class polymon.model.gatv2.lineevo.GATv2LineEvo(num_atom_features: int, hidden_dim: int, num_layers: int, num_heads: int = 8, pred_hidden_dim: int = 128, pred_dropout: float = 0.2, pred_layers: int = 2, activation: str = 'prelu', num_tasks: int = 1, bias: bool = True, dropout: float = 0.1, edge_dim: int = None, num_lineevo_layers: int = 2)[source]

Bases: BaseModel

GATv2 with LineEvo.

Parameters:
  • num_atom_features (int) – The number of atom features.

  • hidden_dim (int) – The number of hidden dimensions.

  • num_layers (int) – The number of layers.

  • num_heads (int) – The number of heads. Default to 8.

  • pred_hidden_dim (int) – The number of hidden dimensions for the prediction MLP. Default to 128.

  • pred_dropout (float) – The dropout rate for the prediction MLP. Default to 0.2.

  • pred_layers (int) – The number of layers for the prediction MLP. Default to 2.

  • activation (str) – The activation function. Default to 'prelu'.

  • num_tasks (int) – The number of tasks. Default to 1.

  • bias (bool) – Whether to use bias. Default to True.

  • dropout (float) – The dropout rate. Default to 0.1.

  • edge_dim (int) – The number of edge dimensions.

  • num_lineevo_layers (int) – The number of LineEvo layers. Default to 2.

forward(data: Data)[source]

Forward pass.

Parameters:

data (Data) – The batch of data.

Returns:

The output tensor.

Return type:

torch.Tensor

gatv2_sage

class polymon.model.gatv2.gatv2_sage.GATv2SAGE(num_atom_features: int, hidden_dim: int, num_layers: int, num_heads: int = 8, pred_hidden_dim: int = 128, pred_dropout: float = 0.2, pred_layers: int = 2, activation: str = 'prelu', num_tasks: int = 1, bias: bool = True, dropout: float = 0.1, edge_dim: int = None, sage_aggr: str = 'mean', sage_normalize: bool = False, sage_project: bool = False)[source]

Bases: BaseModel

GATv2 with SAGEConv.

Parameters:
  • num_atom_features (int) – The number of atom features.

  • hidden_dim (int) – The number of hidden dimensions.

  • num_layers (int) – The number of layers.

  • num_heads (int) – The number of heads. Default to 8.

  • pred_hidden_dim (int) – The number of hidden dimensions for the prediction MLP. Default to 128.

  • pred_dropout (float) – The dropout rate for the prediction MLP. Default to 0.2.

  • pred_layers (int) – The number of layers for the prediction MLP. Default to 2.

  • activation (str) – The activation function. Default to 'prelu'.

  • num_tasks (int) – The number of tasks. Default to 1.

  • bias (bool) – Whether to use bias. Default to True.

  • dropout (float) – The dropout rate. Default to 0.1.

  • edge_dim (int) – The number of edge dimensions.

  • sage_aggr (str) – The aggregation function. Default to 'mean'.

  • sage_normalize (bool) – Whether to normalize the output. Default to False.

  • sage_project (bool) – Whether to project the output. Default to False.

forward(data: Data)[source]

Forward pass.

Parameters:

data (Data) – The batch of data.

Returns:

The output tensor.

Return type:

torch.Tensor

gatv2_source

class polymon.model.gatv2.multi_fidelity.GATv2_Source(num_atom_features: int, hidden_dim: int, num_layers: int, num_heads: int = 8, pred_hidden_dim: int = 128, num_tasks: int = 1, bias: bool = True, dropout: float = 0.1, edge_dim: int = None, source_names: List[int] = [1], **kwargs)[source]

Bases: BaseModel

GATv2 with source-specific heads.

Parameters:
  • num_atom_features (int) – The number of atom features.

  • hidden_dim (int) – The number of hidden dimensions.

  • num_layers (int) – The number of layers.

  • num_heads (int) – The number of heads. Default to 8.

  • pred_hidden_dim (int) – The number of hidden dimensions for the prediction MLP. Default to 128.

  • num_tasks (int) – The number of tasks. Default to 1.

  • bias (bool) – Whether to use bias. Default to True.

  • dropout (float) – The dropout rate. Default to 0.1.

  • edge_dim (int) – The number of edge dimensions.

  • source_names (List[str]) – The names of the sources. Default to ['internal'].

forward(batch: Polymer)[source]

Forward pass.

Parameters:

batch (Polymer) – The batch of data. It should have source attribute.

Returns:

The output tensor.

Return type:

torch.Tensor

gatv2_pe

class polymon.model.gatv2.position_encoding.GATv2_PE(num_atom_features: int, hidden_dim: int, num_layers: int, num_heads: int = 8, pred_hidden_dim: int = 128, pred_dropout: float = 0.2, pred_layers: int = 2, activation: str = 'prelu', num_tasks: int = 1, bias: bool = True, dropout: float = 0.1, edge_dim: int = None, num_descriptors: int = 0, position_encoding_type: Literal['sin', 'rope', 'learned'] = 'sin')[source]

Bases: BaseModel

GATv2 with position encoding.

Parameters:
  • num_atom_features (int) – The number of atom features.

  • hidden_dim (int) – The number of hidden dimensions.

  • num_layers (int) – The number of layers.

  • num_heads (int) – The number of heads. Default to 8.

  • pred_hidden_dim (int) – The number of hidden dimensions for the prediction MLP. Default to 128.

  • pred_dropout (float) – The dropout rate for the prediction MLP. Default to 0.2.

  • pred_layers (int) – The number of layers for the prediction MLP. Default to 2.

  • activation (str) – The activation function. Default to 'prelu'.

  • num_tasks (int) – The number of tasks. Default to 1.

  • bias (bool) – Whether to use bias. Default to True.

  • dropout (float) – The dropout rate. Default to 0.1.

  • edge_dim (int) – The number of edge dimensions.

  • num_descriptors (int) – The number of descriptors. Default to 0.

  • position_encoding_type (Literal['sin', 'rope', 'learned']) – The type of position encoding. Default to 'sin'.

forward(batch: Polymer)[source]

Forward pass.

Parameters:

batch (Polymer) – The batch of data. It should have relative_position attribute.

Returns:

The output tensor.

Return type:

torch.Tensor

gatv2_embed_residual

class polymon.model.gatv2.embed_residual.GATv2EmbedResidual(num_atom_features: int, hidden_dim: int, num_layers: int, num_heads: int = 8, pred_hidden_dim: int = 128, pred_dropout: float = 0.2, pred_layers: int = 2, activation: str = 'prelu', num_tasks: int = 1, bias: bool = True, dropout: float = 0.1, edge_dim: int = None, num_descriptors: int = 0, pretrained_model: GATv2 = None)[source]

Bases: BaseModel

GATv2 with embedding from pretrained model as residual connection.

Parameters:
  • num_atom_features (int) – The number of atom features.

  • hidden_dim (int) – The number of hidden dimensions.

  • num_layers (int) – The number of layers.

  • num_heads (int) – The number of heads. Default to 8.

  • pred_hidden_dim (int) – The number of hidden dimensions for the prediction MLP. Default to 128.

  • pred_dropout (float) – The dropout rate for the prediction MLP. Default to 0.2.

  • pred_layers (int) – The number of layers for the prediction MLP. Default to 2.

  • activation (str) – The activation function. Default to 'prelu'.

  • num_tasks (int) – The number of tasks. Default to 1.

  • bias (bool) – Whether to use bias. Default to True.

  • dropout (float) – The dropout rate. Default to 0.1.

  • edge_dim (int) – The number of edge dimensions.

  • num_descriptors (int) – The number of descriptors. Default to 0.

  • pretrained_model (GATv2) – The pretrained model. Default to None.

forward(batch: Polymer)[source]

Forward pass.

Parameters:

batch (Polymer) – The batch of data.

Returns:

The output tensor.

Return type:

torch.Tensor

kan_gin

class polymon.model.kan.gin.KAN_GIN(num_atom_features: int, hidden_dim: int, num_layers: int, dropout: float = 0.2, n_mlp_layers: int = 2, pred_hidden_dim: int = 128, pred_dropout: float = 0.2, pred_layers: int = 2, grid_size: int = 10)[source]

Bases: BaseModel

KAN-augmented GIN.

Parameters:
  • num_atom_features (int) – The number of atom features.

  • hidden_dim (int) – The number of hidden dimensions.

  • num_layers (int) – The number of layers.

  • dropout (float) – The dropout rate. Default to 0.2.

  • n_mlp_layers (int) – The number of MLP layers. Default to 2.

  • pred_hidden_dim (int) – The number of hidden dimensions for the prediction MLP. Default to 128.

  • pred_dropout (float) – The dropout rate for the prediction MLP. Default to 0.2.

  • pred_layers (int) – The number of layers for the prediction MLP. Default to 2.

  • grid_size (int) – The number of grid points. Default to 10.

forward(batch: Polymer)[source]

Forward pass.

Parameters:

batch (Polymer) – The batch of data.

Returns:

The output tensor.

Return type:

torch.Tensor

fastkan_gin

class polymon.model.kan.gin.FastKAN_GIN(num_atom_features: int, hidden_dim: int, num_layers: int, dropout: float = 0.2, n_mlp_layers: int = 2, pred_hidden_dim: int = 128, grid_min: float = -4.0, grid_max: float = 3.0, num_grids: int = 10)[source]

Bases: BaseModel

Fast KAN-augmented GIN.

Parameters:
  • num_atom_features (int) – The number of atom features.

  • hidden_dim (int) – The number of hidden dimensions.

  • num_layers (int) – The number of layers.

  • dropout (float) – The dropout rate. Default to 0.2.

  • n_mlp_layers (int) – The number of MLP layers. Default to 2.

  • pred_hidden_dim (int) – The number of hidden dimensions for the prediction MLP. Default to 128.

  • grid_min (float) – The minimum value of the grid. Default to -4.0.

  • grid_max (float) – The maximum value of the grid. Default to 3.0.

  • num_grids (int) – The number of grids. Default to 10.

forward(batch: Polymer)[source]

Forward pass.

Parameters:

batch (Polymer) – The batch of data.

Returns:

The output tensor.

Return type:

torch.Tensor

kan_gcn

class polymon.model.kan.gcn.KAN_GCN(num_node_features: int, hidden_dim: int, num_layers: int, grid_size: int = 10, pred_hidden_dim: int = 128, pred_dropout: float = 0.0, pred_layers: int = 2)[source]

Bases: BaseModel

KAN-GCN wrapper.

Parameters:
  • num_node_features (int) – The number of node features.

  • hidden_dim (int) – The number of hidden dimensions.

  • num_layers (int) – The number of layers.

  • grid_size (int) – The number of grid points. Default to 10.

  • pred_hidden_dim (int) – The number of hidden dimensions for the prediction MLP. Default to 128.

  • pred_dropout (float) – The dropout rate for the prediction MLP. Default to 0.0.

  • pred_layers (int) – The number of layers for the prediction MLP. Default to 2.

forward(batch: Polymer)[source]

Forward pass.

Parameters:

batch (Polymer) – The batch of data.

Returns:

The output tensor.

Return type:

torch.Tensor

dmpnn

class polymon.model.dmpnn.DMPNN(mode: str = 'regression', n_classes: int = 3, n_tasks: int = 1, global_features_size: int = 0, atom_fdim: int = 133, bond_fdim: int = 14, hidden_dim: int = 300, num_layers: int = 3, bias: bool = False, enc_activation: str = 'relu', enc_dropout_p: float = 0.0, aggregation: str = 'mean', aggregation_norm: int | float = 100, ffn_hidden: int = 300, ffn_activation: str = 'relu', ffn_layers: int = 3, ffn_dropout_p: float = 0.0, ffn_dropout_at_input_no_act: bool = True)[source]

Bases: BaseModel

Directed Message Passing Neural Network. The implementation is adapted from DeepChem.

Parameters:
  • mode (str) – The mode of the model. Default to regression.

  • n_classes (int) – The number of classes. Default to 3.

  • n_tasks (int) – The number of tasks. Default to 1.

  • global_features_size (int) – The size of the global features. Default to 0.

  • atom_fdim (int) – The number of atom features. Default to 133.

  • bond_fdim (int) – The number of bond features. Default to 14.

  • hidden_dim (int) – The number of hidden dimensions. Default to 300.

  • num_layers (int) – The number of layers. Default to 3.

  • bias (bool) – Whether to use bias. Default to False.

  • enc_activation (str) – The activation function for the encoder. Default to relu.

  • enc_dropout_p (float) – The dropout rate for the encoder. Default to 0.0.

  • aggregation (str) – The aggregation function. Default to mean.

  • aggregation_norm (Union[int, float]) – The normalization factor for the aggregation. Default to 100.

  • ffn_hidden (int) – The number of hidden dimensions for the FFN. Default to 300.

  • ffn_activation (str) – The activation function for the FFN. Default to relu.

  • ffn_layers (int) – The number of layers for the FFN. Default to 3.

  • ffn_dropout_p (float) – The dropout rate for the FFN. Default to 0.0.

  • ffn_dropout_at_input_no_act (bool) – Whether to apply dropout at the input without activation. Default to True.

forward(pyg_batch: Batch) → Tensor | Sequence[Tensor][source]
Parameters:
  • data (Batch) –

    A pytorch-geometric batch containing tensors for:

    • atom_features

    • f_ini_atoms_bonds

    • atom_to_incoming_bonds

    • mapping

    • global_features

  • batch. (The molecules_unbatch_key is also derived from the)

  • batch) ((List containing number of atoms in various molecules of the)

Returns:

output – Predictions for the graphs

Return type:

Union[torch.Tensor, Sequence[torch.Tensor]]

kan_dmpnn

class polymon.model.kan.dmpnn.KAN_DMPNN(mode: str = 'regression', n_classes: int = 3, n_tasks: int = 1, global_features_size: int = 0, atom_fdim: int = 133, bond_fdim: int = 14, hidden_dim: int = 300, num_layers: int = 3, bias: bool = False, enc_activation: str = 'relu', enc_dropout_p: float = 0.0, aggregation: str = 'mean', aggregation_norm: int | float = 100, ffn_hidden: int = 300, ffn_activation: str = 'relu', ffn_layers: int = 3, ffn_dropout_p: float = 0.0, ffn_dropout_at_input_no_act: bool = True, grid_size: int = 3)[source]

Bases: BaseModel

KAN-augmented DMPNN.

Parameters:
  • mode (str) – The mode of the model. Default to regression.

  • n_classes (int) – The number of classes. Default to 3.

  • n_tasks (int) – The number of tasks. Default to 1.

  • global_features_size (int) – The size of the global features. Default to 0.

  • atom_fdim (int) – The number of atom features. Default to 133.

  • bond_fdim (int) – The number of bond features. Default to 14.

  • hidden_dim (int) – The number of hidden dimensions. Default to 300.

  • num_layers (int) – The number of layers. Default to 3.

  • bias (bool) – Whether to use bias. Default to False.

  • enc_activation (str) – The activation function for the encoder. Default to relu.

  • enc_dropout_p (float) – The dropout rate for the encoder. Default to 0.0.

  • aggregation (str) – The aggregation function. Default to mean.

  • aggregation_norm (Union[int, float]) – The normalization factor for the aggregation. Default to 100.

  • ffn_hidden (int) – The number of hidden dimensions for the FFN. Default to 300.

  • ffn_activation (str) – The activation function for the FFN. Default to relu.

  • ffn_layers (int) – The number of layers for the FFN. Default to 3.

  • ffn_dropout_p (float) – The dropout rate for the FFN. Default to 0.0.

  • ffn_dropout_at_input_no_act (bool) – Whether to apply dropout at the input without activation. Default to True.

forward(pyg_batch: Batch) → Tensor | Sequence[Tensor][source]

Forward pass.

Parameters:

pyg_batch (Batch) – The batch of data. It should be preprocessed by polymon.data.utils.DMPNNTransform.

Returns:

The output of the model.

Return type:

Union[torch.Tensor, Sequence[torch.Tensor]]