Browse by type

Active learning for machine-learned interatomic potentials.
Train, explore, select, label, evaluate, and deploy from one modular toolkit.
Why CURATOR · Installation · Quick start · CLI · Configuration · Citation
CURATOR is a config-driven framework for building robust machine-learned interatomic potentials (MLIPs). It brings equivariant neural networks, model fine-tuning and knowledge distillation, atomistic simulation, uncertainty-aware batch selection, first-principles labeling, evaluation, and deployment into a single workflow.
Reference data → Train → Simulate → Select → Label ↺
Evaluate throughout the loop · Deploy when the potential is ready
Every stage can run independently, while MyQueue can connect the stages into autonomous, iteration-aware HPC workflows.
| Where CURATOR is strongest | |
|---|---|
| Models | PaiNN · NequIP · MACE · Allegro · eSEN · MatGL. Train CURATOR's native architectures, attach CURATOR heads to external backbones, or bring pretrained models into the same workflow through dedicated adapters. |
| Training | Train, adapt, or compress. PyTorch Lightning powers distributed training and composite energy/force/virial/Hessian objectives; full, head-only, and LoRA fine-tuning sit alongside hessian-based knowledge distillation. |
| Exploration | Run ASE or TorchSim directly, or use pair_style curator and pair_style mliap unified in LAMMPS. |
| Uncertainty | Uncertainty-aware simulation Ensemble disagreement and Mahalanobis distance can be evaluated at run time, globally or per atom. |
| Selection | Model-aware data curation. Build feature space from gradient-based or learned latent features, then collect structures with active learning algorithms like LCMD or DIRECT/BIRCH, Max-distance, max-determinant, and CUR. |
| Labeling | VASP and GPAW adapters for DFT labeling. |
| Evaluation | Built-in energy/force metrics and diagnostic plots support. |
| Deployment | Produce TorchScript for pair_style curator or ML-IAP models for mliap; augment a trained model with uncertainty (ensemble or Mahalanobis); cuEquivariance or OpenEquivariance backends. |
All stages use the same YAML configuration system without requiring the full workflow to run as one monolith.
Install PyTorch for the CPU or CUDA environment you intend to use, following the official PyTorch instructions, then install CURATOR.
python -m pip install --upgrade pip
python -m pip install curator-torch
Official wheels include CURATOR's native C++ neighbor list, which is the default backend for data loading. Installing from source instead requires a C++17 compiler and Python development headers.
git clone https://github.com/Yangxinsix/curator.git
cd curator
python -m pip install -e .
For the current development branch, Python 3.10 or newer is recommended.
| Extra | Install command | Adds |
|---|---|---|
| Performance extras | python -m pip install "curator-torch[opt]" |
torch-scatter plus the alternative ASAP3 and matscipy neighbor backends |
| cuEquivariance | python -m pip install "curator-torch[cueq]" |
NVIDIA cuEquivariance acceleration |
| TorchSim | python -m pip install "curator-torch[torchsim]" |
TorchSim simulation backend |
[!NOTE] GPU packages are platform-specific. Confirm that the PyTorch, CUDA, and cuEquivariance builds are compatible before installing acceleration extras.
The repository includes a small dataset and checkpoint, so you can verify an installation without training first:
curator-evaluate \
--data example/LiFePO4.traj \
--model example/best_model.ckpt \
--device cpu \
--out runs/evaluate
Metrics and plots are written to runs/evaluate/LiFePO4/:
runs/evaluate/LiFePO4/
├── metrics.json
├── parity_energy.png
├── parity_forces_xyz.png
├── hist_energy_error.png
├── hist_force_error_norm.png
└── bar_force_mae_by_element.png
Add --save-data to also write results.npz, or --no-plot for metrics-only evaluation.
cd example/train
curator-train cfg=config.yaml
Before a production run, review data.datapath, device, batch size, precision, and trainer.max_epochs in the config. Training writes the resolved config, log, checkpoints, and deployable model into run_path.
Each stage consumes a YAML config and hands an artifact to the next stage:
curator-train cfg=train.yaml
curator-simulate cfg=simulate.yaml
curator-select cfg=select.yaml
curator-label cfg=label.yaml
dataset ──▶ train ──▶ model ──▶ simulate ──▶ pool ──▶ select ──▶ indices
▲ │
└──────────────────────────── label ◀──────────────────────────────┘
Ready-to-edit examples live in example/, while reusable defaults and component groups live in curator/configs/.
Installing CURATOR provides the following commands:
| Command | Purpose |
|---|---|
curator-train |
Train, resume, or fine-tune a potential; optionally deploy the best checkpoint. |
curator-simulate |
Run configured MD, optimization, NEB, TorchSim, or LAMMPS exploration. |
curator-select |
Compute features and select an informative batch from a structure pool. |
curator-label |
Label selected structures with a configured electronic-structure calculator. |
curator-evaluate |
Evaluate checkpoints or ensembles and export metrics, plots, and predictions. |
curator-deploy |
Export TorchScript or LAMMPS ML-IAP models, including uncertainty-aware ensembles. |
curator-convert |
Upgrade checkpoints or convert model backends, formats, and domain structure. |
curator-workflow |
Submit the iterative pipeline through MyQueue. |
Two useful post-training commands are:
# Export a TorchScript model
curator-deploy model.ckpt --target_path compiled_model.pt
# Export a LAMMPS ML-IAP model
curator-deploy model.ckpt --mliap \
--element-types Fe Li O P \
--target_path mliap_model.pt
Run any command with --help for its current options. Hydra-driven commands also accept direct overrides such as device=cpu, run_path=runs/train, or trainer.max_epochs=10.
CURATOR composes package defaults with a user YAML supplied through cfg=<path>. The resolved configuration is saved beside the run artifacts, so every experiment remains inspectable and reproducible.
| Stage | Main inputs | Default run artifacts |
|---|---|---|
| Train | Dataset, representation, task, trainer | training.log, config.yaml, model_path/, compiled_model.pt |
| Simulate | Model, initial structures, simulator | simulation.log, config.yaml, trajectories and warning structures |
| Select | Model, pool, optional training set | selection.log, config.yaml, selected.json, optional feature stores and selected.traj |
| Label | Pool, selected indices, annotator | labelling.log, config.yaml, dft_structures.db, appended dataset |
| Evaluate | Model or ensemble, labeled dataset | predict.log, per-dataset metrics.json, plots, optional results.npz |
The most important component groups are:
curator/configs/
├── model/representation/ # PaiNN, NequIP, MACE, Allegro, eSEN
├── data/ # single- and multi-domain datasets
├── finetune/ # full, head-only, LoRA
├── simulator/ # engines, callbacks, uncertainty
├── task/ # objectives, distillation, optimizers, schedulers
├── trainer/ # Lightning runtime, logging, callbacks
└── annotator/ # VASP and GPAW labeling
Found a bug or have an idea for a new model, simulator, selector, or annotator? Please open an issue.
If CURATOR contributes to your research, please cite the CURATOR preprint:
@article{yang2024curator,
title = {CURATOR: Building Robust Machine Learning Potentials for Atomistic
Simulations Autonomously with Batch Active Learning},
author = {Yang, Xin and Petersen, Martin Hoffmann and Sechi, Renata and others},
journal = {ChemRxiv},
year = {2024},
doi = {10.26434/chemrxiv-2024-p5t3l}
}
$ claude mcp add curator \
-- python -m otcore.mcp_server <graph>