We are the AI and computational discovery group at Aurigene Pharmaceutical Services Limited. This organization is our public home for open models, benchmark datasets and browser-based tools that support AI-driven drug discovery β from picking a target, through generating and triaging chemical matter, to planning the synthesis and reading the literature that justifies all of it.
Everything here is open, permissively licensed, and mirrored from the original authors with full attribution.
No GPU, no login, no install. These are static Spaces, so they never sleep and load instantly.
![]() |
![]() |
![]() |
Paste a SMILES string or a compound name. Get the 2D structure, physicochemical descriptors, Lipinski / Veber / Ghose / Egan / Muegge rules, an approximate QED, highlighted structural alerts, nearest approved drugs by Tanimoto similarity, and batch profiling with CSV export. RDKit runs as WebAssembly, so nothing is uploaded.
Enter a UniProt accession, a human gene symbol or a PDB ID. Get the target's annotation, its AlphaFold model coloured by per-residue confidence, every experimental structure with bound ligands, DrugBank drugs that hit it, and full sequence physicochemistry.
A live dashboard of this whole catalogue, mapped onto the five stages of discovery, with copy-paste transformers snippets and adoption statistics pulled from the Hub API.
Understand the protein before you try to drug it.
Generate and screen chemical matter.
Predict properties, refine the series.
Can we actually make it?
Mine the papers that justify the programme.
Standard evaluation sets, mirrored so every model above has data to train and benchmark against. Row counts on each card are computed from the files themselves.
Lead optimization β MoleculeNet ADMET & toxicity
| Dataset | Task | Rows |
|---|---|---|
| MoleculeNet_BBBP | Blood-brain barrier penetration | 2,039 |
| MoleculeNet_BACE | BACE-1 inhibition, an Alzheimer's target | 1,513 |
| MoleculeNet_ClinTox | Clinical toxicity & FDA approval | 1,477 |
| MoleculeNet_Tox21 | 12 toxicity assays | 7,831 |
| MoleculeNet_SIDER | Marketed-drug side effects | 1,427 |
| MoleculeNet_ESOL | Aqueous solubility | 1,128 |
| MoleculeNet_FreeSolv | Hydration free energy | 642 |
| MoleculeNet_Lipophilicity | logD at pH 7.4 | 4,200 |
Hit generation β virtual screening
| Dataset | Task | Rows |
|---|---|---|
| MoleculeNet_HIV | HIV replication inhibition | 41,127 |
Evidence & literature β medical QA
| Dataset | Task | Rows |
|---|---|---|
| MedQA-USMLE-4-options | US licensing exam, 4-option MCQ | 11,451 |
| MedMCQA | Medical entrance exam questions | 193,155 |
All eleven are in the Benchmark Datasets collection.
from transformers import AutoModel, AutoTokenizer
model_id = "Aurigene-AI/MoLFormer-XL-both-10pct"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModel.from_pretrained(model_id, trust_remote_code=True, deterministic_eval=True)
smiles = ["CC(=O)Oc1ccccc1C(=O)O", "CC(C)Cc1ccc(cc1)C(C)C(=O)O"]
embeddings = model(**tokenizer(smiles, padding=True, return_tensors="pt")).pooler_output
print(embeddings.shape) # torch.Size([2, 768])
Browse the catalogue as curated collections: molecular representation & property prediction, generative chemistry & synthesis planning, protein & target modeling, biomedical language models, interactive tools, benchmark datasets.
Every model here is a mirror of an open upstream release. The original authors β IBM Research, Meta AI, Microsoft Research, DeepChem, ChemFM, BioMistral and others β retain all credit, and each repository keeps the upstream model card and licence intact. Please cite the original work.
Models and tools are provided for research use. Rule-based filters and predictions are triage heuristics, not statements about safety or efficacy, and nothing here is a medical device or clinical advice.
π aurigeneservices.com