Free Chemical Databases: A Practical Guide for Students and Researchers
Short answer: The most useful free chemical databases are PubChem (the largest general compound database), ChEMBL (curated bioactivity data for drug discovery), ChEBI (a curated dictionary of small molecules), ZINC (billions of purchasable compounds for virtual screening) and the RCSB PDB (3D protein–ligand structures). For spectra, use nmrshiftdb2 (NMR), MassBank (MS/MS) and the NIST Chemistry WebBook; for natural products, COCONUT; for patents, SureChEMBL; and UniChem links the identifiers between them.
Chemistry students meet the same problem sooner or later: you have a molecule, and you need its properties, its biological activity, a crystal structure, a reference spectrum or a patent. Each of those questions lives in a different database. This guide explains what each major free chemical database is, what it is good for, how big it is, what the licence allows, and how to query it, with real example URLs that we tested before publishing.
All figures below were checked in October 2026 against the official sites or APIs. Database sizes change every release, so treat them as snapshots, not constants.
What is a chemical database?
A chemical database is a searchable collection of chemical structures linked to data about them: names, identifiers, calculated properties, experimental measurements, biological activities, 3D coordinates, spectra or the documents where they were reported. Most modern databases can be searched by name, by identifier (CAS number, InChIKey, SMILES) or by drawing a structure, and many offer a free REST API for programmatic access.
The key thing to understand is that no single database does everything. They differ in focus:
- Compound encyclopedias (PubChem, ChEBI) answer "what is this molecule?"
- Bioactivity databases (ChEMBL, DrugBank) answer "what does this molecule do to a protein?"
- Screening libraries (ZINC) answer "what similar molecules can I buy or dock?"
- Structural databases (RCSB PDB) answer "how does this molecule sit in its target?"
- Spectral databases (nmrshiftdb2, MassBank, NIST WebBook) answer "what should its spectrum look like?"
- Specialist collections (COCONUT for natural products, SureChEMBL for patents) cover niches the big ones miss.
How do the free chemical databases compare?
| Database | Focus | Size (checked Oct 2026) | Best for | API | Licence |
|---|---|---|---|---|---|
| PubChem | General compounds, substances, bioassays | ~124 M compounds, ~349 M substances | Identifiers, properties, broad lookups | PUG-REST, PUG-View | Free; reuse terms vary by depositor |
| ChEMBL | Curated bioactivity | 2.92 M compounds, 24.5 M activities, 18,552 targets (ChEMBL 37) | SAR, potency, drug targets | REST (JSON/XML) | CC BY-SA 3.0 |
| ChEBI | Curated small-molecule dictionary + ontology | 201,471 entries, 61,961 fully curated (release 244) | Definitions, classes, biological roles | REST (ChEBI 2.0) | CC BY 4.0 |
| ZINC-22 | Purchasable "make-on-demand" compounds | ~54.9 B (2D), ~5.9 B (3D) | Virtual screening, analogue buying | Web (CartBlanche), bulk download | Free to use; no redistribution of major portions without permission |
| RCSB PDB | 3D macromolecular structures | 260,724 entries | Protein–ligand complexes, docking prep | Data API, Search API | CC0 1.0 |
| UniChem | Identifier cross-mapping | 25 sources | Converting IDs between databases | REST | Free (EMBL-EBI terms); source licences apply |
| nmrshiftdb2 | NMR spectra + prediction | 271,817 structures, 70,030 measured spectra | ¹H/¹³C NMR reference and prediction | Web; new REST interface | Open content licence |
| MassBank (Europe) | Experimental mass spectra | 139,006 records | MS/MS matching, metabolomics | REST | Per-record Creative Commons |
| COCONUT 2.0 | Natural products | 738,829 unique molecules, 71 collections | Natural product discovery | REST (login required) | CC0 (data) |
| SureChEMBL | Chemistry from patents | 116.6 M patent documents (May 2025) | Prior art, patent landscaping | REST | CC BY 4.0 |
| DrugBank | Drugs and drug targets | n/a (not checked) | Approved-drug details | Paid API | Free academic, CC BY-NC 4.0; commercial licence needed |
| NIST Chemistry WebBook | Thermochemistry, IR, EI-MS | n/a (not checked) | Reference IR/MS, physical data | Web only | NIST Standard Reference Data (copyrighted) |
What is PubChem? The largest free chemical database
PubChem is the world's largest free chemistry database, run by the US National Center for Biotechnology Information (NCBI). It aggregates data from more than a thousand contributors into linked collections: Substance (what depositors submitted), Compound (unique standardized structures), BioAssay, Protein, Gene, Pathway and Patent.
- Size: about 124 million compounds and 349 million substances on the PubChem homepage (checked October 2026). The peer-reviewed PubChem 2025 update reported 118.6 million compounds as of September 2024.
- Used for: looking up names, synonyms, CAS numbers, formulas, computed properties (XLogP, TPSA), safety data, suppliers and literature.
- Who uses it: students, teachers, analytical chemists, toxicologists, cheminformaticians and anyone who needs a quick, reliable identifier.
- Access: web search with a structure drawer, and the free PUG-REST API. PubChem asks programmatic users to stay below 5 requests per second.
- Licence: free to use; reuse rights follow each data source, so check the source before redistributing data.
Plain-language query: "Give me the formula, molecular weight, SMILES and InChIKey of aspirin."
https://pubchem.ncbi.nlm.nih.gov/rest/pug/compound/name/aspirin/property/MolecularFormula,MolecularWeight,SMILES,InChIKey/JSON
Real response (trimmed):
{"PropertyTable": {"Properties": [{
"CID": 2244,
"MolecularFormula": "C9H8O4",
"MolecularWeight": "180.16",
"SMILES": "CC(=O)OC1=CC=CC=C1C(=O)O",
"InChIKey": "BSYNRYMUTXBXSQ-UHFFFAOYSA-N"}]}}
Heads-up: PubChem has renamed its SMILES properties. Asking for the old CanonicalSMILES still works, but the answer now comes back labelled ConnectivitySMILES (no stereochemistry). Ask for SMILES if you want the full, stereo-aware string.
More tested PubChem queries:
- InChIKey to CID:
https://pubchem.ncbi.nlm.nih.gov/rest/pug/compound/inchikey/BSYNRYMUTXBXSQ-UHFFFAOYSA-N/cids/JSONreturns2244. - 2D similarity (≥90%) to cetirizine:
https://pubchem.ncbi.nlm.nih.gov/rest/pug/compound/fastsimilarity_2d/cid/2678/cids/JSON?Threshold=90&MaxRecords=10 - Substructure search from SMILES:
https://pubchem.ncbi.nlm.nih.gov/rest/pug/compound/fastsubstructure/smiles/c1ccc(cc1)C(c1ccc(Cl)cc1)N1CCNCC1/cids/JSON?MaxRecords=10
What is ChEMBL? The go-to bioactivity database
ChEMBL is a manually curated database of bioactive, drug-like molecules, maintained by EMBL's European Bioinformatics Institute (EMBL-EBI). Its core is measured activity data, such as IC50, Ki and EC50 values, extracted from journal articles and deposited datasets, each linked to a compound, an assay and a biological target.
- Size (ChEMBL 37, May 2026): 2,921,148 distinct compounds, 24,527,044 activities, 18,552 targets and 101,100 publications, as reported by the ChEMBL API status endpoint.
- Used for: structure–activity relationships (SAR), finding known actives for a target, building machine-learning training sets, and checking drug approval phase.
- Who uses it: medicinal chemists, pharmacologists, computational chemists and data scientists.
- Access: web interface and a free REST API (JSON or XML), plus full downloads (SQLite, PostgreSQL, SDF).
- Licence: CC BY-SA 3.0. ChEMBL asks that ChEMBL IDs and the release number be preserved when you reuse its data.
Plain-language query: "Find aspirin in ChEMBL."
https://www.ebi.ac.uk/chembl/api/data/molecule/search.json?q=aspirin
The first hit is CHEMBL25 (ASPIRIN, max phase 4, i.e. approved), with SMILES CC(=O)Oc1ccccc1C(=O)O.
Plain-language query: "What potency values does ChEMBL hold for cetirizine (CHEMBL1000) on the human histamine H1 receptor (CHEMBL231)?"
https://www.ebi.ac.uk/chembl/api/data/activity.json?molecule_chembl_id=CHEMBL1000&target_chembl_id=CHEMBL231&pchembl_value__isnull=false
Real response (trimmed to two of the 12 records returned):
{"activities": [
{"standard_type": "Ki", "standard_value": "14.0", "standard_units": "nM",
"pchembl_value": "7.85", "target_pref_name": "Histamine H1 receptor"},
{"standard_type": "Ki", "standard_value": "5.89", "standard_units": "nM",
"pchembl_value": "8.23", "target_pref_name": "Histamine H1 receptor"}]}
PubChem vs ChEMBL: what is the difference?
PubChem is broad; ChEMBL is deep. PubChem tries to include every chemical anyone has deposited, so it is the best place to identify a compound. ChEMBL contains far fewer compounds but curates their biological activity in a consistent format (standard units, pChEMBL values, mapped targets), so it is the better place to ask "how potent is it, and on what?" In practice you use both: PubChem to confirm the structure and identifiers, ChEMBL to read the pharmacology.
What is ChEBI? A curated dictionary of small molecules
ChEBI (Chemical Entities of Biological Interest) is a manually curated dictionary and ontology of small molecules, run by EMBL-EBI. Curated entries have a definition, structure, synonyms and ontology links that say what the molecule is (e.g. "a member of the class of benzoic acids") and what role it has (e.g. "non-steroidal anti-inflammatory drug").
- Size: 201,471 entries, of which 61,961 are fully curated "3-star" entries (release 244, reported in the 2026 Nucleic Acids Research paper).
- Used for: consistent naming and classification, annotating metabolomics and systems-biology data, and teaching functional groups and compound classes.
- Access: ChEBI 2.0 (launched October 2025) has a new website and REST API; the old SOAP services were retired.
- Licence: CC BY 4.0.
Plain-language query: "What is CHEBI:15365?"
https://www.ebi.ac.uk/chebi/backend/api/public/compound/CHEBI:15365/
The response returns "name": "acetylsalicylic acid", formula C9H8O4, monoisotopic mass 180.04226 and the definition text. A text search also works: https://www.ebi.ac.uk/chebi/backend/api/public/es_search/?term=aspirin.
What is ZINC? Billions of compounds for virtual screening
ZINC is a free database of commercially available ("tangible") compounds prepared for virtual screening, built by the Irwin and Shoichet labs at UCSF. ZINC-22 focuses on huge make-on-demand catalogues from Enamine, WuXi and Mcule, while its sister database ZINC20 covers smaller in-stock catalogues. ZINC15 is the older generation, which many tutorials still mention.
- Size: about 54.9 billion molecules in 2D and 5.9 billion in ready-to-dock 3D formats (CartBlanche22 homepage, checked October 2026). The ZINC-22 paper reported 37 billion 2D molecules at publication, which shows how fast it grows.
- Used for: docking campaigns, finding purchasable analogues, and downloading ready-to-dock 3D files with charges and conformations.
- Who uses it: computational chemists, structural biologists and academic screening groups.
- Access: the CartBlanche web interface (similarity, substructure and ID search), bulk downloads, and cloud copies.
- Licence: "free to use and download for everyone"; you may not redistribute major portions without written permission (UCSF ZINC licence).
Plain-language query: "Find purchasable compounds similar to cetirizine." Paste the SMILES OC(=O)COCCN1CCN(CC1)C(c1ccccc1)c1ccc(Cl)cc1 into the similarity search at cartblanche22.docking.org.
Honest note: we could not script ZINC queries reliably while writing this guide. The ZINC15/ZINC20 pages sit behind a captcha, and the CartBlanche22 search tasks we submitted programmatically came back empty. The web interface is the recommended route for most users.
What is the RCSB PDB? 3D structures of proteins and ligands
The Protein Data Bank (PDB) is the single global archive of experimentally determined 3D structures of proteins, nucleic acids and their complexes; RCSB PDB is its US data centre and search portal. For chemists, its value is the bound ligands: you can see exactly how a small molecule sits in its binding pocket.
- Size: 260,724 current entries (RCSB holdings API, checked October 2026).
- Used for: studying binding modes, preparing receptors for docking, teaching protein structure, and finding every structure that contains a given ligand.
- Access: web, the Data API (
data.rcsb.org) and the Search API (search.rcsb.org), which supports text, sequence and chemical searches. - Licence: PDB data are released under CC0 1.0 (public domain).
Plain-language query: "Show me the ligand record for aspirin (component ID AIN)."
https://data.rcsb.org/rest/v1/core/chemcomp/AIN
This returns the name 2-(ACETYLOXY)BENZOIC ACID and formula C9 H8 O4.
Plain-language query: "Which PDB entries contain aspirin?" (Search API; paste it into a browser, which URL-encodes the JSON for you):
https://search.rcsb.org/rcsbsearch/v2/query?json={"query":{"type":"terminal","service":"text","parameters":{"attribute":"rcsb_nonpolymer_instance_annotation.comp_id","operator":"exact_match","value":"AIN"}},"return_type":"entry"}
Real response (trimmed): "total_count": 8, with entries such as 1OXR, 1TGM and 3GCL. A chemical search posted with the SMILES of aspirin (service: "chemical") also returned AIN as the top match. You can load any of these entries by its PDB ID in the MolDraw protein viewer or open it at rcsb.org.
What is UniChem? Translating IDs between databases
UniChem is EMBL-EBI's identifier cross-referencing service: give it one structure identifier and it returns the matching IDs in other databases. It links records through the standard InChI/InChIKey, so it works whenever two databases hold the same structure.
- Size: 25 sources listed by the UniChem API, including ChEMBL, PubChem, ChEBI, DrugBank, RCSB PDB, PDBe, SureChEMBL, HMDB, BindingDB and nmrshiftdb2 (checked October 2026).
- Used for: merging datasets, jumping from a ChEMBL ID to a PDB ligand code, and checking where else a compound appears.
- Access: web and REST API (
/unichem/api/v1/). - Licence: free to use under EMBL-EBI's terms of use; the data behind each mapped ID keeps its source's licence.
Plain-language query: "Where else does aspirin's InChIKey appear?"
curl -X POST https://www.ebi.ac.uk/unichem/api/v1/compounds \
-H "Content-Type: application/json" \
-d '{"type":"inchikey","compound":"BSYNRYMUTXBXSQ-UHFFFAOYSA-N"}'
The response maps it to CHEMBL25, CHEBI:15365, PubChem CID 2244, DrugBank DB00945, PDB ligand AIN and Guide to Pharmacology 4139, among others.
What is nmrshiftdb2? A free NMR spectra database
nmrshiftdb2 is an open, peer-reviewed web database of organic structures and their NMR spectra, with built-in ¹H and ¹³C shift prediction. Its core is fully assigned spectra, many with raw data.
- Size: 271,817 structures, 70,030 measured spectra and 396,583 calculated spectra (homepage counter, checked October 2026).
- Used for: checking an assignment, finding reference shifts for a compound class, and teaching NMR interpretation.
- Access: web search by structure, substructure, spectrum and other properties; the site announced a new REST interface in June 2026 (we did not test it).
- Licence: the software is open source and the data are published under an open content licence.
Plain-language query: "Find measured ¹³C spectra of compounds containing a 4-chlorobenzhydryl group." Draw the fragment in the nmrshiftdb2 search tab and run a substructure search. For a quick estimate before you search, MolDraw's structure to NMR tool predicts a spectrum from your drawing.
What is MassBank? Experimental MS/MS spectra
MassBank is an open repository of experimental mass spectra of small molecules, contributed by laboratories worldwide. The examples here use MassBank Europe (massbank.eu). Each record describes one spectrum: compound, instrument, ionisation, collision energy and the peak list.
- Size: 139,006 records (MassBank API count, checked October 2026).
- Used for: identifying unknowns in LC-MS/MS and metabolomics, comparing fragmentation patterns, and teaching mass spectrometry.
- Access: web search (name, formula, peaks) and a REST API.
- Licence: each record carries its own Creative Commons licence (CC BY by default, but some are CC0, CC BY-SA or non-commercial), so check the
licensefield before reuse.
Plain-language query: "List MassBank spectra for caffeine."
https://massbank.eu/MassBank-api/records/search?compound_name=caffeine
This returned 134 record accessions. Fetching one, https://massbank.eu/MassBank-api/records/MSBNK-ACES_SU-AS000088, gives an LC-APCI Orbitrap MS2 spectrum of [M+H]⁺ with the precursor at m/z 195.088 and a major fragment at m/z 138.067, licensed CC BY.
What is COCONUT? The open natural products database
COCONUT (COlleCtion of Open NatUral producTs) is the largest open collection of natural product structures, curated by the Steinbeck group at Friedrich Schiller University Jena. Version 2.0 links each molecule to its source organisms, literature and stereochemical variants.
- Size: 738,829 unique molecules from 71 collections, with 73,137 organisms (COCONUT statistics page, checked October 2026). The COCONUT 2.0 paper reported 695,133 at its September 2024 release.
- Used for: natural-product dereplication, scaffold ideas, and building screening sets.
- Access: web search (name, structure, substructure, similarity) and full SDF/SQL downloads. The REST API exists but requires a free account login, so we did not run it for this guide.
- Licence: data CC0; code MIT.
Plain-language query: "Show natural products similar to caffeine." Paste caffeine's SMILES CN1C=NC2=C1C(=O)N(C(=O)N2C)C into the structure search at coconut.naturalproducts.net.
What is SureChEMBL? Chemistry from patents
SureChEMBL is EMBL-EBI's free patent chemistry database: it text- and image-mines patent documents and extracts the chemical structures they mention. It is the easiest free way to ask "which patents contain this structure?"
- Size: SureChEMBL 2.0 (May 2025) covers about 116.6 million patent documents from five authorities (CNIPA, EPO, JPO, USPTO, WIPO); USPTO alone contributed about 19.9 million unique compounds.
- Used for: prior-art checks, patent landscaping, and spotting compound series before they appear in journals.
- Access: web search by structure or keyword, a REST API, and bulk Parquet files updated every two weeks.
- Licence: CC BY 4.0.
Plain-language query: "Resolve the name cetirizine to a SureChEMBL compound."
https://www.surechembl.org/api/chemical/name/cetirizine
Real response (trimmed): "chemical_id": "4176", "inchi_key": "ZKLPARSLTMPFCP-UHFFFAOYSA-N". You can then search for patents containing that structure in the web interface.
What about DrugBank and the NIST Chemistry WebBook?
DrugBank is a detailed knowledge base of approved and investigational drugs, their targets, metabolism and interactions. It is free to browse, but downloads require an account; academic datasets are licensed CC BY-NC 4.0 for non-commercial use, and commercial use needs a paid licence. Treat it as "free to read, not free to reuse."
The NIST Chemistry WebBook (NIST Standard Reference Database 69) offers thermochemical data, gas-phase IR spectra and electron-ionisation mass spectra for many small molecules. It is free to search online, for example https://webbook.nist.gov/cgi/cbook.cgi?Name=caffeine&Units=SI, but the data are copyrighted and bulk redistribution needs permission from NIST.
Which chemical database should I use?
| Your question | Start with | Then try |
|---|---|---|
| What is this compound? Name, formula, CAS, properties | PubChem | ChEBI |
| How potent is it, and on which target? | ChEMBL | PubChem BioAssay |
| What class is it, and what is its biological role? | ChEBI | PubChem |
| Which similar compounds can I buy or dock? | ZINC-22 | PubChem (vendors) |
| How does a ligand bind its protein? | RCSB PDB | ChEMBL (target data) |
| I have an ID from database A; what is it in database B? | UniChem | PubChem cross-references |
| What should its ¹H/¹³C NMR look like? | nmrshiftdb2 | MolDraw NMR prediction |
| Does my MS/MS spectrum match a known compound? | MassBank | NIST WebBook (EI-MS) |
| Is it a natural product, and from which organism? | COCONUT | ChEBI |
| Is this structure in a patent? | SureChEMBL | PubChem Patent |
| Details on an approved drug? | DrugBank (browse) | ChEMBL |
| Thermochemistry or reference IR? | NIST WebBook | PubChem |
How do I use these databases with MolDraw?
Most databases search best with a machine-readable identifier rather than a name. MolDraw gives you those identifiers straight from a drawing:
- Draw the structure in the MolDraw editor (or paste a name with name to structure).
- Copy SMILES from the top options row. Copy as… in the context menu also gives InChI and MOL formats (see the SMILES chapter of the MolDraw course).
- Make an InChIKey with the SMILES to InChIKey converter. InChIKeys are the most reliable exact-match key across PubChem, ChEMBL, ChEBI and UniChem.
- Query: paste the SMILES into a database's structure search, or drop the InChIKey into one of the REST URLs above.
- Bring results back: convert a PubChem CID with the CID to SMILES converter or an InChIKey with InChIKey to SMILES, then paste it into MolDraw to edit.
Worked example: how do I compare cetirizine analogues?
Suppose you are studying the antihistamine cetirizine and want to know which close analogues exist, how potent they are, and whether you could buy new ones. Here is a three-database workflow; every number comes from the live queries we ran.
Step 1: Identify it in PubChem. Draw cetirizine in MolDraw, copy the SMILES, and look it up:
https://pubchem.ncbi.nlm.nih.gov/rest/pug/compound/name/cetirizine/property/MolecularFormula,MolecularWeight,InChIKey,XLogP,TPSA/JSON
PubChem returns CID 2678, C21H25ClN2O3, MW 388.9, XLogP 1.7, TPSA 53 and InChIKey ZKLPARSLTMPFCP-UHFFFAOYSA-N.
Step 2: Find analogues in PubChem. The 2D similarity query (threshold 90) returns relatives including hydroxyzine (CID 3658, cetirizine's parent drug) and levocetirizine (CID 1549000, its single enantiomer), plus their salts. A substructure search on the chlorobenzhydryl-piperazine core adds chlorcyclizine, meclizine and buclizine.
Step 3: Compare activity in ChEMBL. Look up the InChIKey in ChEMBL:
https://www.ebi.ac.uk/chembl/api/data/molecule.json?molecule_structures__standard_inchi_key=ZKLPARSLTMPFCP-UHFFFAOYSA-N
This gives CHEMBL1000 (max phase 4, first approved 1995). Its activity records include human H1 receptor Ki values between 5.89 and 79 nM (pChEMBL 7.1–8.2) and a much weaker hERG (KCNH2) IC50 of about 30 µM (pChEMBL about 4.5). That window of roughly three orders of magnitude is the kind of selectivity comparison ChEMBL makes easy. A similarity search, https://www.ebi.ac.uk/chembl/api/data/similarity/ZKLPARSLTMPFCP-UHFFFAOYSA-N/80.json, lists levocetirizine and cetirizine salt forms to compare in the same way.
Step 4: Look for purchasable analogues in ZINC. Paste the SMILES into CartBlanche22's similarity search to see make-on-demand compounds that vary the acid side chain or the aryl rings. Then check any interesting hit's properties in MolDraw, for example with the Lipinski rule calculator.
Interpretation tip: activity values from different papers are measured in different assays. Compare pChEMBL values for the same target and assay type, and treat single measurements with caution.
How do I search chemical databases well?
- Prefer InChIKeys for exact matches and SMILES for similarity or substructure searches. Names are ambiguous ("vitamin C", salts, brand names).
- Watch for salts and stereoisomers. Cetirizine, cetirizine dihydrochloride and levocetirizine are different records in every database.
- Note the release or date. Cite ChEMBL with its release number and record the date you queried PubChem.
- Respect rate limits. PubChem asks for no more than 5 requests per second; batch large jobs or download the bulk files.
- Check the licence before you redistribute. Browsing is free everywhere in this guide; republishing data is not always.
Key takeaways
- PubChem is the largest free chemical database and the best first stop to identify a compound.
- ChEMBL is the best free bioactivity database: curated potency data linked to targets.
- ChEBI classifies molecules; UniChem translates IDs between databases.
- ZINC-22 is for virtual screening and buying analogues; RCSB PDB is for 3D binding modes.
- nmrshiftdb2, MassBank and the NIST WebBook cover NMR, MS/MS and IR/thermochemistry.
- COCONUT covers natural products and SureChEMBL covers patents.
- Draw in MolDraw, copy the SMILES, generate an InChIKey, and you can query all of them.
FAQ
What is the largest free chemical database?
PubChem is the largest free chemical database, with about 124 million unique compounds and 349 million deposited substances (checked October 2026). ZINC-22 lists more molecules (about 54.9 billion), but most of them are virtual make-on-demand compounds for screening rather than characterised substances.
What is the difference between PubChem and ChEMBL?
PubChem is a broad aggregator that tries to cover every deposited chemical, so it is best for identification and properties. ChEMBL is smaller but manually curated for bioactivity, with standardised potency values linked to targets, so it is best for structure–activity and drug-discovery questions.
Is ZINC free for commercial use?
According to the UCSF ZINC licence, ZINC is free to use and download for everyone, including companies. The restriction is redistribution: you may not redistribute major portions of ZINC without written permission from the ZINC team. The compounds themselves must be bought from the listed vendors.
Which free database has NMR spectra?
nmrshiftdb2 is the main free NMR database, with 70,030 measured spectra for 271,817 structures (checked October 2026) and built-in ¹H and ¹³C prediction. PubChem compound pages also link to spectra from other sources, and the NIST Chemistry WebBook covers IR and mass spectra rather than NMR.
Which free database has mass spectra?
MassBank holds 139,006 experimental mass spectrum records (checked October 2026), mostly MS/MS spectra of small molecules, under per-record Creative Commons licences. The NIST Chemistry WebBook offers electron-ionisation mass spectra for many small molecules online.
Can I search chemical databases by structure?
Yes. PubChem, ChEMBL, ChEBI, ZINC, RCSB PDB, nmrshiftdb2, COCONUT and SureChEMBL all support structure search, usually by drawing or pasting SMILES. You can draw the molecule in MolDraw, copy its SMILES, and paste it into the database's structure search, or use exact, substructure or similarity queries through their APIs.
Which chemical database is best for students?
Start with PubChem for names, properties and safety data, ChEBI for definitions and compound classes, and the RCSB PDB for 3D structures. All three are free, need no account and have clear web interfaces. Add ChEMBL when you study pharmacology, and nmrshiftdb2 when you learn NMR.
What is a bioactivity database?
A bioactivity database stores measured effects of compounds on biological targets, such as IC50, Ki or EC50 values, linked to the assay and protein. ChEMBL is the leading free bioactivity database; PubChem BioAssay holds screening results deposited by laboratories and screening centres.
Is DrugBank free?
DrugBank is free to browse online, and academic users can download datasets under a CC BY-NC 4.0 licence after registering. Commercial use, including use in commercial products or services, requires a paid licence from DrugBank.
How do I convert a PubChem CID into a ChEMBL ID?
Use UniChem. Look up the compound's InChIKey (PubChem shows it on every compound page), then send it to the UniChem compounds API, which returns the matching ChEMBL, ChEBI, DrugBank, PDB and other identifiers. For aspirin, InChIKey BSYNRYMUTXBXSQ-UHFFFAOYSA-N maps CID 2244 to CHEMBL25.
Sources
- PubChem: pubchem.ncbi.nlm.nih.gov · PUG-REST documentation · Kim S. et al. PubChem 2025 update, Nucleic Acids Res. 2025, doi:10.1093/nar/gkae1059
- ChEMBL: ebi.ac.uk/chembl · API documentation · Licence and about · ChEMBL 37 release post
- ChEBI: ebi.ac.uk/chebi · ChEBI 2.0 launch · ChEBI: re-engineered for a sustainable future, Nucleic Acids Res. 2026, academic.oup.com/nar/article/54/D1/D1768/8349173
- ZINC: CartBlanche22 · Tingle B. I. et al. ZINC-22, J. Chem. Inf. Model. 2023, doi:10.1021/acs.jcim.2c01253 · UCSF ZINC licence
- RCSB PDB: rcsb.org · Data API · Search API · Usage policy
- UniChem: ebi.ac.uk/unichem · API docs
- nmrshiftdb2: nmrshiftdb.nmr.uni-koeln.de
- MassBank: massbank.eu · MassBank API · Record format and licences
- COCONUT: coconut.naturalproducts.net · Chandrasekhar V. et al. COCONUT 2.0, Nucleic Acids Res. 2025, doi:10.1093/nar/gkae1063
- SureChEMBL: surechembl.org · SureChEMBL 2.0 announcement · FAQ and licence
- DrugBank: go.drugbank.com · Academic licensing
- NIST Chemistry WebBook: webbook.nist.gov/chemistry · NIST SRD licensing