Free Chemical Databases: A Practical Guide for Students and Researchers

Published: 6 Oct 2026 · By Rafeeque Mavoor · Category: Guides · Data checked: October 2026

Short answer: The most useful free chemical databases are PubChem (the largest general compound database), ChEMBL (curated bioactivity data for drug discovery), ChEBI (a curated dictionary of small molecules), ZINC (billions of purchasable compounds for virtual screening) and the RCSB PDB (3D protein–ligand structures). For spectra, use nmrshiftdb2 (NMR), MassBank (MS/MS) and the NIST Chemistry WebBook; for natural products, COCONUT; for patents, SureChEMBL; and UniChem links the identifiers between them.

Chemistry students meet the same problem sooner or later: you have a molecule, and you need its properties, its biological activity, a crystal structure, a reference spectrum or a patent. Each of those questions lives in a different database. This guide explains what each major free chemical database is, what it is good for, how big it is, what the licence allows, and how to query it, with real example URLs that we tested before publishing.

All figures below were checked in October 2026 against the official sites or APIs. Database sizes change every release, so treat them as snapshots, not constants.

Map of free chemical databases grouped by the type of data they hold Seven groups arranged around UniChem, which cross-links identifiers between them. Compounds and identifiers: PubChem and ChEBI. Bioactivity: ChEMBL, plus DrugBank which is licence-restricted. Screening and purchasable compounds: ZINC. Natural products: COCONUT. 3D structures: RCSB PDB. Spectra: nmrshiftdb2 for NMR, MassBank for MS/MS, NIST Chemistry WebBook for IR, mass spectra and thermochemistry. Patents: SureChEMBL. The free chemical database landscape Grouped by the main kind of data each resource holds · MolDraw guide · checked Oct 2026 Compounds & identifiers Names, structures, properties PubChemNIH · free ChEBICC BY 4.0 Bioactivity Potency, targets, drug data ChEMBLCC BY-SA 3.0 DrugBank*licence-restricted Virtual screening Purchasable, dockable molecules ZINC-22 / ZINC20 / ZINC15UCSF · billions of tangible compounds Natural products Structures + organisms COCONUT 2.0Univ. Jena · CC0 data UniChem links IDs across sources by InChIKey 3D structures Proteins, ligands, complexes RCSB PDBwwPDB archive · CC0 Spectra & physical data Experimental reference data for identification nmrshiftdb2NMR spectra · open MassBankMS/MS · per-record CC NIST WebBookIR, EI-MS, thermo Patents Chemistry mined from patent text SureChEMBLEMBL-EBI · patent chemistry * DrugBank is free to browse, but downloads need an account and commercial reuse needs a paid licence.
Free chemical databases grouped by data type: compounds, bioactivity, screening, natural products, 3D structures, spectra and patents, linked by UniChem.

What is a chemical database?

A chemical database is a searchable collection of chemical structures linked to data about them: names, identifiers, calculated properties, experimental measurements, biological activities, 3D coordinates, spectra or the documents where they were reported. Most modern databases can be searched by name, by identifier (CAS number, InChIKey, SMILES) or by drawing a structure, and many offer a free REST API for programmatic access.

The key thing to understand is that no single database does everything. They differ in focus:

How do the free chemical databases compare?

Database Focus Size (checked Oct 2026) Best for API Licence
PubChem General compounds, substances, bioassays ~124 M compounds, ~349 M substances Identifiers, properties, broad lookups PUG-REST, PUG-View Free; reuse terms vary by depositor
ChEMBL Curated bioactivity 2.92 M compounds, 24.5 M activities, 18,552 targets (ChEMBL 37) SAR, potency, drug targets REST (JSON/XML) CC BY-SA 3.0
ChEBI Curated small-molecule dictionary + ontology 201,471 entries, 61,961 fully curated (release 244) Definitions, classes, biological roles REST (ChEBI 2.0) CC BY 4.0
ZINC-22 Purchasable "make-on-demand" compounds ~54.9 B (2D), ~5.9 B (3D) Virtual screening, analogue buying Web (CartBlanche), bulk download Free to use; no redistribution of major portions without permission
RCSB PDB 3D macromolecular structures 260,724 entries Protein–ligand complexes, docking prep Data API, Search API CC0 1.0
UniChem Identifier cross-mapping 25 sources Converting IDs between databases REST Free (EMBL-EBI terms); source licences apply
nmrshiftdb2 NMR spectra + prediction 271,817 structures, 70,030 measured spectra ¹H/¹³C NMR reference and prediction Web; new REST interface Open content licence
MassBank (Europe) Experimental mass spectra 139,006 records MS/MS matching, metabolomics REST Per-record Creative Commons
COCONUT 2.0 Natural products 738,829 unique molecules, 71 collections Natural product discovery REST (login required) CC0 (data)
SureChEMBL Chemistry from patents 116.6 M patent documents (May 2025) Prior art, patent landscaping REST CC BY 4.0
DrugBank Drugs and drug targets n/a (not checked) Approved-drug details Paid API Free academic, CC BY-NC 4.0; commercial licence needed
NIST Chemistry WebBook Thermochemistry, IR, EI-MS n/a (not checked) Reference IR/MS, physical data Web only NIST Standard Reference Data (copyrighted)

What is PubChem? The largest free chemical database

PubChem is the world's largest free chemistry database, run by the US National Center for Biotechnology Information (NCBI). It aggregates data from more than a thousand contributors into linked collections: Substance (what depositors submitted), Compound (unique standardized structures), BioAssay, Protein, Gene, Pathway and Patent.

Plain-language query: "Give me the formula, molecular weight, SMILES and InChIKey of aspirin."

https://pubchem.ncbi.nlm.nih.gov/rest/pug/compound/name/aspirin/property/MolecularFormula,MolecularWeight,SMILES,InChIKey/JSON

Real response (trimmed):

{"PropertyTable": {"Properties": [{
  "CID": 2244,
  "MolecularFormula": "C9H8O4",
  "MolecularWeight": "180.16",
  "SMILES": "CC(=O)OC1=CC=CC=C1C(=O)O",
  "InChIKey": "BSYNRYMUTXBXSQ-UHFFFAOYSA-N"}]}}

Heads-up: PubChem has renamed its SMILES properties. Asking for the old CanonicalSMILES still works, but the answer now comes back labelled ConnectivitySMILES (no stereochemistry). Ask for SMILES if you want the full, stereo-aware string.

More tested PubChem queries:

What is ChEMBL? The go-to bioactivity database

ChEMBL is a manually curated database of bioactive, drug-like molecules, maintained by EMBL's European Bioinformatics Institute (EMBL-EBI). Its core is measured activity data, such as IC50, Ki and EC50 values, extracted from journal articles and deposited datasets, each linked to a compound, an assay and a biological target.

Plain-language query: "Find aspirin in ChEMBL."

https://www.ebi.ac.uk/chembl/api/data/molecule/search.json?q=aspirin

The first hit is CHEMBL25 (ASPIRIN, max phase 4, i.e. approved), with SMILES CC(=O)Oc1ccccc1C(=O)O.

Plain-language query: "What potency values does ChEMBL hold for cetirizine (CHEMBL1000) on the human histamine H1 receptor (CHEMBL231)?"

https://www.ebi.ac.uk/chembl/api/data/activity.json?molecule_chembl_id=CHEMBL1000&target_chembl_id=CHEMBL231&pchembl_value__isnull=false

Real response (trimmed to two of the 12 records returned):

{"activities": [
  {"standard_type": "Ki", "standard_value": "14.0", "standard_units": "nM",
   "pchembl_value": "7.85", "target_pref_name": "Histamine H1 receptor"},
  {"standard_type": "Ki", "standard_value": "5.89", "standard_units": "nM",
   "pchembl_value": "8.23", "target_pref_name": "Histamine H1 receptor"}]}

PubChem vs ChEMBL: what is the difference?

PubChem is broad; ChEMBL is deep. PubChem tries to include every chemical anyone has deposited, so it is the best place to identify a compound. ChEMBL contains far fewer compounds but curates their biological activity in a consistent format (standard units, pChEMBL values, mapped targets), so it is the better place to ask "how potent is it, and on what?" In practice you use both: PubChem to confirm the structure and identifiers, ChEMBL to read the pharmacology.

What is ChEBI? A curated dictionary of small molecules

ChEBI (Chemical Entities of Biological Interest) is a manually curated dictionary and ontology of small molecules, run by EMBL-EBI. Curated entries have a definition, structure, synonyms and ontology links that say what the molecule is (e.g. "a member of the class of benzoic acids") and what role it has (e.g. "non-steroidal anti-inflammatory drug").

Plain-language query: "What is CHEBI:15365?"

https://www.ebi.ac.uk/chebi/backend/api/public/compound/CHEBI:15365/

The response returns "name": "acetylsalicylic acid", formula C9H8O4, monoisotopic mass 180.04226 and the definition text. A text search also works: https://www.ebi.ac.uk/chebi/backend/api/public/es_search/?term=aspirin.

What is ZINC? Billions of compounds for virtual screening

ZINC is a free database of commercially available ("tangible") compounds prepared for virtual screening, built by the Irwin and Shoichet labs at UCSF. ZINC-22 focuses on huge make-on-demand catalogues from Enamine, WuXi and Mcule, while its sister database ZINC20 covers smaller in-stock catalogues. ZINC15 is the older generation, which many tutorials still mention.

Plain-language query: "Find purchasable compounds similar to cetirizine." Paste the SMILES OC(=O)COCCN1CCN(CC1)C(c1ccccc1)c1ccc(Cl)cc1 into the similarity search at cartblanche22.docking.org.

Honest note: we could not script ZINC queries reliably while writing this guide. The ZINC15/ZINC20 pages sit behind a captcha, and the CartBlanche22 search tasks we submitted programmatically came back empty. The web interface is the recommended route for most users.

What is the RCSB PDB? 3D structures of proteins and ligands

The Protein Data Bank (PDB) is the single global archive of experimentally determined 3D structures of proteins, nucleic acids and their complexes; RCSB PDB is its US data centre and search portal. For chemists, its value is the bound ligands: you can see exactly how a small molecule sits in its binding pocket.

Plain-language query: "Show me the ligand record for aspirin (component ID AIN)."

https://data.rcsb.org/rest/v1/core/chemcomp/AIN

This returns the name 2-(ACETYLOXY)BENZOIC ACID and formula C9 H8 O4.

Plain-language query: "Which PDB entries contain aspirin?" (Search API; paste it into a browser, which URL-encodes the JSON for you):

https://search.rcsb.org/rcsbsearch/v2/query?json={"query":{"type":"terminal","service":"text","parameters":{"attribute":"rcsb_nonpolymer_instance_annotation.comp_id","operator":"exact_match","value":"AIN"}},"return_type":"entry"}

Real response (trimmed): "total_count": 8, with entries such as 1OXR, 1TGM and 3GCL. A chemical search posted with the SMILES of aspirin (service: "chemical") also returned AIN as the top match. You can load any of these entries by its PDB ID in the MolDraw protein viewer or open it at rcsb.org.

What is UniChem? Translating IDs between databases

UniChem is EMBL-EBI's identifier cross-referencing service: give it one structure identifier and it returns the matching IDs in other databases. It links records through the standard InChI/InChIKey, so it works whenever two databases hold the same structure.

Plain-language query: "Where else does aspirin's InChIKey appear?"

curl -X POST https://www.ebi.ac.uk/unichem/api/v1/compounds \
  -H "Content-Type: application/json" \
  -d '{"type":"inchikey","compound":"BSYNRYMUTXBXSQ-UHFFFAOYSA-N"}'

The response maps it to CHEMBL25, CHEBI:15365, PubChem CID 2244, DrugBank DB00945, PDB ligand AIN and Guide to Pharmacology 4139, among others.

How UniChem cross-maps one InChIKey to identifiers in many databases The standard InChIKey of aspirin, BSYNRYMUTXBXSQ-UHFFFAOYSA-N, sent to the UniChem API returns matching identifiers in other databases, including ChEMBL CHEMBL25, ChEBI CHEBI:15365, DrugBank DB00945, RCSB PDB ligand AIN, PubChem compound 2244 and Guide to Pharmacology ligand 4139. One InChIKey, many database IDs Real UniChem API result for aspirin (checked Oct 2026) STANDARD INCHIKEY (aspirin) BSYNRYMUTXBXSQ-UHFFFAOYSA-N UniChem CHEMBLCHEMBL25 CHEBICHEBI:15365 PUBCHEM CID2244 RCSB PDB LIGANDAIN DRUGBANKDB00945 GUIDE TO PHARMACOLOGY4139 Also returned: SureChEMBL, BindingDB, HMDB, FooDB, CompTox, DrugCentral, FDA SRS, Wikipedia and more.
UniChem maps aspirin’s InChIKey to its IDs in ChEMBL, ChEBI, PubChem, RCSB PDB, DrugBank and Guide to Pharmacology.

What is nmrshiftdb2? A free NMR spectra database

nmrshiftdb2 is an open, peer-reviewed web database of organic structures and their NMR spectra, with built-in ¹H and ¹³C shift prediction. Its core is fully assigned spectra, many with raw data.

Plain-language query: "Find measured ¹³C spectra of compounds containing a 4-chlorobenzhydryl group." Draw the fragment in the nmrshiftdb2 search tab and run a substructure search. For a quick estimate before you search, MolDraw's structure to NMR tool predicts a spectrum from your drawing.

What is MassBank? Experimental MS/MS spectra

MassBank is an open repository of experimental mass spectra of small molecules, contributed by laboratories worldwide. The examples here use MassBank Europe (massbank.eu). Each record describes one spectrum: compound, instrument, ionisation, collision energy and the peak list.

Plain-language query: "List MassBank spectra for caffeine."

https://massbank.eu/MassBank-api/records/search?compound_name=caffeine

This returned 134 record accessions. Fetching one, https://massbank.eu/MassBank-api/records/MSBNK-ACES_SU-AS000088, gives an LC-APCI Orbitrap MS2 spectrum of [M+H]⁺ with the precursor at m/z 195.088 and a major fragment at m/z 138.067, licensed CC BY.

What is COCONUT? The open natural products database

COCONUT (COlleCtion of Open NatUral producTs) is the largest open collection of natural product structures, curated by the Steinbeck group at Friedrich Schiller University Jena. Version 2.0 links each molecule to its source organisms, literature and stereochemical variants.

Plain-language query: "Show natural products similar to caffeine." Paste caffeine's SMILES CN1C=NC2=C1C(=O)N(C(=O)N2C)C into the structure search at coconut.naturalproducts.net.

What is SureChEMBL? Chemistry from patents

SureChEMBL is EMBL-EBI's free patent chemistry database: it text- and image-mines patent documents and extracts the chemical structures they mention. It is the easiest free way to ask "which patents contain this structure?"

Plain-language query: "Resolve the name cetirizine to a SureChEMBL compound."

https://www.surechembl.org/api/chemical/name/cetirizine

Real response (trimmed): "chemical_id": "4176", "inchi_key": "ZKLPARSLTMPFCP-UHFFFAOYSA-N". You can then search for patents containing that structure in the web interface.

What about DrugBank and the NIST Chemistry WebBook?

DrugBank is a detailed knowledge base of approved and investigational drugs, their targets, metabolism and interactions. It is free to browse, but downloads require an account; academic datasets are licensed CC BY-NC 4.0 for non-commercial use, and commercial use needs a paid licence. Treat it as "free to read, not free to reuse."

The NIST Chemistry WebBook (NIST Standard Reference Database 69) offers thermochemical data, gas-phase IR spectra and electron-ionisation mass spectra for many small molecules. It is free to search online, for example https://webbook.nist.gov/cgi/cbook.cgi?Name=caffeine&Units=SI, but the data are copyrighted and bulk redistribution needs permission from NIST.

Which chemical database should I use?

Your question Start with Then try
What is this compound? Name, formula, CAS, properties PubChem ChEBI
How potent is it, and on which target? ChEMBL PubChem BioAssay
What class is it, and what is its biological role? ChEBI PubChem
Which similar compounds can I buy or dock? ZINC-22 PubChem (vendors)
How does a ligand bind its protein? RCSB PDB ChEMBL (target data)
I have an ID from database A; what is it in database B? UniChem PubChem cross-references
What should its ¹H/¹³C NMR look like? nmrshiftdb2 MolDraw NMR prediction
Does my MS/MS spectrum match a known compound? MassBank NIST WebBook (EI-MS)
Is it a natural product, and from which organism? COCONUT ChEBI
Is this structure in a patent? SureChEMBL PubChem Patent
Details on an approved drug? DrugBank (browse) ChEMBL
Thermochemistry or reference IR? NIST WebBook PubChem

How do I use these databases with MolDraw?

Most databases search best with a machine-readable identifier rather than a name. MolDraw gives you those identifiers straight from a drawing:

  1. Draw the structure in the MolDraw editor (or paste a name with name to structure).
  2. Copy SMILES from the top options row. Copy as… in the context menu also gives InChI and MOL formats (see the SMILES chapter of the MolDraw course).
  3. Make an InChIKey with the SMILES to InChIKey converter. InChIKeys are the most reliable exact-match key across PubChem, ChEMBL, ChEBI and UniChem.
  4. Query: paste the SMILES into a database's structure search, or drop the InChIKey into one of the REST URLs above.
  5. Bring results back: convert a PubChem CID with the CID to SMILES converter or an InChIKey with InChIKey to SMILES, then paste it into MolDraw to edit.
Workflow: from a structure drawn in MolDraw to database results Four steps. One: draw the structure in the MolDraw editor. Two: copy SMILES from the editor, or convert it to an InChIKey with the MolDraw SMILES to InChIKey converter. Three: paste the identifier into a database search box or a REST URL such as PubChem PUG-REST, ChEMBL or UniChem. Four: read the results, such as properties, bioactivity, 3D structures, spectra or patents. Draw → identify → query → read How a MolDraw structure becomes a database search 1 Draw Sketch or paste a structure in the MolDraw editor 2 Get an identifier SMILES: CC(=O)Oc1cc… InChIKey: BSYNRYMU… Copy SMILES in MolDraw; make an InChIKey with the converter 3 Query PubChem PUG-REST ChEMBL API UniChem / RCSB / others Search box or REST URL 4 Read results • properties & synonyms • bioactivity (IC50, Ki) • 3D complexes • NMR / MS spectra • patents, vendors
Workflow: draw in MolDraw, copy SMILES or make an InChIKey, query a database, read the results.

Worked example: how do I compare cetirizine analogues?

Suppose you are studying the antihistamine cetirizine and want to know which close analogues exist, how potent they are, and whether you could buy new ones. Here is a three-database workflow; every number comes from the live queries we ran.

Step 1: Identify it in PubChem. Draw cetirizine in MolDraw, copy the SMILES, and look it up:

https://pubchem.ncbi.nlm.nih.gov/rest/pug/compound/name/cetirizine/property/MolecularFormula,MolecularWeight,InChIKey,XLogP,TPSA/JSON

PubChem returns CID 2678, C21H25ClN2O3, MW 388.9, XLogP 1.7, TPSA 53 and InChIKey ZKLPARSLTMPFCP-UHFFFAOYSA-N.

Step 2: Find analogues in PubChem. The 2D similarity query (threshold 90) returns relatives including hydroxyzine (CID 3658, cetirizine's parent drug) and levocetirizine (CID 1549000, its single enantiomer), plus their salts. A substructure search on the chlorobenzhydryl-piperazine core adds chlorcyclizine, meclizine and buclizine.

Step 3: Compare activity in ChEMBL. Look up the InChIKey in ChEMBL:

https://www.ebi.ac.uk/chembl/api/data/molecule.json?molecule_structures__standard_inchi_key=ZKLPARSLTMPFCP-UHFFFAOYSA-N

This gives CHEMBL1000 (max phase 4, first approved 1995). Its activity records include human H1 receptor Ki values between 5.89 and 79 nM (pChEMBL 7.1–8.2) and a much weaker hERG (KCNH2) IC50 of about 30 µM (pChEMBL about 4.5). That window of roughly three orders of magnitude is the kind of selectivity comparison ChEMBL makes easy. A similarity search, https://www.ebi.ac.uk/chembl/api/data/similarity/ZKLPARSLTMPFCP-UHFFFAOYSA-N/80.json, lists levocetirizine and cetirizine salt forms to compare in the same way.

Step 4: Look for purchasable analogues in ZINC. Paste the SMILES into CartBlanche22's similarity search to see make-on-demand compounds that vary the acid side chain or the aryl rings. Then check any interesting hit's properties in MolDraw, for example with the Lipinski rule calculator.

Interpretation tip: activity values from different papers are measured in different assays. Compare pChEMBL values for the same target and assay type, and treat single measurements with caution.

How do I search chemical databases well?

Key takeaways

  • PubChem is the largest free chemical database and the best first stop to identify a compound.
  • ChEMBL is the best free bioactivity database: curated potency data linked to targets.
  • ChEBI classifies molecules; UniChem translates IDs between databases.
  • ZINC-22 is for virtual screening and buying analogues; RCSB PDB is for 3D binding modes.
  • nmrshiftdb2, MassBank and the NIST WebBook cover NMR, MS/MS and IR/thermochemistry.
  • COCONUT covers natural products and SureChEMBL covers patents.
  • Draw in MolDraw, copy the SMILES, generate an InChIKey, and you can query all of them.

FAQ

What is the largest free chemical database?

PubChem is the largest free chemical database, with about 124 million unique compounds and 349 million deposited substances (checked October 2026). ZINC-22 lists more molecules (about 54.9 billion), but most of them are virtual make-on-demand compounds for screening rather than characterised substances.

What is the difference between PubChem and ChEMBL?

PubChem is a broad aggregator that tries to cover every deposited chemical, so it is best for identification and properties. ChEMBL is smaller but manually curated for bioactivity, with standardised potency values linked to targets, so it is best for structure–activity and drug-discovery questions.

Is ZINC free for commercial use?

According to the UCSF ZINC licence, ZINC is free to use and download for everyone, including companies. The restriction is redistribution: you may not redistribute major portions of ZINC without written permission from the ZINC team. The compounds themselves must be bought from the listed vendors.

Which free database has NMR spectra?

nmrshiftdb2 is the main free NMR database, with 70,030 measured spectra for 271,817 structures (checked October 2026) and built-in ¹H and ¹³C prediction. PubChem compound pages also link to spectra from other sources, and the NIST Chemistry WebBook covers IR and mass spectra rather than NMR.

Which free database has mass spectra?

MassBank holds 139,006 experimental mass spectrum records (checked October 2026), mostly MS/MS spectra of small molecules, under per-record Creative Commons licences. The NIST Chemistry WebBook offers electron-ionisation mass spectra for many small molecules online.

Can I search chemical databases by structure?

Yes. PubChem, ChEMBL, ChEBI, ZINC, RCSB PDB, nmrshiftdb2, COCONUT and SureChEMBL all support structure search, usually by drawing or pasting SMILES. You can draw the molecule in MolDraw, copy its SMILES, and paste it into the database's structure search, or use exact, substructure or similarity queries through their APIs.

Which chemical database is best for students?

Start with PubChem for names, properties and safety data, ChEBI for definitions and compound classes, and the RCSB PDB for 3D structures. All three are free, need no account and have clear web interfaces. Add ChEMBL when you study pharmacology, and nmrshiftdb2 when you learn NMR.

What is a bioactivity database?

A bioactivity database stores measured effects of compounds on biological targets, such as IC50, Ki or EC50 values, linked to the assay and protein. ChEMBL is the leading free bioactivity database; PubChem BioAssay holds screening results deposited by laboratories and screening centres.

Is DrugBank free?

DrugBank is free to browse online, and academic users can download datasets under a CC BY-NC 4.0 licence after registering. Commercial use, including use in commercial products or services, requires a paid licence from DrugBank.

How do I convert a PubChem CID into a ChEMBL ID?

Use UniChem. Look up the compound's InChIKey (PubChem shows it on every compound page), then send it to the UniChem compounds API, which returns the matching ChEMBL, ChEBI, DrugBank, PDB and other identifiers. For aspirin, InChIKey BSYNRYMUTXBXSQ-UHFFFAOYSA-N maps CID 2244 to CHEMBL25.

Sources

Rafeeque Mavoor, product and development lead at MolDraw. Figures were checked against official sites and live APIs in October 2026; see our editorial guidelines.