Kinase Knowledgebase

Kinase Knowledgebase (KKB)

  • 3.34M+ bioactivity data points
  • 579+ kinase targets
  • Comprehensive SAR data
  • Integrated Ketcher editor
  • Advanced search capabilities
  • Activity classifiers for 392 kinases, median ROC-AUC 0.91
  • pIC50 potency regressors for 319 kinases
Four-panel performance summary of the KKB per-kinase activity classifiers: ROC curves grouped by training-set size, the distribution of ROC-AUC across 392 kinases, accuracy rising with the number of measured compounds, and mean AUC by data-size band.

Per-kinase model performance across 392 kinases

Oncology Knowledgebase

Oncology Knowledgebase (OKB)

  • 146,206+ biological activity points
  • Cancer-focused therapeutics
  • Clinical trial data
  • Drug resistance profiles
  • Mutation analysis tools
Family Foundation Model: the compound-preference model ranking two compounds at one target, and the target-preference model comparing two targets from different families for one compound

Family Foundation Model New!

  • The broad instrument: two separately fitted models over 34 protein families
  • Which of two compounds a target prefers, and which of two targets a compound prefers
  • 0.71 on 65,725 held-out comparisons, and 0.75 on 8,689
  • 2,079 targets and 1,879 targets, across 34 families
  • Trained on ChEMBL 37 alone, so both are downloadable
  • The GPCR and kinase models below are family deep dives beside it
The two GPCR models side by side: the potency model ranking two compounds at one target, and the selectivity model comparing two targets for one compound

GPCR Foundation Model New!

  • Two models: potency ranking and selectivity
  • Reads the target sequence; no structure or docking
  • Ranks which of two ligands is more potent at one target
  • Ranks which of two targets a ligand prefers
  • 0.77 on 762,493 held-out potency comparisons; 0.80 on 24,741 selectivity comparisons
  • 0.94 and 0.93 at prediction strength 0.70 and above
  • A deep dive into one family, on ChEMBL together with Eidogen-curated GPCR data
  • 284 targets in the training data; 235 and 211 carry a measured held-out accuracy
Kinase Foundation Model version 2: the potency model ranking two compounds against one kinase, and the selectivity model ranking two kinases for one compound

Kinase Foundation Model New!

  • Two models: potency ranking and selectivity
  • Reads the kinase sequence; no structure or docking
  • Ranks which of two compounds is more potent
  • Ranks which of two kinases a compound prefers
  • 0.69 on 1.8M held-out potency comparisons; 0.75 on 3.1M selectivity comparisons
  • 0.88 and 0.92 at prediction strength 0.70 and above
  • A deep dive into one family: primary measurements from the Eidogen-Sertanty Kinase Knowledgebase, with ChEMBL used to cross-validate
Five panels showing two molecules compared through predicted pharmacophore fingerprints

PharmCast and PharmSim™ New!

  • Predicts a full 10,549-bit PharmPrint™ from flat structure
  • No conformer generation at any point
  • A complete comparison of two molecules in 0.584 milliseconds
  • Against 5.7 seconds for the real calculation
  • Agrees with it at Pearson 0.969
  • True nearest neighbor in the top 100 for 95.3% of queries
Every co-crystal ligand in the Protein Data Bank embedded by its predicted pharmacophore fingerprint, colored by target class

Reverse Screen New!

  • Given a molecule, the proteins it might interact with
  • 27,797 co-crystal ligands indexed by PharmCast fingerprint
  • A query fingerprinted in 4 milliseconds, compared in 40
  • Docking runs only on what retrieval returns
  • Against 40.5 hours to dock every characterized site
The ChIP search cycle in five stages. 1 GENE: a reaction sequence plus the exact building blocks for each reactant slot, one to three steps deep, drawn as chains of colored shapes. 2 ENUMERATE: every block combination is run through the sequence in both orientations, fanning out to many product chains. 3 SCORE: your model filters, featurizes and predicts for every product, and the scored cloud narrows through a funnel from better to best. 4 SELECT: elitist truncation keeps the best genes, selection acting on genes rather than on molecules. 5 VARY: crossover swaps a step's blocks or splices two routes to produce the next generation.

Swipe the diagram → · tap it for the full method

ChIP™ de Novo Design New!

  • Searches reaction space, not molecule space
  • 85 validated reaction transforms
  • 853,409 cataloged building blocks
  • Runs against your model: potency, selectivity, ADME
  • Designs the non-obvious “me-too”: high pharmacophore similarity, low 2D similarity
  • Every design carries its full synthetic route
  • Potency scoring, or a 3D pharmacophore scaffold hop
  • Matched random control on every campaign
Two-panel diagram. Panel one: a compound, with its assay, primary target and potency, is reduced by a one-way SHA-256 hash to an irreversible encoded fingerprint, and the original structure is kept private. Panel two: two parties each hold a private list of encoded fingerprints, compare them, and learn only the percentage that overlaps.

Illustrative example. Figures shown are for explanation only and are not results from our datasets.

Dataset Overlap Analysis New

  • Works on any two SAR datasets, any target family
  • Neither side reveals a structure
  • Irreversible SHA-256 fingerprints, encoded locally
  • Only the overlap percentage is shared
  • Free downloadable toolkit
Predicted against real pharmacophore similarity, colored by two-dimensional similarity

Training Corpora

  • 4,612,044 compounds of the commercial screening collection
  • 135,768 distinct Bemis-Murcko scaffolds per 200,000 molecules
  • Median nearest-neighbor similarity 0.714, measured all against all
  • Plus 1,524,535 peptide loops read from protein structures
F1000Research data note describing the kinase validation sets, version 3

Kinase Validation Set

  • 258,000 structure-activity points, freely available
  • 76,000 unique structures across eight kinase targets
  • Extracted from the KKB, published as a peer reviewed data note

Professional Services

Custom Database Development

Tailored solutions for your specific research needs

Data Integration

Seamless integration with existing workflows

Training & Support

Comprehensive training and ongoing technical support

Consulting Services

Expert guidance for drug discovery projects

Ready to get started?

Contact us for pricing and licensing information