• Kinase Knowledgebase Interface
  • Kinase Foundation Model version 2: the potency model ranking two compounds against one kinase, and the selectivity model ranking two kinases for one compound
  • Dataset Overlap Analysis
  • The ChIP search cycle: gene, enumerate, score, select, vary
  • KKB per-kinase activity classifiers
  • KKB per-kinase potency regressors
  • KKB model performance rises with depth of curated data
  • KKB Sample Data
  • PDF Processing Workflow
  • Data Validation Process
  • Curation Search Interface
  • Oncology Knowledgebase
  • OKB Sample Data
  • ML Training Datasets

Advancing Drug Discovery Through AI-Powered Solutions

Eidogen-Sertanty is dedicated to improving healthspan, medicine, and general well-being through cutting-edge pharmaceutical research tools.


LEARN MORE

Our Products

The two version 2 models side by side. Left: which of two compounds binds this kinase more tightly - ligand A, the kinase sequence and ligand B enter one random forest in that fixed order, each ligand as a 1,024-bit Morgan count fingerprint plus 14 descriptors and the sequence as 480 ESM2 numbers, returning a confidence bar from A binds tighter to B binds tighter. Right: which of two proteins binds this ligand more tightly - sequence A, the ligand and sequence B in fixed order, returning a bar from protein A has greater affinity to protein B has greater affinity.

Version 2: potency ranking (left) and selectivity (right). Click the diagrams to visit kinasefoundationmodel.com

Kinase Foundation ModelNew!

Two models, one question each. Give the first a kinase sequence and two compounds and it says which binds more tightly. Give the second one compound and two kinases and it says which kinase the compound prefers, the selectivity question. Both read the kinase’s amino-acid sequence itself, with no structure, no docking and no binding-site definition, and both are trained on the Kinase Knowledgebase.

Tested on ChEMBL measurements the models never saw: 0.690 over 1,836,100 held-out potency comparisons across 477 targets, and 75.3% over 3,137,588 selectivity comparisons. Acting only on the confident calls raises those to 0.842 and 92.3%. Version 1, the original scorer across 478 kinases, remains published alongside them.

Explore the Kinase Foundation Model →
Kinase Knowledgebase: growth to more than 3 million biological activity data points Four-panel performance summary of the KKB per-kinase activity classifiers: ROC curves grouped by training-set size, the distribution of ROC-AUC across 392 kinases, accuracy rising with the number of measured compounds, and mean AUC by data-size band.

Per-kinase model performance across 392 kinases

Kinase Knowledgebase (KKB)

Currently the Kinase Knowledgebase Q2 2026 Release includes the following data:

  • Journal articles and patents: 10,326
  • Number of Biological Activity Data Points: 3,336,903
  • Number of unique kinase molecules with annotated assay data: 509,691
  • Number of all unique kinase molecules from patents and articles (with or without bio-activity data): 854,437
  • Number of unique kinase targets with assay data: 579
  • Number of annotated assay protocols: 105,916
  • Machine-learning models built from this release’s SAR: activity classifiers across 392 kinase targets, median ROC-AUC 0.91 (0.95 for well-studied kinases) – view model performance report

To show what this depth of curation supports, we build machine-learning models from the KKB SAR and report how they perform: an activity classifier for each of 392 kinase targets and a potency regressor for 319 of them. The panel at left summarises classifier accuracy: ROC curves grouped by how much data each kinase has, the spread of ROC-AUC, and how accuracy rises with the depth of measured chemistry.

Accuracy is highest for compounds chemically related to what a target already has in KKB and declines for novel scaffolds, so every prediction is reported with a similarity score against the model’s own training set. Performance is also measured against published data absent from the knowledgebase.

Search the Kinase Knowledgebase →
The ChIP search cycle in five stages: a gene holds a reaction sequence plus the exact building blocks for each reactant slot; enumeration runs every block combination through the sequence; your model filters, featurises and scores every product; elitist truncation keeps the best genes; crossover and mutation produce the next generation, and the cycle repeats.

How the search works. Click the diagram for the full method.

ChIP™ de Novo Design

Design molecules you can actually make. ChIP searches reaction space, not molecule space: it evolves synthetic protocols — a sequence of validated reactions plus the specific catalogued building blocks entering each one — and scores the molecules those protocols produce with your predictive model. Every design that comes out carries an executable synthesis from purchasable starting materials.

85 validated reaction transforms drawing on 853,409 catalogued building blocks, across 154 reactant slots. Potency, selectivity, an ADME or physicochemical property — any per-molecule model can drive the search, and every campaign is paired with a matched random control.

How ChIP works → See a worked example →
Two-panel diagram. Panel one: a compound, with its assay, primary target and potency, is reduced by a one-way SHA-256 hash to an irreversible encoded fingerprint, and the original structure is kept private. Panel two: two parties each hold a private list of encoded fingerprints, compare them, and learn only the percentage that overlaps.

Illustrative example. Figures shown are for explanation only and are not results from our datasets.

Dataset Overlap AnalysisNew

How much of a dataset do you already have? Our toolkit answers that for any two SAR datasets, in any target family, without either side revealing a structure. Each party encodes its own data locally into irreversible SHA-256 fingerprints; only fingerprints are compared, and only the overlap figure is shared.

How it works → Download the toolkit →
The Oncology Knowledgebase interface, showing a curated compound record beside a panel reading over 127,000 biological activity data points across 1,100 unique oncology targets.

Oncology Knowledgebase (OKB)

Currently the Oncology Knowledgebase Q2 2023 Release includes the following data:

  • Journal articles and patents: 977
  • Number of Biological Activity Data Points: 146,206
  • Number of molecules from patents and articles: 66,198
  • Number of unique oncology targets with assay data: 1,158
  • Number of annotated assay protocols: 5,614
  • Number of disease models: 137
Learn More →