Pharmacophore study

Reverse peptide mimetics

Short peptides that present the same three-dimensional pharmacophore as ligands already observed bound in co-crystal structures. Where a peptide reproduces what a drug presents, it is a starting point for rules of peptide mimicry in either direction.

Every co-crystal ligand with a measured potency was scored against all 3,368,420 capped peptides of one to five residues, with PharmCast pharmacophore fingerprints on both sides of the comparison. The forty closest pairs are below. Open any row for the superposition.

25 of the 38 distinct peptides occur as complete loops in real protein structures. Loops that mimic drugs shows seven of them superposed on their ligand in the conformation the loop adopts in its own crystal.

1,233co-crystal ligands scored
3,368,420peptides scored per ligand
0.916highest pharmacophore Tanimoto
22pairs at 0.800 or better
0.144median Morgan similarity across these forty
1,233co-crystal ligands scored
3,368,420peptides scored per ligand
0.916highest pharmacophore Tanimoto
22pairs at 0.800 or better
0.144median Morgan similarity across these forty

The forty closest pairs

Sorted by pharmacophore Tanimoto. Morgan Tanimoto is the two dimensional chemical similarity of the same pair. The z column is how far that peptide sits above the mean of all 3,368,420 peptides scored against that ligand. pActivity is the strongest measured value for that ligand against that protein among pIC50, pKi, pKd and pEC50, all on a negative log molar scale and all from records with an exact relation. Complex names a Protein Data Bank entry in which that ligand is bound to that protein. Click a row to open the superposition.

LigandPeptidePharmacophore similarityMorganzpActivityComplexTargetFamily
+3J4SGS0.9160.07611.86.734R5A +1P00918LyaseCarbonic anhydrase 2
+ZGWAMVP0.9130.1675.69.598GFUP0DTD1TransferaseReplicase polyprotein 1ab
+L6GQPI0.8980.0716.06.594RN0Q9BY41Epigenetic EraserHistone deacetylase 8
+6FJGCF0.8620.1227.59.145J20 +1P07900Cytosolic OtherHeat shock protein HSP 90-alpha
+FJCAWF0.8580.2474.47.646M0KP0DTD1TransferaseReplicase polyprotein 1ab
+KDQFAAF0.8550.1593.86.326ROTP00734 P09945ProteaseProthrombin, Hirudin variant-2
+WNNTLPF0.8540.1213.59.705B5PP45452ProteaseCollagenase 3
+FSPVIW0.8530.1394.19.081TU6P43235ProteaseCathepsin K
+MFUS0.8460.05072.86.572JDM +2Q9HYN5UnclassifiedFucose-binding lectin PA-IIL
+6FFFSV0.8330.1056.27.865J27P07900Cytosolic OtherHeat shock protein HSP 90-alpha
+ZOAPGY0.8220.1375.87.587MU3P00918LyaseCarbonic anhydrase 2
+287FFVA0.8160.1693.18.662RG6Q16539KinaseMitogen-activated protein kinase 14
+V9HCFG0.8110.1827.56.607OE5P25440Epigenetic ReaderBromodomain-containing protein 2
+779WVL0.8100.1274.07.804R93P56817ProteaseBeta-secretase 1
+UBYEIF0.8100.2605.07.783T74P00800ProteaseThermolysin
+FHRWMP0.8060.2104.68.106LZE +1P0DTD1TransferaseReplicase polyprotein 1ab
+8J9YFD0.8050.4673.47.467Q25 +1P12821ProteaseAngiotensin-converting enzyme
+V9KCFG0.8050.2006.57.307OE6P25440Epigenetic ReaderBromodomain-containing protein 2
+6GCLFP0.8040.1136.48.065J6LP07900Cytosolic OtherHeat shock protein HSP 90-alpha
+87MVFV0.8030.1316.28.005UEQO60885Epigenetic ReaderBromodomain-containing protein 4
+3YHYIA0.8010.1064.37.683V13P00760ProteaseSerine protease 1
+99EGFV0.8010.1677.16.645NU3Q92793Epigenetic WriterCREB-binding protein
+VZIFPPF0.7990.1503.46.228OTMP9WGR1ReductaseEnoyl-[acyl-carrier-protein] reductase [NADH]
+S54KFF0.7900.1783.38.893RM0P00734 P09945ProteaseProthrombin, Hirudin variant-2
+8RZFPVP0.7870.1373.77.305NAWP00746ProteaseComplement factor D
+U4VYPPV0.7850.1173.37.486WKAP00918LyaseCarbonic anhydrase 2
+QQCPHF0.7840.1197.28.528BM2O60674KinaseTyrosine-protein kinase JAK2
+MEYFWI0.7840.1553.57.523OWDP07900Cytosolic OtherHeat shock protein HSP 90-alpha
+P75YAI0.7830.0783.48.036YPWP00918LyaseCarbonic anhydrase 2
+HZKWVP0.7820.1214.99.806QEDP50579ProteaseMethionine aminopeptidase 2
+EOQVFT0.7810.1875.910.396G7AO43570LyaseCarbonic anhydrase 12
+M31HFF0.7780.1564.19.203RMLP00734 P09945ProteaseProthrombin, Hirudin variant-2
+3KULFPP0.7780.1424.07.414R92P56817ProteaseBeta-secretase 1
+6KEFPVI0.7770.1113.38.305J8ZP00918LyaseCarbonic anhydrase 2
+FC4FPVI0.7770.1113.38.305J8ZP00918LyaseCarbonic anhydrase 2
+83PYTY0.7760.1583.29.265U9DQ06187KinaseTyrosine-protein kinase BTK
+S29FVH0.7760.1884.18.283RLYP00734 P09945ProteaseProthrombin, Hirudin variant-2
+IVQPFM0.7760.1506.36.527ZIIP17752ReductaseTryptophan 5-hydroxylase 1
+M32FFH0.7750.1744.28.663RMMP00734 P09945ProteaseProthrombin, Hirudin variant-2
+S28HVF0.7740.1463.98.433RLWP00734 P09945ProteaseProthrombin, Hirudin variant-2

Known biological motifs, scored the same way

Seventeen peptide motifs with documented roles in biology or drug design, looked up by hand and pushed through the identical fingerprint against the same 1,233 ligands. Every z is measured against the same 3,368,420 peptide distribution used everywhere else on this page. The z states how far the value sits above the mean of that ligand’s own scores, so a high similarity on a motif that matches everything cannot pass for a result.

MotifWhat it isBest ligandComplexTargetPharmacophore similarityz
LQSSARS-CoV-2 main protease substrate register, P2 to P1 primeL6G4RN0Q9BY41Histone deacetylase 80.8725.8
AVPFSmac tetrapeptide variant used in IAP antagonist seriesWNN5B5PP45452Collagenase 30.8303.3
AVPISmac and DIABLO N terminal tetrapeptide, the epitope behind the Smac mimetic drug classZGW8GFUP0DTD1Replicase polyprotein 1ab0.8094.7
PPPYWW domain binding motifZOA7MU3P00918Carbonic anhydrase 20.8005.5
GFLGCathepsin B cleavable linker used in antibody drug conjugatesWNN5B5PP45452Collagenase 30.7863.0
FFDiphenylalanine, the minimal aromatic self assembly motifYPB8V9FO60885Bromodomain-containing protein 40.7447.6
ATPFHtrA2 and Omi N terminal IAP binding motifU4V6WKAP00918Carbonic anhydrase 20.6972.6
LVPRThrombin cleavage recognition, P4 to P1A1LZZ8YMEO60885Bromodomain-containing protein 40.6243.2
LDVFibronectin CS-1 motif recognized by integrin alpha4 beta1ZGW8GFUP0DTD1Replicase polyprotein 1ab0.5752.6
KLVFFAmyloid beta 16 to 20, the self recognition motif targeted by aggregation inhibitorsS543RM0P00734 P09945Prothrombin, Hirudin variant-20.5731.5
IKVAVLaminin alpha1 neurite outgrowth motifA1L098Z7EQ96KQ7Histone-lysine N-methyltransferase EHMT20.5491.6
YIGSRLaminin beta1 cell adhesion motifS543RM0P00734 P09945Prothrombin, Hirudin variant-20.5431.3
GPRPFibrin knob A, the polymerization motif9UN4B7PP07900Heat shock protein HSP 90-alpha0.5193.7
NGRAminopeptidase N tumor homing motif9UN4B7PP07900Heat shock protein HSP 90-alpha0.4933.3
DEVDCaspase 3 and caspase 7 cleavage recognition sequenceGHC3GHCP00374Dihydrofolate reductase0.4892.0
PHSRNFibronectin synergy site9UN4B7PP07900Heat shock protein HSP 90-alpha0.4222.2
RGDIntegrin recognition motif of fibronectin and vitronectin9UN4B7PP07900Heat shock protein HSP 90-alpha0.3681.4

Two rows carry independent confirmation. Nirmatrelvir was designed to occupy the main protease substrate register, and the Smac tetrapeptide Ala-Val-Pro-Ile is the epitope that became birinapant and LCL161; both reach the top of their distributions against it. The adhesion and self assembly motifs at the foot of the table are the counterweight: RGD, PHSRN and NGR recognize protein surfaces rather than the enclosed pockets these ligands occupy, and they score accordingly.

Designing from peptide epitopes is established practice

Two bodies of work document taking a short peptide out of a folded protein and turning it into a drug-like molecule.

Protein epitope mimetics

Robinson and colleagues transplant a beta-hairpin loop sequence out of a folded protein onto a hairpin stabilizing template, typically the D-Pro-L-Pro dipeptide, producing a synthetic molecule that reproduces the loop’s conformation and its activity. The approach has produced ligands for CXCR4 and for the bacterial outer membrane protein LptD.

Acc Chem Res 2008 · Drug Discov Today 2008 · Design and applications of protein epitope mimetics · J Pept Sci 2013

Smac mimetics from a four residue epitope

The N terminal tetrapeptide of Smac and DIABLO, Ala-Val-Pro-Ile, binds a groove on the BIR3 domain of XIAP with affinity comparable to the full protein. Structure based design converted it into conformationally constrained small molecules, giving the clinical candidates birinapant, LCL161, GDC-0152 and AT-406.

Mol Cancer Ther 2014, birinapant

These peptides exist as real loops

25 of the 38 distinct peptides above occur as complete loops in real protein structures, out of 1,524,535 loops searched. Seven are superposed on their ligand in the conformation the loop actually adopts, with citations for each, in the companion page.

Open Loops That Mimic Drugs

Method

1. The ligand set

Co-crystal ligands were taken from Protein Data Bank entries in which the component is a drug-like small molecule bound to a characterized protein site. Selection keeps the drug-like organic ligands of those entries. Each ligand was matched to measured activity against its own protein through the InChIKey connectivity block, the first fourteen characters, against the Kinase Knowledgebase Q2 2026 release and ChEMBL 37. Ligand structures come from the Protein Data Bank chemical component dictionary, which is the same source that activity matching agrees with. 1,233 ligands carry a measured potency and were scored here.

2. The peptide library

Every sequence of length one through five over the twenty natural residues was enumerated: 20 plus 400 plus 8,000 plus 160,000 plus 3,200,000, or 3,368,420 peptides. Each was built as a capped peptide, acetyl on the N terminus and N-methylamide on the C terminus, which is what a fragment presented from within a chain looks like. Capping is a deliberate choice: a free peptide carries a charged N terminus and a charged C terminus whose ionic features dominate the fingerprint and inflate similarity against any charged ligand.

Peptide fingerprints do not depend on any ligand, so the library was computed once and reused for every comparison.

3. The fingerprint

PharmCast v10 converts a SMILES into a three dimensional pharmacophore fingerprint, 10,549 bits wide, by predicting what a hundred conformer ensemble would produce. It works from the two-dimensional structure alone. Each bit is one unordered triple of pharmacophore features together with the three binned distances between them, so the fingerprint records three dimensional feature geometry rather than substructure. The six feature types are hydrogen bond acceptor, hydrogen bond donor, positive, negative, aromatic ring and hydrophobic contact.

Both sides of every comparison are PharmCast predictions. Ligand and peptide pass through the identical model, so both carry the same systematic error and the similarity between them is like for like. Similarity is the Tanimoto coefficient over the 10,549 bits.

4. The search is exhaustive

Every one of the 3,368,420 peptides was scored against every ligand. Because the enumeration is complete, each reported best is the true global optimum for that ligand and the complete score distribution behind it is known.

That distribution is what makes a single similarity interpretable. Each pair therefore carries a z, the number of standard deviations its best peptide sits above the mean of all 3,368,420 scores for that same ligand. The forty pairs listed here run from z 3.1 to z 72.8.

5. The superposition

The fingerprint comparison is invariant to coordinate frame and produces no geometry, so the superposition shown in each expanded row is computed separately and stands as an independent check.

  1. Both molecules are embedded with ETKDGv3, twenty five conformers each, minimized with MMFF94, and the lowest energy conformer is kept.
  2. The peptide is fitted onto the ligand with Open3DAlign on Crippen contributions, which aligns on physicochemical character rather than on shared atoms. This is required by the data, since these pairs share almost no substructure and any alignment resting on a common scaffold would fail.
  3. Pharmacophore features are assigned to both molecules with the RDKit base feature definitions.
  4. A feature of one molecule counts as shared with a feature of the other when the two are the same type and their centers fall within 2.5 angstroms. Matching is greedy, closest pair first, so no feature is claimed twice.
  5. A shared triplet is any three of those agreements whose vertices are at least 2.5 angstroms apart, which excludes degenerate triangles. Triplets are ranked by how tightly the two molecules’ copies of the three vertices agree.

6. Reading the viewer

The ligand is drawn heavy and dark, the peptide light and thin, so the two are distinguished by weight as well as by tone. Feature overlap places one cloud at the midpoint of each agreeing feature pair, colored by feature type, so the marks show where the two molecules agree rather than either molecule’s own inventory. Top triplet and top five triplets draw the shared pharmacophore triangles with both molecules’ copies of every vertex, a hairline joining corresponding vertices and dashed triangle edges, so any disagreement between the two copies is visible rather than hidden. The Open3DAlign score reported on each pair is the quality of the fit, not the similarity.

7. Software and versions

PharmCast SCP v10, frozen 1 September 2026, trained on 5,887,229 molecules, fingerprint width 10,549. RDKit for structure handling, conformer generation, MMFF94 minimization, Open3DAlign, feature assignment and Morgan fingerprints at radius 2 and 2,048 bits. 3Dmol.js for the superposition viewer. Activity from the Kinase Knowledgebase Q2 2026 and ChEMBL 37.

Talk to us about peptide mimicry

The fingerprints are PharmCast, the ligands are the co-crystal set behind our structures and potency work, and the kinase activity is Kinase Knowledgebase data. We are happy to walk through the method and what it is useful for.

The method behind PharmCast is described in Muskal, S. M.; McGregor, M. J. PharmCast: rapid generation of three-dimensional pharmacophore fingerprints from two-dimensional structure without conformer generation. bioRxiv 2026, doi: 10.64898/2026.09.02.748999.