Query 2,016,592 binding-affinity measurements in a public bioactivity index and one number jumps out: the top 15 protein targets — 0.23% of the 6,468 distinct targets in the data — account for 167,810 measurements, or 8.3% of the entire dataset. Widen the lens and the pattern holds at every scale: the top 100 targets (1.5% of all targets) carry 32.5% of all measurements, and the top 500 (7.7% of targets) carry 69.5%. Roughly 5,968 targets — the remaining 92% — share the last 30.5% of the data among them.
This isn't a data-quality bug. It's the fossil record of eight decades of pharmacology funding decisions, and it has a direct, underappreciated consequence for the AI drug-discovery models increasingly trained on exactly this kind of public affinity data.
What's actually at the top
The most-measured targets are a familiar list: JAK2 (14,457 measurements), EGFR (13,540), the hERG potassium channel (13,227, mostly cardiotoxicity screening rather than efficacy), the nuclear receptor ROR-gamma (11,683), JAK1 (11,570), VEGFR2 (11,447), BRD4 (11,246), BTK (11,001), BACE1 (10,480), CDK2 (10,166), HDAC1 (10,060), PI3K-delta (9,969), IRAK4 (9,845), the D(2) dopamine receptor (9,753), and the sodium channel Nav1.7 (9,366). These are targets with FDA-approved drugs already on the market — EGFR alone anchors erlotinib, gefitinib and osimertinib (FDA label, osimertinib) — which means every med-chem program that ever tried to beat, extend, or generic-around those drugs generated more affinity data for the same protein. It compounds: well-studied targets attract funding because they're well-characterized, which generates more data, which makes them easier to study, which attracts more funding.
Why this matters for AI models
Structure-based and sequence-based deep learning models for drug-target interaction prediction are trained largely on repositories like BindingDB, ChEMBL and KIBA. A 2024 analysis identified this exact skew as "target prior bias" — models trained on such imbalanced label distributions learn spurious target-specific shortcuts rather than genuine binding chemistry, degrading generalization the moment they're asked to score a target outside the well-trodden set (TAPB debiasing framework, PMC). Separately, a 2023 PMC study found that even state-of-the-art deep learning models for kinase inhibitor affinity — arguably the best-covered target class in the entire dataset — generalize poorly to kinases and chemotypes not already represented in training (Poor Generalization by Current Deep Learning Models).
Caveats
This index snapshots public bioactivity submissions, not the underlying druggable proteome directly — target identifiers can be split across isoforms, mutant variants and organism strains, so 6,468 keys somewhat overstates distinct biological targets. It also skews human (82% of measurements) with rat, mouse and viral targets filling most of the rest, and it says nothing about negative or failed binding data, which is chronically underreported in the underlying literature. Even so, against Finan et al.'s widely cited estimate of ~4,479 druggable human genes (Science Translational Medicine, 2017), the fact that measurement volume is this concentrated within even that smaller universe is the real finding.
As AI-native drug discovery scales, the targets most likely to get a well-calibrated model are the ones that least need one — they already have decades of chemical matter. The frontier where new models are supposed to add the most value, the thousands of thinly measured or entirely dark targets, is exactly where today's training data is weakest.