GSK bets $110M on Relation's biological data — not its AI model
What's the deal? Relation Therapeutics has struck a research collaboration with GSK worth up to $110 million in upfront and success-based milestone payments. Relation will generate large-scale human cellular perturbation datasets covering immunology, inflammation, fibrotic diseases, and osteoarthritis. Alongside the deal, it unveiled MORGAN, its cellular biology foundation model.
What's the twist? GSK is not licensing MORGAN. It is paying Relation to manufacture the biological training data that will make the model work. In AI drug discovery, pharma companies usually pay to use a biotech's platform — this deal inverts that.
What's the endgame? MORGAN — Multi-Omic Regulatory Genomics using Artificial Neural Networks — is designed to predict how any cell responds to a drug or genetic change, across cell types and diseases. The petascale datasets Relation generates will train it and probe biological pathways.
Why now? The race to build a "foundation model of the cell" has drawn well-funded rivals including Isomorphic Labs, Noetik, and insitro. They share one problem: data.
Why data is scarce. Protein structure prediction succeeded partly because the Protein Data Bank held more than 200,000 experimentally determined structures. No such repository exists for cellular perturbation data — the molecular readouts showing what happens across the genome, transcriptome, and proteome when a cell is disrupted.
Experts at the American Academy of Arts and Sciences framed the bottleneck in a May 2026 analysis: "Deep-learning models require data of a scale that is unprecedented in cell biology." Existing public libraries such as LINCS L1000 and Perturb-seq were built before AI training needs were understood.
What could go wrong? Bessemer Venture Partners noted in April 2026 that much public biological data lacks the traits machine learning needs: incomplete annotations, missing experimental context, siloed modalities, and batch effects that make cross-experiment comparison unreliable at scale. Alithea Genomics added that public datasets are "too heterogeneous, sparse in perturbations, and biased in their sample composition to robustly train models that reliably generalize." Relation's answer is to manufacture proprietary data rather than find better public sources.
The signal: The value in AI biology is shifting from the model to the data that trains it. By paying Relation to run automated cellular experiments at scale, GSK is buying into the scarcest resource in the field — a bet that whoever controls the training data controls the next generation of drug discovery.
Image credit: Generated with Gemini