Recursion

Recursion Pharmaceuticals: The Company That Photographs Sick Cells for a Living — and Why That Might Be a Trillion-Dollar Idea

| healthcare
↓ .md Take this into your AI — the full analysis + graph as markdown, ready to paste into ChatGPT, Claude, Gemini or any AI.

Based on 29 related nodes across 8 research explorations in the AI sector


What Does Recursion Actually Do?

Imagine you’re trying to figure out which key opens a lock, but instead of trying each key one at a time, you take a photograph of millions of locks simultaneously and teach a computer to recognize which ones look like they’ve been opened. That’s roughly what Recursion does — except the locks are diseased human cells, and the keys are potential drugs.

Every week, Recursion’s machines generate more than 10 terabytes of photographs of cells — what they look like when they’re healthy, when they’re sick, and when various chemical compounds are introduced to them. Over years, this has produced one of the most detailed biological image libraries ever assembled. The company then uses AI to find patterns in those images that suggest which compounds might work as medicines.

This approach is called “phenomics” — studying what things look like and how they behave (their “phenotype”) rather than just studying their genetic code.


The Mapmaker Analogy

Here’s a useful way to think about Recursion’s structural position: they are the mapmakers.

When European explorers were navigating the world, the companies that made accurate maps didn’t need to own the ships, the trade routes, or the ports. They just needed everyone who did own those things to buy their maps. Recursion’s phenomics data functions similarly. Other AI drug discovery efforts — including the most sophisticated protein-folding models in the world — need biological context that their tools alone cannot provide. Recursion generates that context at industrial scale.

The research graph underlying this analysis shows one of the strongest relationships in the entire dataset pointing in Recursion’s direction: the emerging concept of “virtual cells” — computer simulations that could replace animal testing entirely — depends on Recursion’s platform as a primary data source. That’s not a partnership or a contract. That’s a supply chain relationship. Recursion doesn’t need to win every drug discovery competition; it needs to be the raw material supplier to whoever does win.


Why Their Data Is Hard to Copy

The most non-obvious finding in this analysis is that Recursion’s real competitive advantage is not their AI models. AI models are becoming increasingly easy to build, fine-tune, and replicate. The advantage is the data those models are trained on.

You cannot reverse-engineer a decade of cellular imaging experiments from public scientific literature. You cannot scrape it from the internet. You cannot generate synthetic substitutes that carry the same biological validity. If you want to build a virtual cell model — which the pharmaceutical industry is increasingly desperate to do, especially as regulators move away from requiring animal testing — you need training data that looks like Recursion’s library.

This is why the data is described as a “moat” — something that protects a castle from attack. The castle is Recursion’s business. The moat is the irreplaceable, proprietary, continuously growing dataset that took years and hundreds of millions of dollars to accumulate.


The Closed Loop: What the Exscientia Merger Changed

In 2025, Recursion merged with a company called Exscientia. Before that merger, Recursion was excellent at identifying which biological targets might respond to a drug — but they needed partners to then design and synthesize the actual drug molecules.

After the merger, they can do both. The full cycle now works like this: photograph sick cells, find patterns suggesting a target, design a molecule computationally, synthesize it in-house, test it against the cells, observe the result, update the models, repeat. This loop runs continuously without depending on outside companies at critical steps. In manufacturing terms, this is called vertical integration. In drug discovery terms, it means the timeline from “interesting pattern in cells” to “compound ready for animal-free testing” compresses dramatically.


The Strengths

The data flywheel compounds over time. The more data Recursion collects, the better their models get. The better their models get, the more useful their platform is to pharmaceutical partners. The more pharmaceutical partners use the platform, the more data gets generated. This self-reinforcing cycle is hard to interrupt once it reaches a certain scale, and Recursion has been running it longer than almost anyone.

They are inside the pipeline, not competing with it. Recursion’s platform integrates with AlphaFold3 — the dominant protein structure prediction tool, owned commercially by Google’s Isomorphic Labs — rather than trying to replace it. Structure prediction tells you the shape of a protein. Phenomics tells you what happens to a living cell when you perturb it. These are complementary, not competing, questions. Recursion has positioned itself as a required component of the combined answer.

Regulation is moving in their direction. The FDA announced in 2025 that it is phasing out the requirement for animal testing in certain drug development stages. This is enormously significant for Recursion. If regulators accept AI-generated cellular data as evidence of safety — a transition that is now happening in policy terms — Recursion’s library moves from “useful research tool” to “required regulatory filing input.” That changes the monetization calculus entirely.


The Vulnerabilities

They have demonstrated the problem, not just solved it. This is the most uncomfortable finding in the analysis. The research graph describes Recursion’s own platform as something that “demonstrates” the clinical translation gap in AI drug discovery. That gap refers to a pattern that has emerged across the industry: AI systems that produce brilliant predictions in the lab consistently underperform when those predictions are tested on actual patients in clinical trials.

The concern is not that Recursion is uniquely bad at this. The concern is that the gap is real and difficult to close — and Recursion, despite having one of the most sophisticated platforms in existence, has not yet proven they can consistently bridge it.

Their computers depend on a supply chain with one fragile point. Recursion runs on NVIDIA graphics processing units. NVIDIA chips are manufactured almost exclusively on specialized equipment in Taiwan. The analysis traces a direct chain from geopolitical risk in the Taiwan Strait, through NVIDIA’s supply constraints, to Recursion’s ability to generate 10 terabytes of data per week. If that supply chain is disrupted — by trade war, conflict, or export controls — Recursion’s data flywheel stops. Structure-based competitors, who require less continuous raw compute, would be comparatively unaffected.

A better-funded competitor is building toward the same destination. Xaira Therapeutics launched with $1 billion in committed capital and a Nobel Prize-winning co-founder. Their explicit goal is to build the same kind of virtual cell foundation model that Recursion’s data currently underpins. Recursion has a first-mover advantage in data accumulation, but Xaira has more capital and scientific prestige. The question of how much of a lead Recursion’s head start actually represents is genuinely unresolved.


Bull Case: Why This Could Work Spectacularly

The optimistic scenario connects three things happening at once.

First, regulators formally accept AI-generated cellular evidence as a substitute for animal testing. This is already moving from policy statement toward implementation. Second, virtual cell models mature enough that pharmaceutical companies begin licensing or training on them at scale. Third, Recursion’s phenomics library is established as the primary training data source for those models.

If that chain completes, Recursion transforms from a drug discovery company — which must prove clinical results to generate value — into a biological data infrastructure company, where every pharmaceutical company building an AI platform is a potential customer for the underlying data. This is a fundamentally more durable and scalable business than winning individual drug races.

The analogy would be becoming the satellite imagery provider that every mapping application depends on, rather than building one mapping application and hoping it wins.


Bear Case: Why This Could Go Wrong

The pessimistic scenario is more straightforward. Recursion’s AI finds patterns in cells. Those patterns suggest drugs that should work. The drugs then fail in human trials, repeatedly, because human biology at the whole-organism scale is vastly more complex than cell images capture.

If several high-profile Recursion-derived programs fail clinically over the next few years, the narrative that phenomics data is the key to AI drug discovery becomes damaged. Pharmaceutical partners become skeptical. The data moat argument loses force if the data doesn’t produce clinically translatable insights.

This failure mode is not hypothetical — it has already happened to other AI drug discovery companies. BenevolentAI’s highly publicized atopic dermatitis program failed in clinical trials despite strong computational predictions. The pattern is established enough that the research graph treats Recursion’s own platform as an illustration of the risk rather than a solution to it.

The worst-case version combines a clinical setback with a compute disruption — forcing Recursion to raise capital at a depressed valuation while their data generation is temporarily impaired.


The Leverage Points Worth Watching

The single highest-leverage question is whether the FDA formally establishes a pathway for phenomics-derived safety data to substitute for animal testing results. If that happens on a clear timeline, Recursion’s position crystallizes quickly. If it remains ambiguous, the value of their data moat remains theoretical.

The second question is whether Xaira builds competing phenomics infrastructure or focuses on different approaches. A $1 billion commitment going toward protein design and generative chemistry is a threat Recursion can absorb. That same capital going toward building a competing cellular imaging library would be an existential challenge to the moat thesis.


Bottom Line

Recursion has built something genuinely unusual: a proprietary library of biological data that other people’s AI models increasingly need. They occupy the mapmaker position in an industry where the maps are becoming essential infrastructure. Their data moat is real, their regulatory tailwinds are real, and their integrated pipeline is more complete than almost any competitor’s.

The central uncertainty is whether the data translates to medicine. A library of beautiful, high-resolution photographs of sick cells is only worth what it predicts — and the prediction that matters most, whether a drug will work in a human being, remains unsolved by everyone in this space, Recursion included. The bull case is that solving this problem is a matter of data volume and model sophistication, and Recursion is positioned to win that race. The bear case is that the problem is harder than that, and the confidence that large proprietary datasets solve clinical translation is the same confidence that has repeatedly disappointed the industry.

This is a high-conviction structural position with a genuine unresolved question at its center. That question — whether phenomics data bridges the gap from lab to clinic — is what the next five years of Recursion’s pipeline will answer.