The AI race in biology has moved from models to measurement
A $1.8B coalition of Biohub, DOE, NIH and Big Tech is betting that the bottleneck for predictive biology is data, not algorithms.
On October 7, Biohub, the U.S. Department of Energy and the National Institutes of Health announced a $1.8 billion expansion of the Virtual Biology Initiative, which they call the largest coordinated commitment to generating AI-ready biological data to date. DOE will invest more than $500 million over five years. NIH will coordinate datasets built through more than $500 million in prior federal investment. Google DeepMind, Isomorphic Labs and Meta are collectively adding $300 million. Biohub's founding $500 million anchors the effort.
The Signal: Data Is the Scarce Input
Look at who is paying. Three of the best-known AI-for-science players are putting money not into bigger models but into producing measurements and handing them to everyone. Pushmeet Kohli of Google DeepMind described the goal as an open, standardized data commons. Isomorphic's Max Jaderberg said generating the data requires "scaling past the limits of what any single organization can produce today." Read together, the statements suggest these companies see the limiting factor for predictive biology as data, and that it is too expensive for any one of them to produce alone. Note the caveat: this is how the participants frame it, and the announcement offers no evidence that data is the only constraint.
What a "Virtual Cell" Actually Is
The stated goal is a predictive model of how a cell behaves, so researchers can run experiments digitally before, or instead of, running them at the bench. NIH's Nicole Kleinstreuer described aiming for "universal cell models" able to predict how any cell responds to an intervention.
Such a model learns from examples of cause and effect: change something about a cell, such as adding a drug or disabling a gene, and record what happens. The announcement says the initiative will expand cell response data to interventions across far more cell types and conditions than have yet been studied. Today's data covers a thin slice of that space, so a model trained on it can only interpolate within what it has seen.
The other half of the plan is instruments. Biohub's $400 million for new technology covers cryo-electron tomography, which resolves near-atomic detail inside the cell, and microscopy meant to image millions to billions of cells in living tissue. DOE's contribution draws on exascale supercomputing, X-ray and neutron scattering, and autonomous laboratories across the National Laboratory system.
The unglamorous piece may matter most. Biohub says it is building the layer that makes partners' datasets work together: shared standards, common identifiers and a single point of access. Data from different labs is often incompatible in format and labeling, which makes it hard to train on. Standardizing it is a prerequisite for the models, not an afterthought.
What It Means for Businesses and People
The source supports few near-term claims. The resource is described as open to the research community, which could lower the cost of entry for smaller biotech teams that cannot fund large-scale measurement themselves. The stated payoff, per Kleinstreuer, is substantially faster timelines for medical breakthroughs compared with lab experiments alone. That is an aspiration from a participant, and no timeline or deliverable date is given. Anyone building on this should treat it as a multi-year infrastructure project, not a product.
Questions You Should Be Asking
- What does "open" mean in practice: who can download the data, under what license, and can the companies funding it train proprietary models on it with no obligation to share what they build?
- Who validates that a virtual cell's predictions are right, and what accuracy threshold would count as success? The announcement names none.
- The $1.8 billion combines new money, prior federal spending and in-kind data and computation. How much is newly committed cash?
- NIH's contribution relies on existing datasets. How consistent are they, and who bears the cost of standardizing them for AI training?
- If three companies that compete on drug discovery co-fund the commons, what advantage do they keep, and what stops that advantage from shaping which data gets generated first?
What To Watch Next
The signal to track is the first public release of standardized datasets and the access terms attached to them. If shared identifiers, a single access point and permissive licensing arrive on schedule, the data-first thesis gains credibility. If releases stall or access narrows, the $1.8 billion headline will have described intent more than infrastructure.
- 1If you build in biotech, watch the initiative's data standards and common identifiers early, since aligning your own datasets to them will be cheaper than retrofitting later.
- 2Before relying on any virtual-cell model, ask what cell types and interventions it was trained on, because predictions outside that range are extrapolation.
- 3Read the license terms on any released data before building a commercial product on it.
Ready to implement AI in your business?
Our team builds the AI systems you just read about. Start with a free 30-minute discovery meeting.
