A roadmap for AI disruption in drug discovery
A review of how different business models integrate and a few proposals; Part I. Plus, thoughts on the fee-for-service cell biology cloud lab.
We are frequently asked by aspiring entrepreneurs something along the lines of “is it a good idea to…? ”; whilst the specifics of an answer are seldom general learnings, the broad brush business model sometimes can be — motivating the essay below.
Note this is posted in a personal capacity and this blog does not represent any positions of Relation or Lindus. However, it does include reflections made from conversations with talented colleagues and friends, including Henry Taysom, Patrick Collins, Kevin Foote, Adam Cribbs, and Peter Crane.
Many inefficiencies remain in the decade-long journey to bring a drug to market, and there exist many companies fulfilling various niches within the ecosystem. The story that everyone remembers is that, at the cost of billions of dollars per marketed drug, only one in ~20 drugs programmes is successful, but what about other, less known, endeavours? Who are the other companies that support this path to market? How do the billions get spent?
The “TechBio” portmanteau follows from the concept that a biotech can be built in the style of a software company. Given that, where are the opportunities for TechBio companies to improve the value chain? Having interacted with players both big and small employing a diverse range of business models, it is now time for a (draft) review of how this community fits together. As we progress, we also highlight clear gaps in the ecosystem waiting to be filled.
Broadly speaking, TechBio companies can be imperfectly categorised into: drug discovery (both target discovery and biomolecular design), life sciences tools (including automation), drug development groups, and contract research organisations (including manufacturing and clinical trials), and cutting across this all, some nascent but important SaaS presence. This essay is Part I focused on the discovery stage (above diagram), i.e., before late preclinical and before we enter the clinic. At a later time, we will write Part II on drug development, i.e., late preclinical and once you enter the clinic.
If there were a simple take home from the following essay, it is that it is all about the data generating process, either:
If you own a capability or the underlying IP, and there’s a great business model to be had if your technology truly unlocks novel biology.1
If your primary source of data is public, then the key opportunities relate to building and exploiting networks (and moving towards SaaS-style business models).
Finally, we also consider the incentives for cloud labs and explore why they do not necessarily align as a VC-backable proposition.
Drug discovery
Inevitably, for healthcare to work as an industry, someone has to sell a drug. Invariably, this position is served by pharmaceutical companies, who then sell to pharmacies and hospital networks,2 and increasingly even go direct to consumers (e.g., Eli Lilly launched LillyDirect). Some of this revenue ends up reinvested in R&D, but not all of these activities need to happen internally within pharma companies. Multibillion dollar transactions have minted what we refer to today as the drug discovery industry: the systematic creation and outlicence of potential medicines — or more specifically: their associated IP and regulatory portfolios, see graphic below for the changing dynamics of drug origination. Few drug discovery companies ever decide to go all of the way to market themselves, but partner to both lessen capital requirements and acquire a distributor for their eventual product.
To the layperson, you may imagine that one cooks up new molecules in a lab and then starts testing them in mice, and then perhaps move your favourite one into human trials. At some level this is true, but before we start designing molecules, we must have some specifications for what the molecule needs to do; that is, what biological mechanism the would-be drug manipulates — we refer to this as the drug target, which is often a protein-coding gene. Target-free processes do exist, (referred to as phenotypic drug discovery,3) and were employed to great success historically, but with our ever increasing understanding of patient heterogeneity, the clinical need for most diseases is typically discrete and specific (e.g., non-response to first line therapy, greater efficacy vs standard of care, improved safety profiles, that are challenging to model in vitro), requiring one to incorporate disease biology into the target medicine profile (performance criteria for the eventual drug). Only then do we start thinking about molecular structures.
This split is actually mirrored in how drug discovery organisations position themselves. Startups fall into two (not mutually exclusive) categories: biomolecular design companies (“chemistry” companies) and target discovery companies (“biology” companies). These capabilities do not solely exist within start-ups; pharmaceutical companies have such functions themselves, but in order to move into new disease areas, access the latest technology, or to expand their portfolio, partnering is a natural decision.
With that, let’s discuss Target Discovery.
Target discovery
After various retrospective analysis4 (e.g., AstraZeneca’s 5R’s framework), it can be concluded that many assets were brought into the clinic with data gaps, strategic missteps (or indeed, poor governance), and logical leaps of faith that could have been flagged earlier, resulting in clinical trial failure. Below, historical data show that the primary cause of these losses was due to lack of efficacy and safety signals,5 i.e., our understanding of disease biology was inadequate.6
There’s also been a few shocking case studies demonstrating that for many drugs in clinical trials (mostly limited to small molecules in oncology), our understanding of their mechanism was actually completely wrong and the preclinical data related to a different mode of action, see Science Translational Medicine article below.
These studies (along with many others) have led to the development of heuristics and standards for the target discovery process. Target discovery7 can be approached with various philosophies in mind, from using patient genetics (and other ‘omic technologies), to building disease models that can be perturbed using a range of tool molecules (ideally with well-characterised, known function) and genetic instruments (typically variations of the CRISPR technology). Ultimately, for any line of evidence, historical benchmarking studies can both support and refute the utility of any single technology, meaning that one has to synthesize evidence and “make sense of it all”.
Note, in a future post, Relation will release its “how to” philosophy for target discovery — subscribe to hear it first on the official Relation blog.
What’s been tried?
Target discovery companies broadly fall into 3 buckets:
Novel ‘omic technologies: Imagine you have invented some new way to measure cell behaviour or aspects of human physiology, such assays are what we refer to as ‘omics.8 For most ‘omic technologies, the rights holder does not have the freedom to operate, as their claim really sits under a broader patent family held by another party. With this in mind, instead of selling kits (sets of reagents), some founders decide to use the technology to discover drug targets and outlicence these findings to other parties.
Generally, these approaches have limited success: it’s hard to establish the value of one’s technology without selling it at scale, therefore a frequent solution is to move into drug development as quickly as possible so the company can become valued on their prospective medicines — not their unproved method.9 For examples of this, consider Enhanced Genomics and Nucleome Therapeutics, both own patents relating to the Hi-C protocol (or variations thereof) and operate a target discovery business model.
Therapy area expertise: Another option is to be more agnostic to the employed technology but build expertise in a collection of disease areas joined by some common theme (e.g., cell types, or pathway). Relation achieves this through the systematic (tacit) knowledge in collecting and characterising tissue associated with musculoskeletal disorders, immunology, fibrosis, and dermatology. If you are lucky (with the right clinical protocol), certain economies of scale emerge: Relation so frequently accesses resected tissue from a range of orthopedic surgeries that the effective costs have been drastically driven down (less experimental failure, scaled workflows, faster analyses). However, this comes with the cost of agility — suddenly initiating a new study to access a new disease area requires planning and creates friction. Other TechBio companies with this strategy include: Immunai (immunology), insitro (metabolic, neurology), Cellarity (hematology), Celsius Therapeutics (oncology, IBD; now part of AbbVie).
Amalgamation plays, or “the everything company”: This last bucket is somewhat the “other” category, groups like BenevolentAI who, whilst tied with Recursion as one of the first TechBio companies (and deserve to be credited as such), unfortunately developed technology so broad it is hard to know exactly when it can and can’t be employed. When a company cannot elucidate a narrative for what they’re good at compared to where they can only provide limited value, it is reasonable to conclude a lack of focus.
What could work?
The above states a clear preference for therapy area expertise, but there is value beyond just target discovery in such a strategy: can you reuse the data to find biomarkers to diagnose disease, find patient subpopulations, or even predict who will respond to treatment? In short, the data moat can become multipurpose. Case in point, Relation’s capability in orthopedics can clearly feed discovery efforts in at least 3 complex diseases (so far),10 but has also been designed in a manner to identify potential candidate biomarkers and the accompanying stratification logic.
Two big trends that seem to capture the VC imagination at the moment are the “Virtual Cell” and “AI scientist” concepts. Both areas are still nascent with many claims of victory that more often than not collapse when placed under scrutiny; we seldom have sufficient data (or indeed, the right data) to really bring either of these visions to life in their full generality.
In the case of Virtual Cell companies, there are early signs of life. Consider Noetik who build tissue foundation models for oncology that have clear translational logic. They can virtually perturb cells in silico and appear to recover reasonable gene expression profiles that suggests a certain self-consistency of the architecture, whilst also mapping between different imaging modalities of varying expense (H&E vs multi-omic layers); a technology both for target discovery and clinical practice. The successful companies in the space will be those able to learn the lessons above: choose specific disease areas, learn how to address real clinical shortcomings, and work out how to incorporate this into the underlying foundation model.
(Very recent news story: Relation and GSK have recently announced a partnership in this area!)
The universe of AI scientist companies is less convincing. In all likelihood “we can automate (broadly defined) scientific discovery” will not be addressable in a start-up context but will be partnerships between groups like OpenAI or Anthropic working with automation providers — see automation section.
Biomolecular design
Put simply, biomolecular design is the process by which we determine what specific molecule to progress by balancing a series of complex tradeoffs: from activity (on target vs off-target), immunogenicity (for biologics), to absorption, distribution, metabolism, and excretion (ADME) properties, and even commercial aspects including competition and freedom to operate (FTO) concerns and patent strategies. Various “AlphaFold-for-X” companies have addressed different modalities, including antibodies (BigHat Biosciences, Chai Discovery, Nabla Bio), cell therapy (coding.bio), capsid design (Dyno Therapeutics), of course small molecules (Exscientia, now part of Recursion, Insilico Medicine, Charm Tx, Genesis Molecular AI), and modality agnostic groups (Isomorphic Labs).
Given extended development timelines, as of today in 2026, there are no FDA-approved AI-discovered drugs. Moreover, any proposed molecule is often met with skepticism debating as to whether the design was truly novel or mimicking a known skeleton, or if there is any difference in associated clinical success. The long term view feels the most sensible: these tools appear to be reducing timelines and would we still expect to use the same methodology from the 1980’s decades from now?
What’s been tried?
In the beginning, the first round of biomolecular design companies were fast to raise capital, but perhaps a little slower to realise deals. Often overlooked, one of the biggest challenges in drug discovery is the extended timelines to close big pharma partnerships. As predicted, the world of biomolecular design is gradually becoming commoditised, or more specifically, a trend towards (admittedly expensive) fee-for-service models (see Chai Discovery – Eli Lilly deal) and open source repos (and Boltz associated PBC offering enterprise solutions). There’s a reason why this works: faster to execute, often without stipulations on how the product is used (for example, exclusivities on indications that frequently occur in pharma-biotech dealmaking).
What could work?
The challenge for biomolecular design companies is how to maximally exploit their technology. As more startups get formed, we will continue to see endless technology benchmarking studies as marketing material. For all but the most successful companies, there is a downward price pressure and we can expect the fee-for-service deal trend to continue.
However, these comparisons will never tell the full story, invariably each company may have some advantage in a unique modality or pathway. To avoid the conflict of interest (blanket selling a platform access vs. exploiting its unique advantages), single-asset spin outs will allow specialist executive teams to bring the eventual drug to market. This is the boring expansion opportunity.
In all likelihood, the big business opportunity is upselling and even roll ups. Amazon doesn’t just recommend 3rd party products, it actively looks to offer its own “basic” version, whilst keeping you hooked via a Prime subscription. What does this look like for biomolecular design? An AI company designs you an asset with certain characteristics — what else? Can they propose specific tests to derisk what they’ve created? For example, a specific toxicology study to understand an unusual side chain, a target engagement biomarker for the Phase I study, or a freedom-to-operate (FTO) search that identifies a worrisome patent. Recurring revenue could be secured through programme concierge services, e.g., intellectual property (IP) lifecycle management, and softly advertising acquisition to big pharma. Essentially, if the process of biomolecular design truly has been derisked, the next frontier for an AI-native company is handling the back office operations to bring a product to market.
Life science tools
Because clinical development is so expensive, the preclinical evidence packages are becoming ever more extensive: perhaps some new technology or workflow can further derisk a drug before entering the clinic? From experimental reagents, specialist genomic assays, to giant automated lab infrastructure, drug discovery companies rely heavily on life science tooling — but who to choose?
The core challenge for any tool is to get initial traction. Most of the time a new technology is caught between a rock and a hard place: if the solution has direct competitors, it must be price competitive because the customer will seldom want to change workflows or introduce batch effects into their data; the alternative is that it is a fundamental innovation, and therefore the customer will need to displace budget from elsewhere in their organization — and this displaced budget must be significant enough to keep the startup alive. In short, there is a huge activation energy required for new life science tools companies to become established.
Those companies that have established themselves are dominated by publicly listed multinationals, e.g., ThermoFisher, Becton Dickinson, Danaher. These companies are no longer known for their original products — mostly lab equipment — and sometimes date their founding back to pre-1900’s. In actuality, these conglomerates are multi-decade roll ups of many small companies with the dream that there would be synergies realized within their portfolio, see the graphic below taken from The Republic of Science blog post.
If the business model works as advertised, the concept is simple: identify best-in-class solutions, acquire upstart startups before their price becomes too dear, integrate their products into existing sales channels, and then save on operations via centralized resources.11 What distinguishes the good from the great, is that strong executives go beyond “our portfolio needs a proteomics solution to fill gap X” but ask why a very specific solution will achieve market dominance in years to come, or if there is an option to leverage capabilities within their portfolio. For an example of this, consider ThermoFisher’s acquisition of O-link proteomics: a highly sensitive dual-antibody solution that can leverage Thermo’s biologic production capabilities, but also it has a guaranteed customer base for years to come due to its flagship project with the UK Biobank (meaning it is the reference proteomics dataset for nearly all population genomics going forward). On the other hand, when Standard Biotools merged with SomaLogic (then valued ~$444M; Jan ‘24), it was hard to pluck out the logic from behind the deal; after 2 years SomaLogic was sold to Illumina for up to ~$425M (Jan ‘26).
Whilst one cannot hope to cover the breadth of life sciences tools companies, two areas worth doing a deep dive on are ‘omic technologies and automation.
‘Omic technologies.
If there were one unifying feature of all ‘omic technology companies, it would be the sheer amount of effort they expend building and defending their IP portfolio. As covered in the scTrends market review, much of the history of single-cell and spatial ‘omic technologies is characterised by almost constant litigation. In fact, very few individuals have working knowledge of who holds which patent, what claims are covered, and where there are gaps to exploit.12
A good example is the contested space around single-cell RNA-seq and chromatin-accessibility profiling. 10x Genomics sells Chromium Epi Multiome as an off-the-shelf assay for measuring gene expression and open chromatin from the same nucleus, directly linking RNA and ATAC profiles cell by cell. Parse Biosciences’ split-pool barcoding approach can, in principle, support the same broad modalities — RNA-seq, chromatin accessibility, and potentially paired RNA-plus-accessibility readouts — but its commercial path is shaped by a distinct platform architecture and a different set of patent constraints. This is where the IP picture becomes murky: a company may avoid one product claim while still running into a broader patent family covering a class of chemistries, barcoding strategies, enzymatic steps, or compositions of matter. In practice, the question is rarely “can you measure RNA and ATAC?” but rather by what particular molecular route — and under whose claims?
Strange and wonderful things are happening in the NGS-based spatial transcriptomics space: we are seeing multiple companies converge on very similar technologies but by distinct technical architectures and IP positions. 10x Genomics’ Visium, MGI/STOmics’ Stereo-seq, and Illumina’s emerging spatial technology all aim to produce spatially resolved transcriptome maps, yet they differ in how spatial barcodes are generated, immobilised, read out, and linked back to tissue coordinates. At the research end of the spectrum, methods such as Seq-Scope show that Illumina flow-cell chemistry can even be repurposed to create ultra-dense spatial barcode arrays. So the platforms may look similar at the level of data output (post bioinformatics), but they are not necessarily the same assay, nor necessarily covered by the same patent claims.
What’s been tried?
So, imagine you’re a new ‘omic technologies company, what can you commercialise? Without training as a patent lawyer, probably very little (unless you have an entirely new category of innovation).
In the space of single-cell ‘omics, there’s been remarkably little inroads made by startups vs 10x Genomics. When you run a multi-year study with 100’s of patients, you want consistency with an implicit promise that your supplier will be selling the same thing at the same quality for years. One may switch to a new product line for a new study, but these moments are hard to come by from a sales perspective. One notable break in the 10x sticky moat is Parse Biosciences (now part of QIAGEN), a great David vs Goliath story of some PhD students arguably beating the more “official” spin out from their former lab (ScaleBio, now purchased by 10x Genomics) and fending off a lawsuit(s) from 10x Genomics. A true tale of grit and perseverance to become the 2nd in the market.
Within spatial ‘omics and proteomics, the field is arguably more open: many of the foundational patents have long since expired (e.g., FISH, mass spec), but the newer generations of companies are often cautious to launch products until they’ve lined their ducks in a row with regards to FTO.
What could work?
As platforms take so long to reach maturity with questionable IP portfolios, many groups are staying in stealth for longer periods to engage with a customer base and perfect the product. There is a bit of a race between VCs funding the dream before a bigger player is needed — often to handle the legal battle about to ensue. Two-way licensing deals do occur, but usually this is the result of a dispute, not before it occurs.
The big trend to watch here is the tools vendors are vertically integrating into “selling data and disease insights” to pharma to drive foundation model work and internal target ID. Illumina have been doing this via their Billion Cell Atlas initiative (leveraging their recent acquisition of Fluent Bioscience’s PIP-seq solution) and Parse has been operating a similar model for their larger datasets. It’s likely that in the future many companies in the space such as 10X genomics, PacBio, Illumina, Thermo will have a (pre-)clinical insights division selling these outputs to pharma. The big question that many have raised over these deals is how does this sit with their primary model of all selling razors and razorblades?
A more interesting question long term is: was the 10x/Illumina dual solution a one off? Or are there patent exploits littered across the life science tools ecosystem? The answer is, undoubtedly, yes. From Our World in Data, below is a chart of the rate of pharmaceutical and life science patent filings across the globe.
The number of patent filings is inordinately huge, even though many are ultimately abandoned, especially where universities cannot justify expensive fees for non-revenue-generating assets. Others are later challenged at the Patent Trial and Appeal Board, where in 2024 roughly 78% of claims reaching a final written decision were found unpatentable; even so, there are still far more funded patents than actual new products on the market.
One area ripe for AI disruption is mass-scale patent busting. Historically, this has been the preserve of pharmaceutical companies, generic manufacturers, access-to-medicines groups, and utilising specialist patent lawyers: teams map patent estates, run freedom-to-operate analyses, search for prior art, and challenge blocking claims through litigation or patent-office proceedings. A useful example is Gilead’s sofosbuvir, sold as Sovaldi, a hepatitis C drug launched in the United States at about $84,000 per treatment course. In India, access groups and generic manufacturers challenged one of Gilead’s patent applications, arguing that it covered insufficiently inventive chemistry. The Indian Patent Controller initially refused the application, temporarily denying Gilead that monopoly right and strengthening the case for cheaper Indian-made versions. Gilead later won some Indian patent protection, so this was not a clean “patent busted, drug goes generic everywhere” victory. But it did help open market access: sofosbuvir became available from Indian manufacturers through a mixture of patent challenges, voluntary licences, local regulation, and country-by-country patent rules. The typical lesson is that patent busting rarely collapses an entire IP estate; more often, it removes or weakens one obstacle, creating leverage for cheaper products to enter specific markets. Could this work differently for ‘omics product development?
In pharmaceuticals, patent busting can create room for cheaper manufacture of an existing molecule; in ‘omics, it could enable entirely new assay architectures. Agents could decompose platform patents into their functional primitives — sample preparation, capture chemistry, barcoding strategy, enzyme choice, surface chemistry, sequencing readout, imaging readout, and computational registration — then search across papers, old protocols, conference posters, abandoned patents, instrument manuals, and archived product documentation for prior art or design-around routes. The output would not be an automated legal opinion, but a set of ranked claim charts and “white-space maps” for scientists and patent counsel: here are the claims that look weak, here are the ones that are probably blocking, and here are the molecular routes that appear underexplored. For a new ‘omic technologies company, that changes the question from “can we sell our new product” to “can we reach the same biological output through a sufficiently different technical path without walking through someone else’s claims?” In other words, patent busting as part of a product-discovery engine.
Automation
Automation is quite a different beast. Whilst large-scale automation projects are exceptionally expensive (~10’s of millions of dollars), they are mostly running on expired patents. Attending last year’s SLAS, many vendors are offering subtly different, but mostly very similar solutions — so much so that many people in the field cannot distinguish the relative merits between rigs. The high upfront costs and product similarity motivates “integrator” companies who constantly survey the market, find the best solutions, add them to their portfolio of supported devices, and put your lab together for you offsite, and then ship the whole installation to your facility ‘ready to go’ as a turnkey solution. Even with these integrators, the key operational challenges still relate to interoperability (and associated vendor lock in), coordination, and crucially — adaptability. Large integrated automation systems are typically built for a specific scientific endeavor, for example a high throughput screen. Adapting that setup as scientific and business needs change is challenging, slow, and costly.
For examples of interoperability challenges, consider how each vendor may use the same geometry for a 96-well plate (i.e., the positioning of the wells), but one solution may use a barcode labelling with a specific rim for the plate, another uses a numbering system combined with magnets, but then you want to dispense drugs via acoustic liquid handling and that needs a whole different solution. One liquid handler will use plastic tips sourced in the USA, but another uses a very slightly different plastic composition from Asia, and so on, and so on — a business model that is reminiscent of paper printers and printer cartridges.
Beyond straight automation, this really complicates bringing together multiple technologies. For example, company X offering some ‘omic technology (above) will make very specific claims on their performance when using liquid handler Y. Oh no! You have already bought liquid handler Z, then company X will say “you’re on your own, we’ve never tried this”. Now this is true (it’s not been tested in exactly this way), but it’s likely not an issue, but if you’re a pharmaceutical company and you’re investing in an automation build for some ‘omics technology, how much risk do you want to take on? You would prefer to only run as few QC pipelines internally as strictly necessary — remember, within big pharma, every project has to be documented properly with the correct SOP paperwork (often to maintain institutional knowledge and regulatory requirements).
On the dry-lab side, if you want to connect a device that an integrator company has never seen before, you will have to pay a premium — a great revenue stream for the integrator as they are then able to add said device to their portfolio of supported devices. Attaching a device and communicating with it is just the start. At a lower level, plates move through the system via scheduling software, which often only keeps track of where plates are going. Knowing what is in your plate, at the well-level for reagents, samples and volumes, is up to you! Connecting to your LIMS, parsing logs, updating inventory, and linking the experimental design through to experimental results still requires a load of custom patchwork. And then everything breaks if you want to perform some different science — bringing us back to the adaptability problem. The dream is an automation platform with a software and hardware stack that can adapt to changing business and scientific needs.
What’s been tried?
Taking inspiration from Owl Posting’s automation heuristics piece, various strategies have tried (interpretation ours):
The translation layer: By turning protocols into machine-actionable instructions, we can reduce friction between scientific desires and data generation (e.g., Synthace, Briefly Bio, Tetsuwan Scientific).
The hardware layer: By building intermediary proprietary hardware that connects 3rd party devices, we can achieve greater integration between established solutions (e.g., Automata, Ginkgo Bioworks).
The intelligence layer: By improving a system’s perception of its own state (a world model?), we can increase capacity and reduce failure across automated build outs (e.g., Medra, Zeon Systems).
Two groups who were conspicuously absent from the ecosystem worldmap include the key “integrators” HighRes Biosolutions and Biosero, who do all three layers above. A family-owned business and now operating for over 20 years, HighRes have performed over 500 automation installations and partnered with all of the top 20 pharma companies, meaning that whilst startups may have the “mind share”, the people actually delivering appear as unknown entities. However, they’re family owned and most of their big pharma relationships have been long established — they don’t need to advertise. Having seen early demonstrations for how AI is transforming their business (specifically: upload simple protocol into an LLM, get automated implementation with some inventory tracking), much of the automation world feels like it will soon become a solved problem.
What could work?
Owl Posting’s essay then ends predicting cloud labs as the ultimate destination for “where the field is going”; agreed, but is it a venture backable startup problem with huge upside?
On the large-scale fee-for-service build out, everyone wants “Anduril-for-automation” (i.e., ground up, new hardware) whilst wildly gesturing at Flagship’s Lila healthy fundraise (~$200M) to predict this will be the new model.
This triggered some thinking.
Aside: the challenge of the fee-for-service cell biology cloud lab
The case for cloud labs is a nuanced one, depending on what data you want to generate. Historically, the early automation use cases were all related to chemistry, screening small molecule libraries, a very low margin business that was difficult to scale. But the recent focus is all biology (advanced imaging and ‘omic readouts) to power the training of large foundation models. Whilst some applications have been extraordinarily effective (optimization of growth conditions for biomanufacturing), the use cases that move the needle for drug discovery groups are incredibly artisanal: complex model systems that are highly heterogeneous (e.g., organoids, organ-on-a-chip).
Could this be built, a cloud lab delivering such use cases would be hugely enabling for early stage TechBio companies: log onto a website, upload a protocol for interpretation by an LLM, pay for reagents, and receive an email when your data is ready for download. It would allow companies to rapidly build proprietary models and generate target IP, and thus speedrun to the next round of venture financing.
For the avoidance of doubt, the following relates to the challenges in operating a model that is: 1.) fee-for-service, and 2.) focused on cell biology. Essentially, there is a tension: current installations are designed to meet specific use cases, but with multiple customers, one requires adaptability that is hard to deliver on a fee-for-service basis. Therefore, on one hand, someone needs to fund both a build out and product R&D, potentially requiring decades (organoids take months to grow!), and on the other hand, the platform needs to resist attempts to become captured by special interest groups, aka remaining “apolitical”.
The hurdles:
Short-term price vs. long-term value: The alternative to automation is human labour. Whilst automation-generated data can be assumed to be of higher quality with greater quality control. If a manual CRO will charge reagents plus salary against some multiplier, there is some upper limit for how much one can charge for automated science.
However with regards to long-term value, suppose the artisanal craft of “how to do science” is learnable at some abstract level encodable into a foundation model;13 then who gets to retain that information and how? Imagine you let the customer directly utilise your proprietary automated-science foundation model: it will be expensive to operate and the benefits will be diffuse and not immediately tangible.14 In contrast, if the cloud lab keeps all of the experimental learnings to commercially exploit them later, there is an instant conflict with their customers. How does the customer know you won’t compete with them?15 This cannot be accounted for on a fee-for-service basis — remember, partnerships with FTO restrictions will be slower to negotiate and will limit future customer acquisition. So the cloud lab needs to not “have a horse in the race” so to say – but this is expensive.
The disincentive to externalise: Losing a data generation capability has a number of knock on effects: one loses intuition, which then means missing out on opportunities for more interesting information-rich experimental designs — something the AI scientist narrative has yet to deliver, but may do in years to come.
Moreover, many CROs have preferred customers: pharma companies are more important sources of revenue when compared to startups, and whilst not explicitly verbalised, capacity is often directed towards the customer who is guaranteed to still be doing business a decade from now. In the case of Lila, hypothetically Flagship Pioneering (their multibillion dollar venture creator) could feed all of their portfolio’s automation work through Lila.
Put bluntly, are you a real company or an outsourced ML team ripe for an acqua-hire? Perhaps they will turn the data tap off if you disagree to their terms!
Solution space:
To subvert Ronald Reagan, maybe government is the solution. Accepting that we will need to perform product development and we want to “build in the open” so customers can benefit from its learnings, the next-gen cloud lab should be capitalised via a governmental organisation or an “apolitical group”, e.g., NVIDIA, Microsoft, Amazon etc.16 No one worries that AWS is serving your competition, but startup executives should rightfully be concerned if key platform data generation capability all comes from a specific VC-backed cloud lab — especially if that group is much better capitalised than you (this is actually bizarrely similar to what happened in the early days of Ginkgo Bioworks where they invested in their customers and were subsequently the target of shortsellors).
The other question is on what needs to be built: bluntly, it is not clear where a large unfilled technical gap is beyond specific use cases, e.g., some complex coculture capability. There’s various small hardware and scheduling optimizations that can be made, but there’s providers (like HighRes) who can already integrate lots of diverse systems. Moreover, with so many coding agents, as with biomolecular design companies, the open source universe will eventually fill this niche and the margins will collapse. Anything relating to automating protocols will be the remit of OpenAI, Anthropic, DeepMind etc, and almost all solutions will be a wrapper around their agents. What does need to be built is the open source infrastructure to rapidly redeploy labs for new use cases. With this, one may have a foundation for actual cloud labs that could scale as cloud compute does.
Without being specific on actual schematics, two thoughts worth contemplating on: modularity and franchise. The development within automation that’s most exciting is the transition from rails (where there is a fixed route through the system) to autonomous mobile robots (AMRs) to move between static systems, which can then be upgraded and switched out depending on demand.
You only need to get this working once — and we know this is technically possible. Thereafter, you then have a system that can be deployed across multiple jurisdictions around the globe (Boston, West Coast, London, Singapore etc) via a franchise model. Data robustness is trivial in a sense: does Boston’s data agree with London’s data for the same experiment? This replication then can be used to estimate your confidence intervals. Capacity can also be answered by offshoring experimental work.
Unfortunately, this is hugely capital intensive and for the reasons above, this cloud lab network is not easily a VC investable model with a 10✕ return, but perhaps a 2-5✕ return suitable for private-equity – but with a non-trivial risk profile of a de novo startup building an operational capability. However, from an ecosystem perspective such a proposal is hugely appealing — especially capital constrained TechBio groups outside Boston and San Francisco — we unlock a community-led data generation capability that will allow fledgling startups access the next wave of capital.
That’s it for Part I, but hopefully that gives a feel for what’s going on during the discovery phase.
We realise we’ve not really touched on the impact of China. Approximately a quarter of drug candidates under active development originate from China with 46% of new drug molecules entering clinical trials in 25H1 from Chinese companies. Unfortunately, this is going to get pushed to Part II. Subscribe to hear more!
The caveat here is that the patent system is possibly unstable in the long run, see Peter Crane’s commentary here.
Reimbursed by insurance companies; mediated by pharmacy benefit managers — but that’s for another day.
Most effectively used in cancer and infectious disease where simple phenotypes can be measured at high throughput (proliferation rate, cell death etc).
And various conversations with insiders, who shall remain nameless.
Unfortunately, this analysis does not distinguish between on-target vs off-target toxicity.
Historically, pharmacokinetics was a huge driver of attrition but this was on the whole solved by better preclinical data packages.
We are avoiding making the distinction between target identification and validation as it is seldom helpful.
We really need a better name, but this essentially covers all high-dimensional experimental biological data generation pertaining to the genome (genomics), RNA (transcriptomics), proteins (proteomics) etc.
Humorous point: industry commentators often refer to how a positive clinical trial validates the technology… these things are only weakly correlated!
Not to mention the plethora of rare bone disorders.
There’s also the darker version: identify competition early, purchase them along with the IP, discontinue and dissolve!
This is compounded by the fact that early patents are very broad with later patents much more specialised.
Whilst a data leak or two could seriously tarnish the whole enterprise, imagine perfect data security and solutions like federated learning are being employed.
Pharma/biotech are often unduly painted as unsophisticated, this is rarely the case. Using some automated science platform, many of the results will in fact recover known relationships (but may not be in the public domain). Relevant earlier blog post here.
Offhand, I’ve heard many companies reticent to give Anthropic their most valuable data due to their recent announcement to work in drug discovery.
Note the latest AWS-Ginkgo Bioworks press release, but this does not appear to answer the cell biology problem.








