Within Bio Uplift

Why Good AI Advice Still Fails in the Lab

Wet-lab studies suggest that strong AI explanations do not yet reliably overcome contamination, judgement and hands-on execution problems.

37 sources 3 graphics
Preview for Why Good AI Advice Still Fails in the Lab

On this page

  • The tacit skills missing from written instructions
  • What early wet lab uplift studies found
  • Which practical failures remain hardest to solve

Introduction

A recurring question in debates about whether AI could help non-experts build a pandemic pathogen is whether increasingly capable AI advice is enough to bridge the gap between reading about biology and successfully doing biology. The evidence so far suggests that it is not. Although frontier AI systems have become dramatically better at explaining laboratory concepts, generating protocols and troubleshooting scientific questions, real laboratory work still depends on physical judgement, manual dexterity, repeated trial and error, and recognising subtle problems that are difficult to describe in text alone. Recent studies show that AI can substantially improve scientific reasoning and provide useful laboratory guidance, but they also reinforce an important distinction: improving knowledge is not the same as reliably improving hands-on execution. That remaining “wet-lab gap” is one reason many researchers argue that current AI has lowered some barriers to biological work without eliminating the practical constraints that still limit inexperienced users.[arxiv.org]arxiv.orgarXiv LLM Novice Uplift on Dual-Use, In Silico Biology TasksLLM Novice Uplift on Dual-Use, In Silico Biology TasksFebruary 26, 2026…Published: February 26, 2026

Wet Lab Gap illustration 1

The tacit skills that written instructions cannot teach

The hardest parts of laboratory biology are often not the formal procedures written into protocols but the practical judgements that experienced researchers acquire through months or years at the bench.

Scientists commonly describe these as tacit skills: knowledge that is learned through repeated practice rather than explicit instruction. A protocol may explain what should happen, but it cannot always prepare someone for what actually happens when equipment behaves unexpectedly, reagents vary in quality or experimental results become ambiguous.

Examples include recognising when:

  • contamination has occurred before obvious laboratory tests confirm it;
  • pipetting or mixing errors have subtly affected a result;
  • cells or cultures look abnormal despite appearing technically viable;
  • an instrument is drifting out of calibration;
  • an apparently successful experiment is producing misleading data.

Many of these decisions depend on visual pattern recognition, physical handling, accumulated experience and continual comparison with previous experiments. They are difficult to capture fully in written guidance because experts often perform them automatically without consciously articulating every judgement they make. Even excellent AI explanations therefore cannot instantly substitute for practical experience acquired over hundreds of laboratory repetitions.[aisi.gov.uk]aisi.gov.ukLong-Form Tasks | AISI WorkLong-Form Tasks | AISI Work

This distinction helps explain why laboratory training typically involves supervision, observation and repeated practice rather than simply reading protocols. Success depends as much on recognising and correcting mistakes as on following instructions correctly the first time.

47:52

What early wet-lab uplift studies actually found

For many years, discussions about AI biosecurity relied largely on digital benchmarks: could a model answer biology questions, explain molecular techniques or solve scientific reasoning problems?

Researchers increasingly regard this as insufficient because the relevant policy question is whether AI helps people complete real laboratory work.

One early attempt to measure this directly was a pilot study led by researchers at Los Alamos National Laboratory. Participants with no previous wet-lab experience attempted a genuine laboratory workflow while either receiving AI assistance or relying on conventional internet resources. Rather than measuring exam-style knowledge, the study observed how participants interacted with equipment, responded to mistakes and progressed through experimental stages. The researchers presented the work as a feasibility study for measuring “skills-based uplift”, emphasising that laboratory skill—not simply biological knowledge—often remains the limiting factor for inexperienced users.[arXiv]arxiv.orgarXiv Measuring skill-based uplift from AI in a real biological laboratoryMeasuring skill-based uplift from AI in a real biological laboratoryOctober 29, 2025…Published: October 29, 2025

Separately, the UK AI Security Institute has developed evaluations that move beyond factual questions by testing multimodal laboratory troubleshooting. These assessments ask models to interpret photographs, identify contamination, explain unexpected observations and advise users confronting realistic laboratory problems. Recent frontier models now outperform expert comparison baselines on several of these evaluation tasks, indicating genuine progress in providing useful scientific assistance. However, the Institute also distinguishes between better troubleshooting advice and demonstrated success in carrying out complete laboratory projects. Behavioural studies examining real-world laboratory performance remain an active area of research rather than a solved measurement problem.[aisi.gov.uk]aisi.gov.ukFrontier AI Trends Report by The AI Security Institute (AISIFrontier AI Trends Report by The AI Security Institute (AISI

Taken together, these studies suggest two simultaneous conclusions:

  • AI is becoming increasingly helpful during laboratory work.
  • Demonstrating improved advice is easier than demonstrating consistently successful experimental execution.

That distinction is central to understanding current biosecurity risk.

Which practical failures remain hardest to solve

Even when AI provides technically accurate advice, several categories of laboratory failure remain stubbornly difficult for inexperienced researchers.

Wet Lab Gap illustration 2

Detecting problems before they become obvious

Laboratory failures rarely announce themselves clearly. Small mistakes made early in an experiment can invalidate results hours or days later.

Experienced researchers often recognise subtle warning signs from appearance, smell, instrument behaviour or previous experience long before formal quality-control checks detect a problem. An AI assistant can suggest possible explanations, but it cannot directly perceive every aspect of a changing physical experiment unless specialised sensors, imaging systems and feedback loops are available.

This matters because successful biology often depends less on performing a protocol perfectly than on recovering intelligently when something unexpected happens.

Translating instructions into precise physical actions

Many laboratory procedures require fine motor skills and consistent technique.

Two people following identical written instructions may produce different outcomes because of differences in timing, pressure, handling or coordination that protocols do not specify in complete detail.

Small variations can accumulate across multiple experimental steps, especially in longer workflows where early errors affect everything that follows.

Current language models can explain these procedures well, but explanation alone does not guarantee reproducible execution.

Knowing when results should not be trusted

Experienced scientists spend considerable effort deciding whether an experiment has genuinely worked.

Unexpected results may reflect contamination, equipment malfunction, statistical noise or unnoticed procedural mistakes rather than genuine scientific discovery.

This judgement requires combining multiple sources of evidence, laboratory history and domain-specific intuition. AI can propose hypotheses, but deciding whether a physical result deserves confidence still relies heavily on experimental validation and quality control.

Better AI advice does not automatically remove laboratory bottlenecks

One of the more surprising findings from recent research is that AI assistance often improves knowledge faster than it improves overall task completion.

Human uplift studies show that novices perform substantially better on complex biology reasoning tasks with frontier language models than with internet search alone. However, researchers also found that users frequently failed to extract the models’ full capabilities, and that standalone models sometimes outperformed AI-assisted humans because people did not ask the most effective follow-up questions or interpret answers optimally.[arXiv]arxiv.orgarXiv LLM Novice Uplift on Dual-Use, In Silico Biology TasksLLM Novice Uplift on Dual-Use, In Silico Biology TasksFebruary 26, 2026…Published: February 26, 2026

This illustrates a broader point about laboratory work: even if AI continues becoming more capable, successful use still depends on the user’s ability to recognise when advice applies, identify mistakes, ask productive questions and integrate recommendations into an evolving experiment.

Improving the assistant therefore does not automatically eliminate human implementation challenges.

Wet Lab Gap illustration 3

What this means for AI doom arguments

Within discussions of AI doom and catastrophic misuse, the wet-lab gap represents an important source of uncertainty.

Those concerned about existential biological misuse argue that progressively stronger AI systems may continue lowering knowledge barriers, making sophisticated scientific advice available to many more people than before. Improvements in protocol generation, multimodal troubleshooting and scientific reasoning all point in that direction.[aisi.gov.uk]aisi.gov.ukFrontier AI Trends Report by The AI Security Institute (AISIFrontier AI Trends Report by The AI Security Institute (AISI

At the same time, current evidence does not support the simpler claim that accurate laboratory advice alone enables inexperienced individuals to perform complex biological work reliably. Physical execution, iterative troubleshooting, quality assurance and tacit laboratory judgement remain meaningful constraints that have not disappeared.

This distinction matters because estimates of future misuse risk depend not only on how much AI knows, but on how much practical capability that knowledge transfers into real-world laboratory settings. If future systems become integrated with increasingly sophisticated laboratory automation, robotics or continuous experimental feedback, today’s implementation barriers could shrink. Conversely, if tacit human expertise remains difficult to replace, improvements in digital advice may translate into more modest increases in real-world biological capability than benchmark scores alone would suggest. Current evidence is consistent with both possibilities, making continued empirical wet-lab uplift studies an important part of assessing long-term biosecurity risks.[openai.com]OpenAIaccelerating biological research in the wet labMeasuring AI’s capability to accelerate biological research in the wet lab | OpenAIDecember 16, 2025…Published: December 16, 2025

Amazon book picks

Further Reading

Books and field guides related to Why Good AI Advice Still Fails in the Lab. Use these as the next step if you want deeper reading beyond the article.

BookCover for The Genesis Machine

The Genesis Machine

By Amy Webb, Andrew Hessel

A New Yorker Best Book of the Year The next frontier in technology is inside our own bodies. The breakthrough science of synthetic biolog...

BookCover for The Craft of Research

The Craft of Research

By Wayne C. Booth, Gregory G. Colomb et al.

Rating: 4.3/5 from 6 Google Books ratings

Since 1995, students, researchers, and professionals have turned to The Craft of Research for clear and helpful guidance on how to conduc...

eBay marketplace picks

Marketplace Samples

Live-tested eBay searches with available results related to this page.

UsingUSA

Selected fromDNA wall art oneBay.co.uk.

Endnotes

1. Source: arxiv.org
Title: arXiv LLM Novice Uplift on Dual-Use, In Silico Biology Tasks
Link:https://arxiv.org/abs/2602.23329

Source snippet

LLM Novice Uplift on Dual-Use, In Silico Biology TasksFebruary 26, 2026...

Published: February 26, 2026

2. Source: aisi.gov.uk
Title: Long-Form Tasks | AISI Work
Link:https://www.aisi.gov.uk/blog/long-form-tasks

3. Source: aisi.gov.uk
Title: Frontier AI Trends Report by The AI Security Institute (AISI)
Link:https://www.aisi.gov.uk/frontier-ai-trends-report

4. Source: arxiv.org
Title: arXiv Measuring skill-based uplift from AI in a real biological laboratory
Link:https://arxiv.org/abs/2512.10960

Source snippet

Measuring skill-based uplift from AI in a real biological laboratoryOctober 29, 2025...

Published: October 29, 2025

5. Source: OpenAI
Title: accelerating biological research in the wet lab
Link:https://openai.com/index/accelerating-biological-research-in-the-wet-lab/

Source snippet

Measuring AI’s capability to accelerate biological research in the wet lab | OpenAIDecember 16, 2025...

Published: December 16, 2025

6. Source: OpenAI
Title: accelerating biological research in the wet lab
Link:https://openai.com/sr-RS/index/accelerating-biological-research-in-the-wet-lab/

7. Source: GOV.UK
Title: www.gov.uk A I Safety Institute approach to evaluations
Link:https://www.gov.uk/government/publications/ai-safety-institute-approach-to-evaluations/ai-safety-institute-approach-to-evaluations

8. Source: GOV.UK
Title: www.gov.uk Introducing the AI Safety Institute
Link:https://www.gov.uk/government/publications/ai-safety-institute-overview/introducing-the-ai-safety-institute?gh_src=v3scng1

9. Source: aisi.gov.uk
Link:https://www.aisi.gov.uk/research-agenda

10. Source: aisi.gov.uk
Link:https://www.aisi.gov.uk/about

11. Source: aisi.gov.uk
Link:https://www.aisi.gov.uk/blog/early-lessons-from-evaluating-frontier-ai-systems

12. Source: aisi.gov.uk
Link:https://www.aisi.gov.uk/blog/advanced-ai-evaluations-may-update?trk=public_post_comment-text

13. Source: aisi.gov.uk
Link:https://www.aisi.gov.uk/blog/5-key-findings-from-our-first-frontier-ai-trends-report

Additional References

14. Source: alphaxiv.org
Link:https://www.alphaxiv.org/abs/2606.31763v2

Source snippet

A Self-Evolving Agentic System for [Automated]({{ 'full-research-loop/' | relative_url }}) Generation and Execution of Biological Protocols | alphaXivJuly 2, 2026 — A SELF-EVOLVING AG...

Published: July 2, 2026

15. Source: nti.org
Title: Without defined risk tolerances or clear red lines, evaluation results cannot be
Link:https://www.nti.org/analysis/articles/aixbio-horizon-scan-spring-2026/

Source snippet

AIxBio Horizon Scan: Spring 2026May 13, 2026 — The field has focused heavily on producing evaluation metrics without first establishing r...

Published: May 13, 2026

16. Source: frontiersin.org
Link:https://www.frontiersin.org/journals/microbiology/articles/10.3389/fmicb.2026.1832401/full

Source snippet

Researchers can already design experimental protocols using software, submit them to remote facilities at s...

17. Source: youtube.com
Link:https://www.youtube.com/watch?v=UUELqwathns

Source snippet

AI biosecurity wet lab biological risk protocol Working in the BSL-4 laboratory Robert Koch-Institut...

18. Source: youtube.com
Link:https://www.youtube.com/watch?v=Jw47VMof4ic

Source snippet

Ep 28 - The Real Risks, and Real Promise, of AI in Biotech from the Perspective of the National A...

19. Source: youtube.com
Link:https://www.youtube.com/watch?v=acnmEQip6Jk

Source snippet

Scaling Laws: Why Data Governance Is the Key to AI Biosecurity, with Jassi Pannu and Doni Bloomfield...

20. Source: youtube.com
Link:https://www.youtube.com/watch?v=SvPcxRiFauY

Source snippet

The Silent Lab: The Rise of Autonomous Biology & Machine-Led Risks | Code Blue AI...

21. Source: youtube.com
Title: Bryce Cai
Link:https://www.youtube.com/watch?v=oh4AoiAOwfs

Source snippet

Ep 59 - A New Study Taking Responsible [Innovation]({{ 'false-positives/' | relative_url }}) From Benchmarks to Benchwork...

22. Source: scale.com
Title: llm novice uplift
Link:https://scale.com/blog/llm-novice-uplift

Source snippet

Understanding New Biosecurity Risks Posed by LLMs | Scale AIApril 16, 2026 — UNDERSTANDING NEW BIOSECURITY RISKS POSED BY LLMS By Matthew...

Published: April 16, 2026

23. Source: securebio.org
Link:https://securebio.org/benchmarks/about/