01

How the agent coalition works

Google DeepMind describes Co-Scientist as a system of specialized Gemini-based agents coordinated by a supervisor. Roles cover hypothesis generation, clustering related ideas, reflection resembling peer review, pairwise ranking, evolution of leading proposals and a final meta-review. Ranking uses an ELO-style tournament, so the system need not accept the first fluent answer it produces. Most computation is intended for criticism, verification and selection rather than simply generating a larger volume of text.

The system connects literature and web search with databases such as ChEMBL and UniProt and selected specialist tools, including experimental use of AlphaFold. Its intended output is not an automatically accepted discovery but a structured research proposal with arguments, references and an experimental direction. Google’s source describes the architecture and case studies. The conclusion that the idea tournament may reduce premature commitment to one hypothesis is Mateusz’s editorial interpretation, not a separately measured outcome reported by DeepMind.

02

What the laboratory tests show

Google and partner laboratories describe hypotheses involving liver fibrosis, ALS, cellular ageing and infectious disease. In one liver-fibrosis experiment, a selected candidate blocked 91% of the measured scarring-linked response. That external contact with a physical experiment is stronger evidence than a textual score alone. It remains, however, one reported case with one specific laboratory readout rather than a general measure of Co-Scientist’s effectiveness across research domains.

DeepMind also reports that more than one hundred institutions have been involved in development and evaluation. That breadth can diversify the questions tested, but partner count does not replace accuracy data for all generated hypotheses. A fuller assessment needs denominators: how many proposals were rejected, entered a laboratory, produced a positive result and survived replication. The current evidence supports the feasibility of the workflow and several promising cases, not a universal discovery rate.

Data view

A coalition of six roles plus a supervisor

A hypothesis passes through three phases and returns to a person as a proposal to evaluate.

  1. 01
    Generation

    Proposes directions and hypotheses.

  2. 02
    Proximity

    Clusters ideas and preserves diversity.

  3. 03
    Reflection

    Acts as a virtual peer reviewer.

  4. 04
    Ranking

    Runs a pairwise idea tournament.

  5. 05
    Evolution

    Combines and develops leading proposals.

  6. 06
    Meta-review

    Synthesises debate into a research proposal.

  7. 07
    Supervisor

    Plans adaptively and coordinates parallel work.

People choose the problem, evidential standard and whether an idea moves into the laboratory.Source: Google DeepMind · Co-Scientist
03

What the system cannot decide

Co-Scientist can compare more directions than one researcher, but it does not independently select the socially appropriate goal, acceptable risk or required standard of evidence. An ELO ranking gives a relative order within the pool of generated hypotheses; it does not prove that the winner is true or genuinely novel. The system may also inherit omissions in the literature, database errors and biases in its underlying models. A defensible output therefore needs visible sources, contradictions and uncertainty.

A positive result in cells or another laboratory model does not imply an effective treatment for people. Replication, mechanism studies, safety, dosing and multiple regulatory stages remain between a hypothesis and clinical use. The system also cannot replace the experimental expertise needed to notice a flawed protocol. The evidence supports describing Co-Scientist as a partner for generating and organizing hypotheses. It does not justify claiming that AI autonomously discovers therapies ready for use.

04

Mateusz’s proposed decision process

Mateusz’s framework moves a hypothesis through five gates: support in the sources, distinction between novelty and repetition, biological plausibility, risk review and a pre-specified discriminating experiment. An agent can prepare evidence for every gate, but a qualified team decides whether the proposal advances. This is not a formal protocol published by Google. It is Mateusz’s model for preserving human accountability while benefiting from the system’s broad exploration of possible directions.

The next useful evidence includes prospective comparisons with human teams, the complete number of generated candidates and the share of hypotheses surviving independent replication. Time and cost from question to decisive experiment matter more than the polish of the final report. Genuine progress will be visible when the system repeatedly proposes directions that are sound, non-obvious and testable, while unsuccessful experiments are reported as transparently as the selected successes.

05

From a Gemini 2.0 prototype to a partner developed with scientists

Google Research introduced the first AI Co-Scientist in February 2025 as a multi-agent system built on Gemini 2.0. Its ambition went beyond summarizing literature: from a scientist’s objective, the system would create hypotheses, supporting rationale and experimental proposals, while allowing a person to add constraints and steer later iterations. In 2026, the architecture and validations appeared in a Nature paper and DeepMind presented further laboratory collaborations. This sequence matters because it separates the initial prototype and its benchmark evidence from newer use cases. A later case study should not be silently treated as evidence available at the original launch.

DeepMind reports collaboration with researchers from more than one hundred institutions and gradual access for individual researchers through Gemini for Science and for Google Cloud partners. Institution count measures the breadth of development and evaluation, not the number of independently confirmed discoveries. Co-Scientist is best described as an experimental collaboration platform whose evidence is accumulating in layers. Product access, publication of an architecture and a positive wet-lab result are three different events, each answering a different question about readiness. Combining them into one headline obscures what has actually been established.

Data view

One reported laboratory result

In a liver-fibrosis study, a highlighted candidate blocked part of a scarring-linked response.

Scarring-linked response blocked in lab test91%
Institutions involved in development>100
This is one case and one laboratory readout, not overall Co-Scientist effectiveness or a clinical outcome.Figures reported by the source author: Google DeepMind
06

A supervisor turns the research goal into a specialist queue

A scientist begins with a natural-language objective and may provide seed ideas, preferences and constraints. A Supervisor agent turns this material into a research-plan configuration, assigns jobs to an asynchronous queue and allocates resources among Generation, Reflection, Ranking, Evolution, Proximity and Meta-review. Generation proposes candidates, Proximity clusters related ideas, Reflection searches for weaknesses, Ranking compares hypotheses pairwise, Evolution combines and extends leading directions, and Meta-review prepares a synthesis. These roles describe functions in a workflow; they are not independent experts with distinct training, institutional incentives or disciplinary backgrounds.

The asynchronous queue lets the system scale computation without maintaining one monolithic conversation. One agent’s output can become material for criticism or evolution in later rounds, and the scientist can add knowledge or correct the goal. The architecture resembles an organized hypothesis workshop, but its roles still rely on related foundation models and may share the same blind spots. A plurality of system voices is not equivalent to independent methods, laboratories or schools of thought. Diversity must also enter through source selection, datasets, experimental techniques and reviewers who are outside the generation loop.

07

Test-time compute and Elo aid selection but do not prove truth

Co-Scientist spends additional computation after receiving a question: agents debate, judge pairs of hypotheses and evolve higher-ranked proposals. An Elo-style rating provides feedback that guides later rounds towards promising directions. In Google’s original report, higher Elo scores correlated with greater accuracy on GPQA questions, and automated quality ratings rose with additional compute. That supports the ranking mechanism as a useful selection signal. Elo remains an internal auto-evaluation, however, not a measurement of nature. A system can become increasingly confident in the best option available to it while every candidate in its pool is incomplete.

The study used fifteen expert-curated open research goals, while human assessment of novelty and impact covered a smaller subset of eleven. The authors report that experts preferred Co-Scientist outputs over comparison systems, but the sample is small and evaluators were judging the potential of written hypotheses before a complete experimental programme. More computation may improve reasoning and remove obvious contradictions; it cannot guarantee that the candidate set contains a true answer. A tournament ranks what the system generated. It does not create an independent ground truth or eliminate shared errors in the judge and contestants.

08

An evidence ladder: benchmark, rediscovery and wet laboratory

Evidence for the system comes at several levels. Benchmark questions measure reasoning where an answer is already known. Expert review asks whether a proposal appears novel, important and testable. In work on antimicrobial-resistance gene transfer, the system recovered a mechanism that the research group already knew from results not yet public. That is a strong hidden-answer test, but it is still rediscovery rather than a direction first found by AI. The strongest next step is a prospectively recorded hypothesis predicting an outcome that the model could not retrieve from public literature, followed by an experiment capable of distinguishing it from credible alternatives.

The original study described three biomedical settings — drug repurposing, new-target discovery and mechanisms of antimicrobial resistance — all with experts in the loop. Newer material includes a liver-fibrosis candidate that blocked 91% of one measured scarring-linked response in a laboratory assay. That is meaningful evidence about one hypothesis and one readout; it is not “91% system accuracy” and it is not a clinical outcome. Moving from a computational proposal to cells, organisms and patients introduces a new evidential burden at each stage, including replication, mechanism, safety, dosing and regulation.

09

Grounding in sources introduces its own failure modes

The system can use web search, ChEMBL, UniProt and selected specialist models, including AlphaFold in some collaborations. Citations and databases help establish contact with previous results and test whether a proposal is an obvious repetition. They do not guarantee completeness. Scientific literature contains conflicting experiments, errors, publication bias against negative results and delays between discovery and indexing. A chemical or protein database may be current at a different date from a paper, while biological names and identifiers can be ambiguous. Better retrieval reduces one source of error without turning retrieved material into ground truth.

A defensible report should therefore expose the exact source for each material claim, database version, search date, strongest counterargument and important evidence that could not be found. Novelty requires more than failing to locate the same sentence: synonyms, adjacent mechanisms, patents and not-yet-indexed work must be considered. Grounding is an audit process, not a row of decorative links. An agent can greatly expand the first scan, but a domain specialist still decides whether a change in language represents a new mechanism or merely a restatement of an established idea.

10

Safety and research purpose cannot be delegated to a ranking

The Nature paper describes default criteria including alignment with the stated goal, plausibility, novelty, testability and safety. DeepMind also reports internal and external misuse evaluations across chemical, biological, radiological and nuclear domains and the development of custom classifiers intended to flag unethical goals. This is an important component of the system, but the announcement does not provide a complete quantitative measure of residual risk for every new model, database and tool combination. A classifier can reduce exposure; it does not replace an institution’s research governance or create permission to perform a procedure.

Even without malicious intent, a system may recommend an experiment with unacceptable biological risk, improper use of patient data or an undisclosed conflict of interest. A hypothesis ranking does not automatically account for ethics approval, laboratory competence, animal welfare, clinical requirements or social cost. People must retain decisions about the objective, data access, material acquisition and execution. The most productive proposal is not necessarily permissible, and a safe refusal should be part of normal operation rather than an exception added after an incident. Scientific judgment includes choosing what ought not be optimized.

11

Integrating Co-Scientist into an auditable process

A team should first record its question, prior knowledge, novelty criterion, constraints and experimental budget. Co-Scientist may generate a broad pool, but rejected candidates and ranking reasons should also be retained. A specialist then checks sources, conflicts and feasibility, and the discriminating experiment is specified before its result is known. Where possible, hypothesis assessment should be blinded to authorship and compared against proposals from people or a simpler system. This converts a selected success story into a measurable test and helps reveal whether the multi-agent process adds value beyond more computation or more polished prose.

The decisive metrics are time and cost to a discriminating experiment, candidates required per success, negative-result rate, replication and eventual mechanistic quality. Publishing only the strongest stories makes process value impossible to calculate. A critical account should also separate the contributions of the model, databases, scientist, laboratory team and inherited literature. Co-Scientist may be a valuable amplifier of search breadth; maturity will be demonstrated by recurring prospective results that independent teams can reproduce, not by the fluency of its final research proposal.

12

Credibility begins before the result

With networks of research agents, the integrity risk begins before interpretation, when the system can reframe the question, alter success criteria, or retain only paths supporting an appealing conclusion. Before the experiment starts, the team should preregister the hypothesis, primary outcome, exclusion rules, analysis plan, and stopping condition. The record needs a timestamp, version, and accountable owner; later amendments should remain visible with their rationale. This does not prohibit exploration. It separates analysis planned in advance from discoveries made after inspecting evidence, so readers can distinguish a test from a new hypothesis.

Negative, inconclusive, and expectation-defying outcomes deserve the same discipline. An agent should not erase failed attempts because they weaken a polished story. The report can explain which approaches were rejected, what limitations appeared, and whether the outcome suggests no effect, inadequate design, or insufficient evidence. Preserving the full path reduces selective reporting and helps others avoid the same dead ends. When privacy, security, or licensing prevents disclosure, the project should still acknowledge the result and describe what was withheld. Silence must not masquerade as an absence of contrary evidence.

A practical standard can connect a public plan registry, an execution log, and a final deviation report. An independent reviewer or a separate control agent can then check whether the reported outcome still answers the declared question and whether omitted experiments have been accounted for. Prompt versions, tool settings, analysis code, and human decisions should also be retained, not to manufacture procedural theatre but to make the reasoning traceable. The resulting record lets others assess whether the conclusion survives a change of operator or model. A trustworthy research system does not pretend that every run succeeds. It shows how failure narrowed the space of plausible answers and why the final claim remains proportionate to the evidence.

Questions and answers

Frequently asked questions

Does a multi-agent debate prove that the winning hypothesis is true?

No. The Elo tournament ranks proposals inside the system-generated pool. It may improve selection, but a well-designed experiment and replication determine whether a biological claim survives.

What does the reported 91% liver-fibrosis result mean?

It concerns one selected candidate blocking a specific scarring-linked response in a laboratory assay. It is not overall Co-Scientist accuracy and it is not an outcome in patients.

Have the system and its results been peer reviewed?

The architecture and foundational validations were published in Nature, while newer cases have different levels of maturity. Each hypothesis and experiment needs to be judged according to its own status.

Can Co-Scientist operate without a domain expert?

It should not independently own high-risk research decisions. Experts define the goal, audit sources and feasibility, approve experiments, and interpret and replicate the results.

Primary sources

Check the evidence

  1. Google DeepMind — Co-Scientist
  2. Google Research · AI Co-Scientist report
  3. Nature · AI Co-Scientist research paper