Google DeepMind has expanded Co-Scientist from a hypothesis generator into a lab-integrated research system, according to The Decoder, which attributes the claims to Google and the researchers behind the work. The system is described as a Gemini-based multi-agent setup that can now plan experiments, write code, control lab equipment, analyze results, and generate scientific manuscripts. The reported technical shift is the move toward a closed-loop research workflow. The Decoder says Co-Scientist starts from a research question, derives hypotheses, produces experimental plans or machine-readable lab protocols, executes or supports the work, checks results, and drafts manuscripts. Verification modules are also said to compare numerical claims in generated text against execution logs from generated code, an effort aimed at reducing fabricated results. This is an expansion of a system Google first introduced in February 2025, when Co-Scientist was based on Gemini 2.0 and still had shortcomings in fact-checking and literature review, according to the report. The new version was tested across three fields with different levels of autonomy: materials science, biology, and computer science. In materials synthesis, Co-Scientist was paired with a semi-automated high-temperature furnace. The Decoder reports that the system identified a safer pathway for a sought-after two-dimensional material previously produced mainly through hazardous etching, then generated growth recipes adapted to the lab’s equipment. After 25 rounds with human refinement, the team produced layered structures whose properties resembled the target material, though the report notes that definitive confirmation of the atomic structure is still pending. A second materials experiment went further on direct equipment control. The Decoder says three semiconductor thin films were synthesized on the first try, with Co-Scientist using Gemini 3 Deep Think to control equipment directly. The reported benefit was a reduction in recipe-development time from days to minutes, but the limits are important: humans still had to load samples and precursor materials manually, the faster mode produced smaller and less uniform crystals than carefully optimized recipes would, and lead author Samuel Schmidgall said transferability to other labs remains open. In biology, Co-Scientist reportedly built an image-analysis pipeline to predict which patterns genetically engineered E. coli colonies form at different chemical concentrations. The Decoder says predictions generated with Gemini 3 Pro Image matched unpublished lab results for three of four shape features. The caveat is that the system was reasoning between known conditions, not predicting behavior in entirely new systems. The most autonomous reported test came in computer science. Co-Scientist designed a medical AI architecture called Agent_H without human involvement beyond the initial setup, according to The Decoder. Agent_H classifies incoming queries, generates multiple candidate responses in parallel, and refines them; after correcting for overly long responses, it reportedly outperformed six frontier models on health benchmarks, including GPT-5 and Claude Opus 5. But the same report says the benchmark results did not hold up under human evaluation by three board-certified physicians across nine categories, a significant caution for any claim of medical superiority. For now, this is a developing research story rather than a settled product launch. The single report describes meaningful progress toward AI systems that can participate in experimental science, but it also preserves the most important limitations: human intervention remains necessary in parts of the materials workflow, some results are not fully confirmed, biological generalization is narrow, and benchmark wins in medical AI did not translate cleanly to physician review. Who benefits: Research groups working with semi-automated lab equipment could benefit if systems like Co-Scientist reliably shorten experiment-planning cycles. Teams building scientific software and lab automation may also see stronger demand for machine-readable protocols and execution logging. Who's exposed: Organizations relying on benchmark-only evaluations are exposed by the Agent_H result, where reported benchmark gains did not hold up in physician review. Labs adopting similar systems also face uncertainty around transferability, manual steps, and incomplete experimental confirmation.