Google DeepMind's Co-Scientist paper: Gemini agents matched unpublished wet-lab E. coli measurements and beat six frontier models on HealthBench
Samuel Schmidgall and 34 co-authors at DeepMind published arXiv 2608.26701 on August 27, covering a Gemini multi-agent system that runs hypothesis generation, experimentation and manuscript writing end to end. It predicted emergent swarming phenotypes of engineered E. coli across IPTG gradients from sparse imaging data, quantitatively matching unpublished wet-lab morphology, and autonomously discovered an inference-time scaling architecture that outperformed six frontier models on HealthBench Hard and Professional under blinded physician review. A double-blind study with 30 domain experts across 450 reviews is the part worth reading: the reliability modules are what cut hallucination and plagiarism, not the base model.
↳ Follow the thread