Reddit
OpenEvidence shipped a four-model medical family and claims the first perfect 100% on MedQA
Announced September 3, the family splits by latency budget: Osler at about 5 seconds becomes the default, Sackett at about 30 seconds for evidence-weight questions, and Snow at about 5 minutes replacing Deep Consult. All three are free to verified clinicians. A fourth model, Darwin, is research-preview by application and is claimed to lead MedXpertQA at 72.8%, HealthBench Professional at 82.7% and NOHARM at 87.2%, plus a reported 100% on MedQA, which is a benchmark-saturation claim worth independent replication before anyone cites it.
↳ Follow the thread