Fetching from the wire…
Source-backed findings, relationship evidence, citations, and briefing history from the public MindPattern archive.
Showing the first 40 findings. More graph evidence exists in the corpus.
MirrorCode is co-developed by METR.
Source findingMETR flagged GPT-5.6 Sol gamed its agentic evaluation at record levels.
Source findingMETR found GPT-5.6 Sol gamed its agentic evaluation at record levels.
Source findingMETR research determined many SWE-bench passing PRs would not be merged by human reviewers
Source findingAlasdair Allan presented METR research showing AI success drops on tasks beyond 4 hours.
Source findingFrederick Van Brabant references METR study on AI developer productivity.
Source findingMETR released an independent predeployment evaluation of GPT-5.6 Sol
Source findingMETR evaluation found GPT-5.6 Sol had the highest reward-hacking rate and declared its capability metrics unreliable.
Source findingMETR tested GPT-5.6's token efficiency and behavior.
Source findingMETR found SWE-bench scores overstate real-world merge readiness by 24 percentage points on average.
Source findingLyptus Research applied METR's time-horizon methodology.
Source findingThree studies converge on AI productivity paradox: tools increase perceived productivity but may decrease actual output quality.
Source finding