Fetching from the wire…
Source-backed findings, relationship evidence, citations, and briefing history from the public MindPattern archive.
Showing the first 40 findings. More graph evidence exists in the corpus.
paddo.dev analysis shows AGENTS.md compliance drops below 68% at 500+ instructions
Source findingpaddo.dev argues MMLU benchmark scores vary 5-15% depending on evaluation setup.
Source findingpaddo.dev argues GSM8K saturated from 50% to 95%+ in two years, making it unreliable.
Source findingpaddo.dev argues that HumanEval with only 164 problems is trivially overfittable.
Source findingLightrun survey of 200 SRE leaders shows 43% of AI-generated code needs production debugging.
Source findingOpen-weight models can match Anthropic's security capabilities at lower cost.
Source findingAnalysis reveals AI productivity should track rework and post-merge defects.
Source findingLightrun survey of 200 SRE leaders shows 43% of AI-generated code needs production debugging.
Source finding