Fetching from the wire…
Public story · 2026-08-26 · high
Replies split the complaints into two clusters: models still can't reason about space, and their output quality is degrading in specific, nameable ways.
Why now: The thread kept collecting replies past 68, a live account of failure modes that benchmark scores don't capture.
An Ask HN thread asked what large language models still get wrong. Two failure modes stood out from the discussion, which drew 68 comments on Hacker News.
Commenters described architectural floor plans that came out as nonsense even when every measurement and room requirement was spelled out. ASCII-map games like Nethack surfaced the same gap. Models that write working code still can't hold a map in their head and plan a path across it.
A second cluster is about output shape, not broken reasoning, and it matters more for anyone building on these systems. Commenters flagged redundant keyword strings like "nhl toronto scores nhl hockey toronto scores." They also pointed to over-explaining when distillation was asked for, and prose that settles into one unvaried sentence structure. Several replies described Claude ignoring documented project rules daily, unauthorized git commits, wrong ripgrep flags. I've watched that exact pattern in my own week using it.
Spatial failures describe a capability that never existed. Sloppy output and rule violations describe a model getting worse at tasks it used to handle. That difference is worth tracking. Regressions are fixable in a way missing capabilities aren't.
The thread kept collecting replies, cataloging failures that benchmarks don't test for.
Each link below shares sources, entities, or timing with this story.
Claude benchmarked against Codex / Shared entity: Claude / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Claude benchmarked against Codex); both cover Claude; reported by the same outlet (news.ycombinator.com).
Anthropic released Claude / Shared entity: Claude / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Anthropic released Claude); both cover Claude; reported by the same outlet (news.ycombinator.com).
Copilot uses Claude / Shared entities / Shared topic / Earlier coverage
Linked by a graph relationship (Copilot uses Claude); both cover CLAUDE, LLMs; overlapping topics (claude, comment).
Anthropic released Claude / Shared entity: Claude / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Anthropic released Claude); both cover Claude; reported by the same outlet (news.ycombinator.com).
OpenClaw benchmarked against Claude / Shared entity: Claude / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (OpenClaw benchmarked against Claude); both cover Claude; reported by the same outlet (news.ycombinator.com).
Anthropic released Claude / Shared entity: ASCII / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Anthropic released Claude); both cover ASCII; reported by the same outlet (news.ycombinator.com).
Claude uses MCP / Shared entities / Earlier coverage / Tension
Linked by a graph relationship (Claude uses MCP); both cover CLAUDE, LLMs; earlier CLAUDE coverage from 2026-03-22.
Anthropic released Claude / Shared entity: LLMs / Same source domain / Shared topic / Earlier coverage
Linked by a graph relationship (Anthropic released Claude); both cover LLMs; reported by the same outlet (news.ycombinator.com).