Fetching from the wire…
Public story · 2026-08-25 · high
Bartholomew scored 0.175 versus 0.127 for GPT-1900, on a benchmark Unbounded Labs had to invent.
Why now: Unbounded Labs published the build log with every cost line and score listed, so outside researchers can check the arithmetic themselves.
Unbounded Labs trained a 2.82 billion-parameter language model from scratch on 20.1 billion tokens of English text written before 1931, for $757 total.
That's cheap enough for a solo researcher to rerun the whole experiment without a lab budget, and the build log makes the entire process checkable line by line.
The premise came from a Hassabis suggestion. Train a model that never sees anything published after 1930, then check whether it reaches the conclusions that era's scientists reached on their own.
The training corpus started as Harvard's Institutional Books 1.0, 242 billion tokens. Unbounded Labs filtered that down to 25.7 billion tokens with a 0.90 OCR threshold and a pass to strip anachronisms. They trained the 32-layer decoder-only model on 20.1 billion of those tokens, using rotary embeddings and Flash Attention 3.
No benchmark existed for scoring a period-bounded model, so Unbounded Labs built one. Bartholomew scored 0.175 centered accuracy on it, against 0.127 for a comparison model called GPT-1900, despite having fewer parameters and fewer training tokens.
The $757 splits into a $227 GPU lease, $235 in API costs, and $180 on tooling.
Each link below shares sources, entities, or timing with this story.
Claude Code benchmarked against GPT / Shared entity: GPT / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (Claude Code benchmarked against GPT); both cover GPT; overlapping topics (against, model).
GPT competes with Claude / Shared entity: GPT / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (GPT competes with Claude); both cover GPT; overlapping topics (against, cost, model).
GPT competes with Claude / Shared entity: GPT / Shared topic / Earlier coverage
Linked by a graph relationship (GPT competes with Claude); both cover GPT; overlapping topics (accuracy, attention, model, token).
Claude Code benchmarked against GPT / Shared entity: GPT / Shared topic / Earlier coverage
Linked by a graph relationship (Claude Code benchmarked against GPT); both cover GPT; overlapping topics (cost, model, token).
Linked by a graph relationship (Claude Code benchmarked against GPT); both cover GPT; overlapping topics (cost, model, token).
Linked by a graph relationship (Claude Code benchmarked against GPT); both cover GPT; overlapping topics (cost, model).
GPT competes with Claude / Shared entity: GPU / Shared topic / Earlier coverage
Linked by a graph relationship (GPT competes with Claude); both cover GPU; overlapping topics (cost, model, token).
GPT competes with Claude / Shared entity: GPT / Shared topic / Earlier coverage / Tension
Linked by a graph relationship (GPT competes with Claude); both cover GPT; overlapping topics (against, model).