← The Wire
Source trail

Unbounded Labs, via r/LocalLLaMA and r/MachineLearning

Public MindPattern findings, entities, and graph evidence that cite this source.

Findings
1
All-time hits
1
High value
0
Last seen
2026-08-25

Related findings

  1. 2026-08-25 / REDDITBartholomew Is a 2.82B Model Trained From Scratch on 20.1B Tokens of Pre-1931 English for $757Unbounded Labs published the full build log on 2026-08-22 for a 32-layer decoder-only model with rotary embeddings and Flash Attention 3, trained on a corpus filtered from Harvard's Institutional Books 1.0 down from 242B tokens to 25.7B using a 0.90 OCR threshold plus anachronism removal. Total cost was about $757 ($227 GPU lease, $235 API, $180 tooling), with projected first-week hosting at $8 to $16. They also had to build the eval, Vintage CORE, since none existed, and Bartholomew scored 0.175 centered accuracy against GPT-1900's 0.127 on the filtered bundle despite fewer parameters and tokens. The premise, from a Hassabis suggestion, is whether a model trained only on pre-1931 text can reach conclusions the scientists of that era reached.
Open latest cited source