Skills
Swapping just 12% of your long-context training mix for dependency-rich cross-repo code lifts retrieval, state tracking, and agentic performance
OctoLong uses AST parsers, language server backends, and package managers to synthesize code contexts with genuine long-distance dependencies — the thing standard long-context corpora lack. Mid-training 600M–14B models on ~50B tokens (including ~6.2B OctoLong tokens and ~10B instruction tuning) and comparing against 18 state-of-the-art open-weight long-context LMs, replacing only 12% of the traditional context-extension corpus produced substantial gains in long-range retrieval, long-term state tracking, repository-level comprehension, and agentic tasks. Short-context API-usage performance improved too, so it isn't a long-context/short-context tradeoff.
↳ Follow the thread