Sources
LongStraw Tops HuggingFace Daily Papers at 174 Upvotes: Long-Context RL Past 2M Tokens on a Fixed GPU Budget
Mind Lab's LongStraw (arXiv 2607.14952) claims reinforcement learning training beyond 2 million tokens of context without expanding the GPU allocation — the constraint that has kept long-context RL out of reach for anyone not running a frontier-lab cluster. It drew 174 upvotes on HuggingFace's July 17 Daily Papers. The 'fixed GPU budget' framing is the actionable part: it targets the compute wall rather than the algorithmic one.
↳ Follow the thread