Research
Intern-S2-Mobius Splits Memory From Reasoning in the Architecture and Gets ~4x Inference Speedup at Parity
Mobius-v0 restructures the transformer into one globally shared Memory (FFN) holding knowledge vectors plus multiple Reasoners (self-attention) that repeatedly query it, using hidden states as cache and carrier. Trained from scratch, a 7B Mobius matches a 7B transformer baseline on downstream scores using 62.6% of the baseline's training data; continually pretrained from Qwen3.5-35B, Intern-S2-Mobius matches downstream score with nearly 4x end-to-end inference speedup. It is currently the #4 trending paper on HuggingFace Daily Papers with 26 upvotes.
↳ Follow the thread