ResearchFlashAttention-4 1613 TFLOPs Blackwell 71 Percent UtilizationarXiv·high signalXBlueskyLinkedInCopy linkNew attention kernel for Blackwell GPUs. 1.3x over cuDNN, 2.7x over Triton. Written in CuTe-DSL Python. 4x perf at long sequences.SourceSource pagearXiv↳ Follow the threadStack layer / ContrastSplitting Decode by Attention Type Instead of by Operator Buys 31-56% More Tokens Per JoulearXiv 2609.13134Stack layer / Threat pattern787,562 Function Pairs Show AI Code Is Half the Size of Human Code With Different Defect Classes, Not FewerarXiv 2609.12708Stack layer / ContrastDistilled Byte Models Match a Token Model's Accuracy on One-Sixth the Data and Cut Logit Storage to a FiftharXiv 2609.12303Stack layer / ContrastOdin Runs All 32 Llama-3-8B Transformer Layers Under FHE in 366 Seconds on One H100, 4.51x Faster Than THORarXiv 2609.12378Stack layer / Update threadSplitting a CTF Task Into Isolated Sub-Contexts Lets a Local gemma-4 Solve 18.52% of Challenges Standard Agent Loops FailarXiv 2609.12839Stack layer / Update threadMarket Concentration Barely Changes Model Collapse: Pushing One Model to 90% Share Moves Five-Generation Endpoints a Few PercentarXiv 2609.11146Policy dependency / Stack layerRIPPLE: an edit confined to one prompt-policy segment changes downstream behavior, so replay candidate edits after previously accepted ones before persistingarXiv 2609.12127Policy dependency / Follow-up threadMOSAIC Picks a GraphRAG Traversal Policy per Query and Beats the Best Fixed Policy by 9.96 PointsarXiv 2609.11065