DeskDeepSeek V4 Benchmark LeaksHHumai Blog·high signalXBlueskyLinkedInCopy link90% HumanEval, 80%+ SWE-bench. Engram memory. Targets consumer hardware. Open-weight. Expected Feb 17.↳ Follow the threadStack layer / Update threadLG Ships K-EXAONE 2.0, a 750B-Parameter MoE, Under Apache 2.0 — South Korea's Largest ModelThe Korea TimesStack layer / Update threadDeepSeek-V4-Flash-0731 Ships MIT-Licensed at 304B Params, Beats DeepSeek-V4-Pro, and Costs $0.14 per Million Input TokensDeepSeek (HuggingFace model card), corroborated by Simon Willison and Latent Space AINewsStack layer / Follow-up threadDeepSeek Ships V4-Flash-0731 With Native Responses API Support and Codex Adaptation, Says V4-Pro 'Will Follow Soon'DeepSeek API Docs (surfaced via r/LocalLLaMA, 752 upvotes / 283 comments; 445 points / 227 comments on Hacker News)Stack layer / ContrastInfoOps Bench: Model Refusal on State-Backed Influence Ops Ranges From 8.8% to 94.5%, Unexplained by Model SizearXivThreat pattern / Update threadWillison on Oxide and Friends: the Open-Weight Revolution Episode Was Already Obsolete on PublicationOxide and Friends (via Simon Willison)Stack layer / Update threadcodex-vision-proxy Takes 181 Stars in a Day Giving Text-Only Models Working Access to Codex's view_image ToolGitHubStack layer / Follow-up threadMicrosoft Research's EvoLib Turns Inference-Time Experience Into Reusable Skills Without Ground-Truth LabelsMicrosoft ResearchStack layer / Update threadSigma-Mem Gives Multi-Agent Systems a Reliability Memory That Tracks Which Peers to Trust and WhenarXiv