ResearchCL4SE Context Learning Benchmark for SE Tasks 24.7% ImprovementarXiv·high signalXBlueskyLinkedInCopy linkFirst standardized eval for context engineering in coding. 13K+ samples, 24.7% avg improvement. Tells which context types matter for which tasks.SourceSource pagearXiv↳ Follow the threadStack layer / Threat patternVibe Coding Cut Task Time 27% and Raised Security Vulnerabilities in the Same TrialarXiv 2609.09560Policy dependency / Stack layerRetrieval that crosses into your dependencies' source, not just your repo, adds up to 6.3% pass@1 and survives version changesarXiv 2609.09987Policy dependency / Threat patternAgents hit 80% F1 deciding whether a dependency CVE is exploitable, but fall below 70% explaining whyarXiv 2609.08040Stack layer / Threat patternAgentAudit attaches to a running agent and scores its trace on ten dimensions, exposing 95.1 vs 22.6 trust spreads at similar task completionarXivStack layer / Threat patternComments help LLM code generation only when they leak correct solution content, and comments from a different problem cut pass@1 by 20.8%arXiv 2609.09242Stack layer / ContrastDetecting benchmark leakage with code-specific features beat perplexity-only methods on eight code benchmarksarXiv 2609.09865Stack layer / ContrastA capability-scoped harness cut prompt-injection execution from 33-47/75 runs to 3/75 without asking the model to spot malicious textarXiv 2609.08371Stack layer / Update threadSkillAdam Borrows Adam's Two Moments to Stop Agent Skill Files From ThrashingarXiv 2609.08944