VoicesWillison Red Green TDD for AI Agentssimonwillison.net·high signalXBlueskyLinkedInCopy linkAppend Use red/green TDD to agent prompts. Three words, immediate quality improvement. Published as structured guide at agentic-engineering-patterns.SourceSource pagesimonwillison.net↳ Follow the threadStack layerLaurie Voss, via Simon Willison: the cost of reviewing and operating code is collapsing next, and what's left is finding out what people wantSimon Willison's WeblogPolicy dependency / Threat patternA Fine-Tuned RoBERTa-Large Permission Gate Matches Claude Haiku 4.5 at Deciding What an Agent May ToucharXiv 2609.15422Stack layer / Threat patternUnlearning Methods That Pass TOFU and MUSE Still Leak the Secret on 22-86% of Queries Once the Model Is an AgentarXiv 2609.12808Policy dependency / Stack layerCROSS-CATEGORY: Salesforce, Zendesk and Workable All Shipped Named Agent Portfolios on Sept 14, Each Metered in a Different UnitZendesk newsroom, Salesforce press release and Workable via GlobeNewswire (three independent Sept 14 announcements)Stack layer / Threat patternGemini CLI ships an external-context processor to stop indirect prompt injection through build filesGitHubStack layer / Threat patternSnyk put its agent-skill scanner behind a free web page called Skill InspectorSnyk LabsStack layer / Threat patternA 50-PR code review benchmark puts GPT-5.6 Luna at 28x cheaper than Astra and 74% precision against 96%EntelligencePolicy dependency / Stack layerCodeBLEU Scored 91% for Both RAG Strategies While One of Them Hallucinated APIs 56.4% of the TimearXiv 2609.12464