Verifier-First Terraform Study: Retrieval Takes Qwen2.5-Coder 7B From 14.0% to 45.7% pass@1, Verifier Feedback to 62.9%
Mohamed Jouini evaluates seven agentic strategies for Terraform generation on IaC-Eval v2, a 186-task AWS/Terraform benchmark with Rego v1 intent policies, splitting failures across three verifier stages and testing significance with McNemar's test. ReAct agents with MCP or ChromaDB-backed RAG lift Qwen2.5-Coder 7B from 14.0% to 45.7% pass@1; iterative refinement on verifier feedback reaches 62.9% for the 7B model and 84.4% for GPT-4o; GEPA reflective instruction optimization adds +7.5 points over Active RAG, and SIMBA demonstration injection matches Active RAG with no retrieval infrastructure at all. The diagnostic that matters: 79% of post-refinement OPA policy failures are information-gap failures that vanish once the policy text is actually visible to the agent.
↳ Follow the thread