RealBench: Repo-Level Code Generation Benchmark Aligned with Real-World Software Development Practices
arXiv·medium signal
New benchmark evaluates LLMs on code generation tasks that match actual software development workflows, going beyond isolated function completion. RealBench requires LLMs to work with full repository context, addressing the gap between synthetic benchmarks like HumanEval and the complexity developers face in production codebases. Tests show significant performance drops compared to simpler benchmarks.