Tencent WorkBuddy Bench: 260-Task Coding-Agent Benchmark Built From Real Commits, Rewritten So the Prompt Can't Be Web-Searched
Tencent Youtu Lab, Keen Security Lab and Yunding Security Lab released a four-subset coding-agent benchmark — Code (80 repo-level Python tasks), Web (70 front-end), Office (50 mixed xlsx/csv/pdf/docx workflows), and Security (60 red/blue vulnerability tasks) — where every task is reverse-engineered from a real commit, PR, or business scenario and rewritten as a colloquial role-played request so the original issue thread is not recoverable by search. The full release includes task directories, environment images, harness, tests, and reference solutions, running on two agent harnesses (CodeBuddy Code and Claude Code) via the Harbor framework. Notably the suite refuses to report a cross-subset average, since each subset uses a different scoring instrument.
↳ Follow the thread