MASEval: Extending Multi-Agent Evaluation from Models to Systems
arXiv 2603.08835·medium signal
Evaluation framework extending LLM benchmarking to full multi-agent systems with per-agent execution tracing, cross-framework compatibility, and system-level metrics that capture emergent coordination failures not attributable to any single model. Addresses the gap between single-model evals (SWE-bench, GAIA) and production multi-agent deployments where failure modes are architectural rather than capability-based.