Agents
Meituan's ATLAS grades tool-use agents on two horizons because final-outcome scoring hides where the failure started
ATLAS separates a request horizon, trajectory-wise signals tying deficiencies to specific execution locations and capability concerns, from an interaction horizon, user-wise signals checking whether later service stays aligned with earlier context. LLM judge interfaces are calibrated against high-confidence references from real business logs and then distilled into smaller diagnostic models for lower-latency scoring. Deployed on Meituan Xiaotuan production traffic, offline experiments assess signal fidelity and replay-based policy improvement while online A/B tests show concurrent gains in user engagement, downstream business outcomes and sampled human-audit quality.
Source
↳ Follow the thread