BTS-AgentBench compiles read-only industrial telemetry into deterministic, replayable multi-turn agent episodes
Posted 27 August, BTS-AgentBench addresses the gap that industrial sites hold large volumes of read-only telemetry but no method exists for turning those records into executable agent tasks. The pipeline normalizes base-station metadata and raw histories into a read-only tool store, compiles static tasks with tool-derived gold answers and evidence, then lifts retained tasks into typed bounded operator-facing episodes adding clarification, goal revision, timestamp policy, quality-gated reporting and evidence attribution. Two independent raw-to-episode builds reproduced the released 356/87/89 split exactly, and applying the same path to XAI4HEAT produced 204 episodes, so the reproducibility claim is the artifact rather than the score.
Source
↳ Follow the thread