Research
ARGUS Names the Correct Kubernetes Root Cause in All Ten Injected Faults, but On-Call Engineers Distrust Its Fixes
ARGUS connects a commercial LLM to live Kubernetes observability data through standardised MCP servers covering cluster state, Prometheus metrics, Loki logs, and NATS messaging, delivering structured diagnostics into the Slack incident channel where on-call engineers already work. Across controlled fault injection on ten incident scenarios it named the correct root cause in all ten with an aggregate MCP success ratio of 0.91. Interviews with six on-call engineers at an industrial partner surfaced a diagnostic/prescriptive asymmetry: practitioners trusted what the system said went wrong but were consistently sceptical of its recommended fixes.
↳ Follow the thread