Activity Frames: Compiling a Day of Screen Capture Into Agent Memory Deterministically — 86x Smaller in 68ms, 98.4% Answer Accuracy vs 66-80% for LLM Summaries
An arXiv paper submitted August 6 (2608.05784) proposes a pipeline with no model in the loop that compiles passively captured screen activity into a prompt-ready context block, making outputs byte-identical, cacheable and mechanically auditable. On a corpus of 128,756 frames across 51 active days from one professional user, it compresses a full day 86x in 68ms and an agent answering questions about that activity reaches 98.4% accuracy (Wilson 95% CI 91.7-99.7%) against 66-80% for LLM-generated summaries, plus deterministic routine replay at zero model tokens on matching scenarios. It also quantifies Routine Overhead Ratios of 60-343x and a delegable-recurrence rate of 9.0% in-sample / 7.7% out-of-sample.
↳ Follow the thread