Watch, Remember, Reason: Human-View Video Understanding with MLLMs
arXiv·medium signal
As multimodal LLM video research moves from short clips to long, multimodal human-view (egocentric) footage, this work proposes a watch-remember-reason pipeline pairing memory with reasoning for long-video understanding. The memory-plus-reasoning decomposition is the actionable angle for builders working on agentic video, surveillance, or assistant systems that must retain context across extended footage.