A-MAR: Agent-Based Multimodal Art Retrieval Uses Multi-Step Reasoning Over Visual and Cultural Context
arXiv·medium signal
Wang, Zhu, and Huang present A-MAR, an agent architecture that performs fine-grained artwork understanding through multi-step reasoning over visual content, cultural context, and historical/stylistic information. Unlike MLLMs that rely on implicit reasoning and internalized knowledge (which hallucinates on domain-specific art queries), A-MAR decomposes understanding into retrieval-augmented reasoning steps. Demonstrates a pattern applicable to any domain requiring multi-hop reasoning over heterogeneous knowledge.