Tools
agent-vision-toolkit Gives Text-Only Models Working Eyes — 345 Stars, MIT, Ships as a Skill
Anionex/agent-vision-toolkit (created 2026-08-01, 345 stars, Python, MIT, pushed 2026-08-07) is a vision toolbox and skill for text-only LLMs: multi-image understanding, image Q&A, long-screenshot OCR, frontend UI reconstruction, and GUI automation, with drop-in integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode. It matters because cheap text-only models are often the cost-effective choice for agent loops, and this externalizes vision as a tool call instead of forcing an upgrade to a multimodal tier for the 5% of steps that need to see something.
Source
↳ Follow the thread