Agents
HiVLA: Visual-Grounded Hierarchical System Enables Generalizable Robotic Agent Manipulation Without Per-Task Fine-Tuning
Researchers introduce HiVLA, a hierarchical embodied manipulation system that addresses the key limitation of end-to-end Vision-Language-Action models: degraded generalization when fine-tuned on narrow task distributions. HiVLA uses visual grounding as its central organizing principle, enabling robotic agents to generalize across manipulation tasks without per-task fine-tuning. For builders working on embodied agents, this offers a practical architecture for deploying manipulation agents that don't require retraining for each new environment.
Source
↳ Follow the thread