Sources
ReferTrack Hits 89.4% Single-Target Success on EVT-Bench With One Camera by Grounding Before Tracking
ReferTrack (arXiv 2607.20061, 22 July, Tencent) replaces abstract reasoning about a tracking target with a referring-then-tracking paradigm: first identify the target from bounding boxes, then generate waypoints from that grounded decision, maintaining a sliding-window queue of prior boxes encoded as 'temporal-viewpoint-bbox indicator tokens.' On EVT-Bench with a single camera it reaches 89.4% on single-target tracking, 73.3% under distraction and 74.1% on ambiguous cases — matching or beating several multi-camera baselines on identification-focused tasks. It was validated with sim-to-real transfer on both legged and humanoid robots, with code released.
↳ Follow the thread