TraVEL Uses Ego-Trajectory Similarity as a GRPO Reward to Fix Motion-Blind Driving Video Retrieval
General-purpose multimodal embedding models retrieve driving clips using static scene shortcuts and fail to distinguish motion-centric events like turning left versus right or accelerating versus decelerating. TraVEL fine-tunes Qwen3-VL-Embedding first with InfoNCE on nuReasoning clip-and-reasoning pairs, then applies Group Relative Policy Optimization with ego-trajectory similarity as the reward — trajectories act purely as privileged training supervision, so retrieval still runs on single-vector embeddings with no ego poses, expert rules, or perception pipeline. Relative to SFT it raised longitudinal and lateral mAP by 9.8 and 4.7 points at 2B and 7.2 and 1.5 points at 8B on a new nuReasoning-derived benchmark.
↳ Follow the thread