DriveTok: 3D Driving Scene Tokenization Unifies Multi-View Reconstruction and Understanding for VLMs
arXiv 2603.19219·medium signal
DriveTok addresses the scalable image tokenization bottleneck for vision-language-action models in autonomous driving by jointly learning tokens that support both multi-view reconstruction and scene understanding. The unified tokenization enables world models and VLAs to share representations across tasks without task-specific encoders. This approach targets the growing adoption of VLAs in production autonomous driving systems.