STARE: Hierarchical RL Framework Treats Diffusion Denoising Trajectory as Attack Surface for VLM Red-Teaming
arXiv·medium signal
STARE introduces a hierarchical reinforcement learning framework that treats the image denoising trajectory itself as the attack surface for red-teaming vision-language models. Unlike existing black-box approaches that only return terminal toxicity scores, STARE exposes when and how toxic semantics emerge during multi-step synthesis under a direct white-box T2I setting. Practitioners building VLM safety filters need to account for trajectory-level attacks, not just final-output screening.