AlphaGRPO: Self-Reflective Multimodal Generation via Decompositional Verifiable Reward
arXiv·medium signal
Runhui Huang et al. propose AlphaGRPO, applying Group Relative Policy Optimization to AR-Diffusion Unified Multimodal Models for self-reflective multimodal generation. The framework introduces decompositional verifiable rewards that allow the model to assess and improve its own text-and-image outputs without human feedback. Extends GRPO beyond text-only RLHF into unified vision-language generation, potentially enabling self-improving multimodal agents.