FLUX-mimic Turns the FLUX 3 Backbone Into a Robot Controller Running on Audi Production Lines at 101ms Reaction Time
Announced alongside FLUX 3 on July 23, FLUX-mimic bolts a lightweight action decoder onto intermediate features of the FLUX 3 backbone to produce a video-action model, and BFL reports it beats prior vision-language-action models even with the backbone frozen, reaching state-of-the-art success rates when the backbone is finetuned. It is deployed at Audi facilities doing kitting, component insertion and flexible-material manipulation, with a 101ms reaction time BFL compares to human visual response, and exhibits failure recovery it was never explicitly shown. The builder takeaway is that generative video pretraining is now being harvested as a robotics world model rather than trained separately — no public release date was given.
↳ Follow the thread