AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis With Fine-Grained Acoustic Transfer
arXiv 2603.15597·medium signal
AC-Foley replaces text prompts in video-to-audio generation with a reference audio sample, enabling fine-grained acoustic feature transfer that text labels cannot capture due to semantic granularity gaps (e.g., distinguishing multiple acoustically distinct sounds grouped under one label). The model addresses two fundamental V2A bottlenecks — coarse training data labels and textual ambiguity — by learning to transfer micro-acoustic characteristics from a reference clip to synthesized audio. Directly actionable for multimodal content generation pipelines.