AURA: Audio-Geometry Conditioned U-Net Refinement with Flow Matching for High-Fidelity Monaural-to-Binaural Synthesis

Wenjie Zhang* Changjun He* , Yinghan Cao , Shiyun Xu , Mingjiang Wang**
Harbin Institute of Technology, Harbin Institute of Technology (Shenzhen)
Submit to Interspeech

*These authors contributed equally

**Corresponding Author

Abstract

Spatial audio is crucial for immersive listening, and binaural rendering is particularly appealing due to its alignment with human auditory perception. However, synthesizing high-quality binaural audio from monaural input remains challenging, as spatial cues must be inferred while preserving both spatial realism and fine-grained timbral detail, limiting its use in VR/AR and telepresence. To address this issue, we propose AURA, a two-stage binaural synthesis framework that explicitly models the relationship between monaural audio and source direction. The first stage combines global context modeling and local texture extraction to produce a coarse binaural estimate. The second stage refines this estimate with a generative mechanism, enhancing spatial cues and spectral details for high-fidelity output. Extensive experiments show that AURA outperforms state-of-the-art methods in both subjective evaluations and objective metrics (Wave-L2: 0.123 vs. 0.128).

Samples from Binaural Speech Datasets

Sample1 Sample2 Sample3 Sample4
Mono
Groundtruth
WaveNet
WarpNet
NFS
BinauralGrad
DPATFNet
AURA

⚠️ Please use headphones to listen to these audios.