Spatial audio is crucial for immersive listening, and binaural rendering is particularly appealing due to its alignment with human auditory perception. However, synthesizing high-quality binaural audio from monaural input remains challenging, as spatial cues must be inferred while preserving both spatial realism and fine-grained timbral detail, limiting its use in VR/AR and telepresence. To address this issue, we propose AURA, a two-stage binaural synthesis framework that explicitly models the relationship between monaural audio and source direction. The first stage combines global context modeling and local texture extraction to produce a coarse binaural estimate. The second stage refines this estimate with a generative mechanism, enhancing spatial cues and spectral details for high-fidelity output. Extensive experiments show that AURA outperforms state-of-the-art methods in both subjective evaluations and objective metrics (Wave-L2: 0.123 vs. 0.128).
| Sample1 | Sample2 | Sample3 | Sample4 | |
|---|---|---|---|---|
| Mono | ||||
| Groundtruth | ||||
| WaveNet | ||||
| WarpNet | ||||
| NFS | ||||
| BinauralGrad | ||||
| DPATFNet | ||||
| AURA |
⚠️ Please use headphones to listen to these audios.