UniAudio 1.0
One modeling language for audio
Audio generation had largely developed as separate systems for speech, music and environmental sound. UniAudio asks whether their knowledge can be shared in one model. It expresses task conditions and target audio as token sequences, and uses a Global–Local Transformer to separate long-range temporal modeling from acoustic detail within each frame. The contribution is a reusable formulation and architecture for multi-task audio generation.
Its architectural influence spans speech and music: Fish Speech explicitly links Dual AR to UniAudio; Moshi cites it among its hierarchical audio-modeling predecessors; Google DeepMind’s Live Music Models describes a similar method citing UniAudio and Moshi. HeartMuLa brings Global–Local modeling into song generation.