← Back to home

PHD RESEARCH / A CONNECTED SET OF QUESTIONS

Toward unified
audio foundation models.

Can one model understand and generate speech, music, and sound—and adapt to new tasks through examples?

My PhD explores this question through three connected directions: audio representations, scalable unified modeling, and high-fidelity waveform generation. Together, they form a path from specialized audio systems toward a shared modeling framework.

01 / REPRESENT

Audio tokens are a modeling choice.

A codec can reconstruct audio well while producing tokens that are difficult for a language model to learn. For a unified audio model, the representation must also be compact, modelable, and connected to language.

LLM-Codec / UniAudio 1.5

What if audio could use a text model's existing vocabulary? LLM-Codec constrains its codebooks to pretrained LLM embeddings and explores few-shot audio tasks through a frozen language model.

ALMTokenizer

Learnable queries aggregate context across audio frames, creating compact, semantic-rich representations. The goal is to make tokens useful for downstream modeling as well as reconstruction.

Factorized tokenization in UniAudio 2.0

The public paper introduces ReasoningCodec: separate token streams serve text-aligned analysis and fine-grained waveform reconstruction. This gives understanding and generation distinct, complementary interfaces.

The underlying trade-off: shorter sequences reduce modeling cost, but preserving meaning and reconstructing acoustic detail place different demands on the representation.

02 / UNIFY

From many generators to a shared model.

Speech synthesis, voice conversion, sound generation, and music generation are often treated as separate problems. UniAudio explores a shared autoregressive objective across these tasks.

UniAudio: scalable multi-task generation

The Global–Local Transformer separates modeling across time from prediction within each audio frame. This helps manage multiple codec codebooks without flattening everything into one long sequence.

Overview of UniAudio's shared framework for audio generation tasks
UniAudio: a shared framework for diverse audio generation tasks.

UniAudio 2.0: understanding and generation

Adding understanding changes the problem. High-level interpretation and detailed acoustic generation need different capabilities. UniAudio 2.0 combines factorized audio tokens with functionally specialized Transformer layers and staged audio–text training.

Its training also uses related sequences of audio and text—auditory sentences—to learn dependencies across segments and support adaptation to new task compositions.

Beyond our own systems: the Fish Speech team describes its Dual AR choice in relation to UniAudio. Moshi also discusses and extends related nested-Transformer approaches. Fish Speech ↗ · Moshi ↗

03 / GENERATE

Compact representations, rich sound.

Reducing the audio token rate makes generation more manageable for a language model. It also makes the final waveform reconstruction more demanding.

SimpleSpeech: a bounded scalar latent

SimpleSpeech introduces scalar-quantized representations and Transformer diffusion for non-autoregressive speech synthesis. SimpleSpeech 2 develops this direction with flow matching and a simpler training pipeline, without requiring phoneme-level duration alignment.

Connecting speech generation to detokenization

My thesis studies how this scalar-latent approach can serve as a waveform reconstruction component downstream of compact audio tokens. This connects the original speech-synthesis work to the broader audio foundation model pipeline.

The original papers present speech-synthesis systems. The thesis organizes and extends the underlying ideas around high-fidelity detokenization. These are related contributions with different scopes.

04 / CONNECT

The components have to work together.

Representation affects sequence length and predictability. Modeling determines what information the tokens need to preserve. Decoding determines how faithfully predicted tokens become sound. These choices cannot be optimized in isolation.

Audio in → shared modeling → text or audio out

For understanding, the model reads audio and produces text. For generation, it predicts audio tokens that a waveform decoder turns into sound. UniAudio 2.0 provides a concrete system in which these capabilities meet.

The current work leaves important questions open: scaling beyond the explored model sizes, balancing speech with music and sound data, and reducing end-to-end latency. These are opportunities to deepen the framework rather than simply add more tasks.

05 / A WIDER TRAJECTORY

The PhD is part of a longer story.

My research also includes work outside the thesis: early sound perception, language-controlled generation, open codec tools, and music systems. The milestones below use first public release years to show that broader trajectory.

Learning to detect sounds from limited examples

Few-shot sound event detection and reference-conditioned target sound detection.

Diffsound

Exploring text-to-sound generation with discrete diffusion.

Controllable generation, open codecs, and unification

InstructTTS (January 2023) explores speaking-style control through natural-language descriptions. The year also includes HiFi-Codec / AcademiCodec, Make-An-Audio, and the first UniAudio release.

Audio as language; speech in scalar latents

UniAudio 1.5, SimpleSpeech, and SimpleSpeech 2.

Representations designed for audio language models

ALMTokenizer, alongside work on long-form speech and unified speech modeling.

Unified audio and open music systems

UniAudio 2.0 and HeartMuLa: complementary directions in task generality and music generation.

From research methods to music systems

HeartMuLa applies ALMTokenizer’s query-based compression, UniAudio’s Global–Local modeling, and SimpleSpeech’s SQ latent space in a music generation system. See the research-to-application case and listen to a generated song ↗.