Five connected research lines: unified audio models, text-to-audio generation, audio tokenizers, speech generation, and omni models. Explore each work’s motivation, method, evidence, and significance.
1.0 → 1.5 → 2.0: from unified audio generation to learning tasks from examples, then joint understanding and generation.
ICML 2024 · released 2023
UniAudio 1.0
One modeling language for audio
Audio generation had largely developed as separate systems for speech, music and environmental sound. UniAudio asks whether their knowledge can be shared in one model. It expresses task conditions and target audio as token sequences, and uses a Global–Local Transformer to separate long-range temporal modeling from acoustic detail within each frame. The contribution is a reusable formulation and architecture for multi-task audio generation.
Its architectural influence spans speech and music: Fish Speech explicitly links Dual AR to UniAudio; Moshi cites it among its hierarchical audio-modeling predecessors; Google DeepMind’s Live Music Models describes a similar method citing UniAudio and Moshi. HeartMuLa brings Global–Local modeling into song generation.
Explore the idea Problem · method · evidence · significance
RESEARCH EXPLAINED4 slides · explore at your pace
UNIAUDIO 1.0 / IDEA
Different audio tasks. One modeling language.
Instead of designing a separate generator for every task, UniAudio turns both the conditions and the audio to generate into token sequences.
Phonemes→Speech
Noisy speech→Clean speech
Text description→Sound / music
SHARED SEQUENCE FORMAT
Task IDConditionTarget audio →
One autoregressive model learns to predict the next token across tasks.
7 jointly trained tasks+4 tasks added by fine-tuning
The question: can audio tasks benefit from learning together?
UNIAUDIO 1.0 / METHOD
Long context globally. Acoustic detail locally.
Audio codecs produce several tokens per frame. Flattening them all into one stream makes the expensive, long-context transformer process a much longer sequence.
Flattened baselineLong-context sequence: K × T
1.11.21.32.12.22.33.13.23.34.14.24.3
Time and codebook detail occupy the same long stream.
UniAudio · Global–LocalLong-context sequence: T
GLOBAL across framesFrame 1→Frame 2→Frame 3→Frame 4
LOCAL · within each frame
Code 1→Code 2→Code 3
A small autoregressive model, conditioned on the global state.
Shorten the expensive sequence without dropping within-frame dependencies.
UNIAUDIO 1.0 / EVIDENCE
Sharing helps. Hierarchy cuts training cost.
Two separate ablations test the two central ideas: learning across tasks, and modeling codec tokens at different temporal scales.
01 / JOINT VS. SINGLE-TASK TRAINING
Better sound & music generation
Same backbone · FAD ↓
Single taskJoint training
Text → sound
3.84
3.12
Text → music
5.24
3.65
02 / GLOBAL–LOCAL VS. FLATTENING
Less memory. Less training time.
Matched 3-codebook setup · similar parameter budget
Metric
Flat
G–L
GPU memory
36.7 GB
19.4 GB
Time / iteration
1.63 s
0.73 s
Speech MOS ↑
3.80 ± .09
3.77 ± .05
47%less memory55%less time / iteration
UNIAUDIO 1.0 / INFLUENCE
From unified audio generation to real-time systems.
UniAudio’s separation of temporal context and within-frame detail became part of a broader architecture lineage spanning speech, dialogue, and music.
UniAudio · Global–LocalModel across frames globally; generate acoustic tokens locally.
SPEECH GENERATION · 2024
Fish Speech
Its developers identify Dual AR as “slow-fast (UniAudio)” and report it as their most reliable tested decoding strategy.
Teach a frozen language model an audio task through examples
UniAudio 1.5 changes the question from training a model on many tasks to giving a frozen language model a new audio task through examples. Its LLM-Codec maps audio into token IDs from the LLM’s existing vocabulary. A prompt can then contain labeled audio examples followed by a query. The codec is trained, but the downstream language model does not receive task-specific weight updates.
The importance is the learning mechanism: an audio representation can expose in-context learning already present in a text model. The evidence is a proof of concept in controlled tasks, with clear room to scale.
Explore the idea Problem · method · evidence · significance
RESEARCH EXPLAINED4 slides · explore at your pace
UNIAUDIO 1.5 / LLM-CODEC / QUESTION
Can examples teach a frozen text LLM an audio task?
The challenge is the interface: ordinary codec IDs have no connection to a text model’s vocabulary.
IN-CONTEXT TASK · schematic, not model output
Example AAudio tokens for a dog bark→ “dog”
Example BAudio tokens for a bell→ “bell”
New inputAnother audio token sequence→ ?
Change the demonstrations. Keep the downstream LLM weights fixed.
UNIAUDIO 1.5 / LLM-CODEC / INTERFACE
Quantize audio into the LLM’s existing vocabulary.
LLM-Codec learns the audio interface while reusing frozen vocabulary embeddings.
Semantic tokens can be used alone for understanding; acoustic layers retain information for reconstruction.
UNIAUDIO 1.5 / LLM-CODEC / EVIDENCE
A shared vocabulary enables measurable few-shot transfer.
Frozen LLaMA 2 7B · 2-way classification · 1 example per class · task induction · no repeats.
Sound-event accuracy ↑
BLSP
47%
LLM-Codec
60%
Emotion accuracy ↑
BLSP
29%
LLM-Codec
53%
The same frozen backbone can use examples through the learned audio interface.
UNIAUDIO 1.5 / LLM-CODEC / MEANING
The contribution is a new adaptation mechanism.
Tasks are specified by input–output demonstrations rather than a task-specific backbone update.
LEARN ONCEAudio → text vocabulary
Train the codec to make audio accessible.
CHANGE AT INFERENCEExamples → task behavior
Replace the demonstration pairs in the prompt.
1.0Shared trainingMany generation tasks
1.5Context learningFrozen text backbone
2.0Unified audio modelUnderstanding + generation
A proof of concept on simple tasks, including digit speech and denoising; broader task generalization remains the next challenge.
1 / 4
2026 · preprint
UniAudio 2.0
Connect understanding, generation, and learning from context
UniAudio 2.0 studies how a shared audio language model can understand, generate, and adapt to new tasks. ReasoningCodec separates text-aligned reasoning tokens from acoustic reconstruction tokens. The model assigns different roles to its lower, middle, and upper layers: audio understanding, cross-modal language modeling, and audio generation. Auditory sentences organize related audio and text segments into task-rich training sequences; evaluation covers seen, few-shot, and zero-shot tasks.
This advances the thesis’s central question: how should representation and modeling be designed together so that a model can both interpret and produce audio?
DiffSound and Make-An-Audio explore how text can specify environmental sound, through discrete diffusion and prompt-enhanced latent diffusion.
2022 → TASLP 2023
DiffSound
An early route from language to environmental sound
Diffsound explores generating environmental sound from text. Discrete diffusion predicts and repeatedly refines acoustic tokens jointly, addressing the directional bias, error accumulation and sequential cost of autoregressive decoding.
It establishes an early part of my research trajectory: language becomes a condition for creating sound. The method is evaluated against an autoregressive baseline for both quality and speed.
The top lane constructs compositional supervision; the bottom lane describes inference.
MAKE-AN-AUDIO / EVIDENCE
Quality and prompt alignment improve together.
Reported AudioCaps comparison; DiffSound is the reference system.
FID ↓ · audio distribution
DiffSound
7.17
Make-An-Audio
4.61
CLAP ↑ · text–audio alignment
DiffSound
0.42
Make-An-Audio
0.482
CLAP-conditioned variant. This system comparison changes multiple components; it does not isolate prompt enhancement alone.
MAKE-AN-AUDIO / EXTENSIONS
The generator also becomes an editing tool.
The paper explores ways to condition the same audio-generation approach beyond a caption alone.
TEXT + EXISTING AUDIOModify a soundAdd noise, then denoise toward a new prompt.
MASKED AUDIOFill missing regions
KnownFillKnown
Inpainting uses additional fine-tuning.
IMAGE / VIDEODescribe, then synthesizeConvert visual information into a text condition.
1 / 4
Audio Tokenizer series
HiFi-Codec, ALMTokenizer, and ReasoningCodec develop the audio interface: compact coding, semantic-rich compression, then separate roles for reasoning and reconstruction.
2023 · codec & open-source toolkit
HiFi-Codec
Make audio tokens easier to model—and codecs easier to reproduce
HiFi-Codec treats codec design as part of the generation problem: too many codebooks increase the generator’s prediction burden. Group-residual vector quantization targets high-fidelity reconstruction with four codebooks.
AcademiCodec complements the method with open training code and pretrained codec models, making the representation layer easier for other researchers to reproduce and extend.
Compress the sequence without discarding what the model needs
ALMTokenizer asks what makes a codec useful to an audio language model, beyond reconstructing a waveform. Learnable queries compress information across frames, while semantic objectives encourage tokens to retain information useful for understanding. The result is a low-bitrate interface evaluated both as a codec and inside a shared audio language-model backbone.
This ties compression to downstream usefulness. Its query-based compression later becomes part of HeartCodec, giving a direct path from representation research to music generation.
An audio interface for both understanding and generation
ReasoningCodec separates audio into reasoning tokens and reconstruction tokens, connecting text-aligned analysis with high-fidelity synthesis.
Presented here as the tokenizer contribution within UniAudio 2.0, it extends the representation line from compact coding toward a shared understanding-and-generation interface.
Explore the idea Problem · method · evidence · significance
RESEARCH EXPLAINED4 slides · explore at your pace
REASONINGCODEC / QUESTION
Understanding and reconstruction ask different things of a token.
ReasoningCodec is the representation contribution inside UniAudio 2.0.
MEANING
What does the audio tell us?
Words, events, musical structure, and descriptions.
SIGNAL
What must the decoder recover?
Voice, timbre, texture, and fine acoustic cues.
Factorize these roles while allowing the two representations to communicate.
REASONINGCODEC / METHOD
Two branches, with semantic conditioning between them.
The reasoning branch aligns with language; the reconstruction branch preserves semantic-rich acoustic information.
FiLM scales and shifts reconstruction features using information from the reasoning branch.
REASONINGCODEC / EVIDENCE
The two streams contribute differently to understanding.
Branch ablation in the same downstream evaluation; lower ASR error and higher classification accuracy are better.
Available tokens
ASR WER ↓
Emotion accuracy ↑
Sound accuracy ↑
Reasoning only
10.1
50.2
33.0
Reconstruction only
16.3
42.1
48.7
Both
9.0
56.4
63.3
The combined representation improves all three shown tasks over either stream alone.
REASONINGCODEC / FIDELITY
Semantic abstraction still needs a reconstructable audio stream.
Speech reconstruction at the paper’s matched token-rate setting.
Wideband PESQ ↑
DAC
2.1
ALMTokenizer
2
ReasoningCodec
2.36
PLACE IN THE CODEC SERIES
Change what the representation is for
HiFi-Codec · fewer codebooks
ALMTokenizer · compact, semantic-rich tokens
ReasoningCodec · complementary functional roles
Codec-system comparison; the ASR/classification ablation on the previous page separately tests branch utility.
1 / 4
Speech generative models
InstructTTS studies natural-language style control; SimpleSpeech and SimpleSpeech 2 develop speech synthesis in scalar latent space.
January 2023 → TASLP 2024
InstructTTS
Language controls the delivery, too
InstructTTS studies an early natural-language interface for expressive speech: specify both the words to speak and a free-form description of how they should sound. The work pairs speech with style descriptions, aligns language and speech-style representations, and conditions discrete diffusion on the resulting style embedding. Disentanglement addresses the risk that a style representation also captures unwanted content or speaker information.
The contribution is the connection between an intuitive user instruction and acoustic control. The fixed-text listening comparison makes that connection audible.
Explore the idea Problem · method · evidence · significance
RESEARCH EXPLAINED4 slides · explore at your pace
INSTRUCTTTS / INTERFACE
Describe how to speak, as well as what to say.
InstructTTS explores free-form natural-language style control in an early January 2023 work.
TWO INDEPENDENT USER INPUTS · illustrative
Content“I thought you would be here.”words
Style“Speak softly, with a calm, steady tone.”delivery
A style request can combine emotion, energy, and speaking manner.
INSTRUCTTTS / METHOD
Learn a language–style connection before generating speech.
Paired descriptions and speech teach the style representation; content remains a separate condition.
INSTRUCTTTS / EVIDENCE
Measure naturalness and style relevance separately.
Human ratings with 95% confidence intervals; the baseline uses the same style encoder.
System
Naturalness MOS ↑
Style relevance RMOS ↑
Adapted StyleSpeech baseline
4.04 ± 0.08
3.85 ± 0.10
InstructTTS · Mel
4.35 ± 0.07
4.22 ± 0.09
InstructTTS · Wave / GRVQ
3.95 ± 0.05
4.32 ± 0.07
The strongest style relevance and strongest naturalness occur in different variants.
INSTRUCTTTS / LISTEN
Keep the words fixed; change the style request.
Listen for the requested change in energy and delivery.
学长今天还说他喜欢我呢。你不珍惜我,我就跟别人跑了。
Calm & steady
镇定从容,语气平和,语调稳定
Angry & forceful
语调高昂,声音宏亮,内心非常愤慨
Selected project samples: a qualitative illustration of control. See the paper for systematic evaluation.
1 / 4
Interspeech 2024 · Best Student Paper Award
SimpleSpeech
Simplify speech generation by choosing a better latent space
SimpleSpeech asks whether speech synthesis becomes simpler when the latent space itself is easier to model. SQ-Codec bounds and scalar-quantizes the representation; a Transformer diffusion model generates speech in that space. Together with plain-text input and no phoneme-level duration alignment, this reduces several sources of complexity in non-autoregressive synthesis.
The key result is a representation–generator pairing. Tokenizer-swap experiments test the latent design through downstream speech generation, and SQ-Codec later supports the reconstruction thread in the thesis and HeartCodec.
Training transcripts are obtained with ASR. Sentence-level duration sets output length; word-level alignment is learned implicitly.
SIMPLESPEECH / EVIDENCE
The latent choice affects the generated speech.
Tokenizer-swap ablation: compare VAE and SQ-Codec targets in the synthesis setup.
TTS word error rate ↓
VAE
5.1%
SQ-Codec
3.3%
Predicted quality · DNS-MOS ↑
VAE
3.49
SQ-Codec
3.9
A representation should be evaluated through the generator that uses it.
SIMPLESPEECH / CONTINUATION
The scalar space becomes a reusable reconstruction target.
The method connects a speech synthesis experiment to later audio decoding work.
SIMPLESPEECHSQ-Codec + diffusionGenerate speech in scalar space.
SIMPLESPEECH 2Flow-based generationDevelop the generator and conditioning.
HEARTCODECMusic reconstructionDecode compact music tokens through SQ latents.
1 / 4
2024 → TASLP 2025
SimpleSpeech 2
Turn scalar-latent synthesis into a simpler, faster TTS system
SimpleSpeech 2 develops flow-based synthesis in SQ-Codec’s scalar latent space. Its Transformer uses Time-MoE to make computation depend on the diffusion timestep, alongside voice conditioning that separates speaker identity from recording attributes. Plain-text conditioning and sentence-level duration avoid phoneme-duration annotations. The paper evaluates the individual design choices as well as the complete system’s quality and generation speed.
It strengthens the link between representation and synthesis efficiency. Flow-based scalar generation also provides the methodological basis for the thesis’s reconstruction work.
English zero-shot TTS system comparison, with 95% confidence intervals.
System
Naturalness MOS ↑
Real-time factor ↓
SimpleSpeech
3.97 ± 0.21
1.60
SimpleSpeech 2
4.28 ± 0.12
0.25
RTF 0.25 means 4 seconds of audio per second of compute in this setup.
This is a system comparison, not a claim that flow matching alone causes the entire speed gain.
1 / 4
Omni models
My participation in LongCat-Omni and my first-author work Omni-AutoThink extend this research toward multimodal interaction and adaptive reasoning.
2025 · participation in the LongCat-Omni team
LongCat-Flash-Omni
Bring multimodal understanding into real-time interaction
I participated in LongCat-Flash-Omni, an open omni-modal model for real-time audio-visual interaction. The team combines multimodal perception, a sparse MoE backbone and speech reconstruction, with progressive training across modalities.
This extends my experience from audio models into a large multimodal system where perception and spoken response must work together.
Learn when multimodal reasoning is worth the effort
My first-author paper Omni-AutoThink studies how an omni model can decide when to reason. Adaptive SFT and Adaptive GRPO train the model to choose between a direct answer and a reasoning response across text, audio and visual inputs.
The work adds a decision-making dimension to multimodal capability: allocate reasoning where it helps. Its benchmark measures accuracy and thinking behavior across difficulty levels.
Explore the idea Problem · method · evidence · significance
RESEARCH EXPLAINED4 slides · explore at your pace
OMNI-AUTOTHINK / QUESTION
An omni model should learn when reasoning is useful.
My first-author work studies adaptive reasoning across text, audio, and vision.
Multimodal question
DIRECT ANSWERWhen the evidence is sufficient
Avoid unnecessary reasoning.
THINK, THEN ANSWERWhen inference is needed
Spend reasoning effort on the harder case.
The choice of response mode becomes learned behavior.
OMNI-AUTOTHINK / LEARNING
Reward correct direct answers without rewarding blind shortcuts.
Adaptive SFT initializes both modes; Adaptive GRPO samples them and learns from correctness.
CorrectIncorrectNo-think+2−1Think+10
Adaptive sampling/rejection changes which rollouts contribute to the update.
OMNI-AUTOTHINK / BEHAVIOR
Thinking increases with audio-question difficulty.
Text–audio benchmark · five calibrated difficulty levels.
Responses using thinking mode
L1 · easiest
17%
L2
37%
L3
60%
L4
69%
L5 · hardest
71%
Overall: 73% accuracy with 47% thinking rate.
OMNI-AUTOTHINK / ABLATION
SFT and RL together unlock adaptation across modalities.
Same base model · reported training-stage ablation · Pass@1 accuracy.
Training
Text
Audio + text
Audio + vision + text
Base
35%
64%
48%
+ SFT
49%
64%
50%
+ RL only
60%
66%
50%
+ SFT + RL
66%
73%
69%
The accompanying benchmark evaluates four modality groups over five difficulty levels.
1 / 4
RESEARCH IN PRACTICE
From methods to a music system
2026 · research in application
HeartMuLa
Research methods become a working system
HeartMuLa is a concrete application of methods from my earlier research. ALMTokenizer’s query-based compression contributes to HeartCodec’s token interface; UniAudio’s Global–Local architecture supports hierarchical music modeling; and SQ-Codec from the SimpleSpeech line provides a latent space for flow-based reconstruction.
The case shows cumulative research: representation, modeling and reconstruction become complementary components of a team-built open music system.
My recent work extends audio modeling into multimodal systems, adaptive reasoning, and ongoing research on full-duplex interaction.
ONGOING RESEARCH
Full-duplex omni-modal interaction
I am currently working on full-duplex interaction models across modalities. This direction brings together audio generation, multimodal understanding and interaction over time.
The research goal: models that can keep perceiving while responding, and adjust as the interaction evolves.
PHD RESEARCH
Toward unified audio foundation models
My thesis studies three connected questions: how to represent audio for language modeling, how to unify audio tasks, and how to reconstruct high-fidelity sound from compact tokens.
It brings together the UniAudio series, LLM-Codec, ALMTokenizer, and scalar-latent generation. Earlier work in sound perception and language-controlled generation forms a broader research trajectory.