Dongchao Yang杨东超

Audio foundation models & omni-modal interaction

I build models that understand and generate audio. My research connects language-based control, audio representations, and unified audio modeling. HeartMuLa brings methods from these threads into music generation. I also work on omni models, including LongCat-Omni and Omni-AutoThink, and am currently exploring full-duplex omni-modal interaction.

PhD student at The Chinese University of Hong Kong,
advised by Prof. Helen Meng. Previously, a master's student at Peking University.

Portrait of Dongchao Yang

Recently↘

UniSRM accepted to ACL 2026 Main; nominated for Best Paper.

UniAudio 2.0 — unified understanding and generation across speech, music, and sound.

HeartMuLa — our open-source family of music foundation models.

Five connected research lines: unified audio models, text-to-audio generation, audio tokenizers, speech generation, and omni models. Explore each work’s motivation, method, evidence, and significance.

UniAudio series

1.0 → 1.5 → 2.0: from unified audio generation to learning tasks from examples, then joint understanding and generation.

UniAudio 1.0

One modeling language for audio

Audio generation had largely developed as separate systems for speech, music and environmental sound. UniAudio asks whether their knowledge can be shared in one model. It expresses task conditions and target audio as token sequences, and uses a Global–Local Transformer to separate long-range temporal modeling from acoustic detail within each frame. The contribution is a reusable formulation and architecture for multi-task audio generation.

Its architectural influence spans speech and music: Fish Speech explicitly links Dual AR to UniAudio; Moshi cites it among its hierarchical audio-modeling predecessors; Google DeepMind’s Live Music Models describes a similar method citing UniAudio and Moshi. HeartMuLa brings Global–Local modeling into song generation.

Explore the idea Problem · method · evidence · significance
RESEARCH EXPLAINED4 slides · explore at your pace

UNIAUDIO 1.0 / IDEA

Different audio tasks. One modeling language.

Instead of designing a separate generator for every task, UniAudio turns both the conditions and the audio to generate into token sequences.

Phonemes→Speech
Noisy speech→Clean speech
Text description→Sound / music
SHARED SEQUENCE FORMAT
Task IDConditionTarget audio →

One autoregressive model learns to predict the next token across tasks.

7 jointly trained tasks+4 tasks added by fine-tuning

The question: can audio tasks benefit from learning together?

UniAudio paper ↗ · Task formulation and training setup. Examples shown; not an exhaustive task list.

UNIAUDIO 1.0 / METHOD

Long context globally. Acoustic detail locally.

Audio codecs produce several tokens per frame. Flattening them all into one stream makes the expensive, long-context transformer process a much longer sequence.

Flattened baselineLong-context sequence: K × T
1.11.21.32.12.22.33.13.23.34.14.24.3

Time and codebook detail occupy the same long stream.

UniAudio · Global–LocalLong-context sequence: T
GLOBAL
across frames
Frame 1→Frame 2→Frame 3→Frame 4
LOCAL · within each frame
Code 1→Code 2→Code 3
A small autoregressive model, conditioned on the global state.

Shorten the expensive sequence without dropping within-frame dependencies.

UniAudio paper ↗ · §2.3. Schematic: T = 4 frames, K = 3 codebooks. Both levels are autoregressive; the local decoder is not parallel.

UNIAUDIO 1.0 / EVIDENCE

Sharing helps. Hierarchy cuts training cost.

Two separate ablations test the two central ideas: learning across tasks, and modeling codec tokens at different temporal scales.

01 / JOINT VS. SINGLE-TASK TRAINING
Better sound & music generation

Same backbone · FAD ↓

Single taskJoint training
Text → sound
3.84
3.12
Text → music
5.24
3.65
02 / GLOBAL–LOCAL VS. FLATTENING
Less memory. Less training time.

Matched 3-codebook setup · similar parameter budget

MetricFlatG–L
GPU memory36.7 GB19.4 GB
Time / iteration1.63 s0.73 s
Speech MOS ↑3.80 ± .093.77 ± .05
47%less memory55%less time / iteration
UniAudio paper ↗ · Multi-task ablation and Table 4. Architecture benchmark: LibriTTS, 20-second clips, 100-trial averages; authors’ baseline implementations. Percent reductions calculated from reported values; timing measures training, not inference.

UNIAUDIO 1.0 / INFLUENCE

From unified audio generation to real-time systems.

UniAudio’s separation of temporal context and within-frame detail became part of a broader architecture lineage spanning speech, dialogue, and music.

UniAudio · Global–LocalModel across frames globally; generate acoustic tokens locally.
SPEECH GENERATION · 2024
Fish Speech

Its developers identify Dual AR as “slow-fast (UniAudio)” and report it as their most reliable tested decoding strategy.

Developer report ↗
REAL-TIME DIALOGUE · 2024
Moshi · Kyutai

Uses Temporal and Depth Transformers; explicitly cites UniAudio among the audio-modeling predecessors it extends.

§3.4 · architecture ↗
LIVE MUSIC · 2025
Magenta RT · Google DeepMind

Live Music Models describes its efficient modeling method as similar to UniAudio and Moshi, citing both directly.

§2 · method, refs. 21–22 ↗
SONG GENERATION · 2026
HeartMuLa

Its hierarchical music generator follows UniAudio’s factorization: global structure first, then local acoustic tokens.

§3.1 · architecture ↗
Sources linked per example. Moshi also credits RQ-Transformer and other predecessors; Live Music Models describes a similar method. These are documented architectural connections, not claims of exclusive origin.

UniAudio 1.5 / LLM-Codec

Teach a frozen language model an audio task through examples

UniAudio 1.5 changes the question from training a model on many tasks to giving a frozen language model a new audio task through examples. Its LLM-Codec maps audio into token IDs from the LLM’s existing vocabulary. A prompt can then contain labeled audio examples followed by a query. The codec is trained, but the downstream language model does not receive task-specific weight updates.

The importance is the learning mechanism: an audio representation can expose in-context learning already present in a text model. The evidence is a proof of concept in controlled tasks, with clear room to scale.

Explore the idea Problem · method · evidence · significance
RESEARCH EXPLAINED4 slides · explore at your pace

UNIAUDIO 1.5 / LLM-CODEC / QUESTION

Can examples teach a frozen text LLM an audio task?

The challenge is the interface: ordinary codec IDs have no connection to a text model’s vocabulary.

IN-CONTEXT TASK · schematic, not model output
Example AAudio tokens for a dog bark→ “dog”
Example BAudio tokens for a bell→ “bell”
New inputAnother audio token sequence→ ?

Change the demonstrations. Keep the downstream LLM weights fixed.

UNIAUDIO 1.5 / LLM-CODEC / INTERFACE

Quantize audio into the LLM’s existing vocabulary.

LLM-Codec learns the audio interface while reusing frozen vocabulary embeddings.

AudioTrained encoderMulti-scale quantizationSemantic · coarse · residualLLM vocabulary embeddingsVocabulary token IDsCodec trained / LLaMA 2 7B frozen

Semantic tokens can be used alone for understanding; acoustic layers retain information for reconstruction.

UNIAUDIO 1.5 / LLM-CODEC / EVIDENCE

A shared vocabulary enables measurable few-shot transfer.

Frozen LLaMA 2 7B · 2-way classification · 1 example per class · task induction · no repeats.

Sound-event accuracy ↑
BLSP
47%
LLM-Codec
60%
Emotion accuracy ↑
BLSP
29%
LLM-Codec
53%

The same frozen backbone can use examples through the learned audio interface.

Table 2 ↗ · Semantic-layer configuration; controlled few-shot tasks.

UNIAUDIO 1.5 / LLM-CODEC / MEANING

The contribution is a new adaptation mechanism.

Tasks are specified by input–output demonstrations rather than a task-specific backbone update.

LEARN ONCEAudio → text vocabulary

Train the codec to make audio accessible.

CHANGE AT INFERENCEExamples → task behavior

Replace the demonstration pairs in the prompt.

1.0Shared trainingMany generation tasks
1.5Context learningFrozen text backbone
2.0Unified audio modelUnderstanding + generation

A proof of concept on simple tasks, including digit speech and denoising; broader task generalization remains the next challenge.

UniAudio 2.0

Connect understanding, generation, and learning from context

UniAudio 2.0 studies how a shared audio language model can understand, generate, and adapt to new tasks. ReasoningCodec separates text-aligned reasoning tokens from acoustic reconstruction tokens. The model assigns different roles to its lower, middle, and upper layers: audio understanding, cross-modal language modeling, and audio generation. Auditory sentences organize related audio and text segments into task-rich training sequences; evaluation covers seen, few-shot, and zero-shot tasks.

This advances the thesis’s central question: how should representation and modeling be designed together so that a model can both interpret and produce audio?

Explore the idea Problem · method · evidence · significance
RESEARCH EXPLAINED4 slides · explore at your pace

UNIAUDIO 2.0 / QUESTION

A foundation model should learn more than a task catalog.

UniAudio 2.0 connects audio understanding, generation, and adaptation within a shared text–audio model.

UNDERSTANDAudio → textTranscription · captions · lyrics
GENERATEText → audioSpeech · sound · music · song
ADAPTExamples → new mappingDenoising · conversion · classification

The system must connect meanings, acoustic detail, and task structure.

UNIAUDIO 2.0 / ARCHITECTURE

Give different layers different jobs.

ReasoningCodec provides the tokens; specialized model layers connect perception, language, and acoustic generation.

Text + audio streamsLower: audio understandingMiddle: cross-modal LLMUpper: audio generationText outputLocal decoder → audio tokensOne autoregressive framework

UNIAUDIO 2.0 / TRANSFER

Test task adaptation explicitly, beyond seen-task scores.

Selected reported evaluations distinguish learning from examples from zero-shot instruction transfer.

FEW-SHOT · ENGLISH EMOTION
One example per class
Accuracy ↑
UniAudio 1.5
45%
MiMo-Audio
42.5%
UniAudio 2.0
67%
ZERO-SHOT · DYSARTHRIC ASR
Recognize atypical speech
Word error rate ↓
Qwen2.5-Omni
80.6%
UniAudio 2.0
19.4%

These are selected strengths, not universal superiority: MiMo-Audio is stronger on 2-shot voice-conversion WER (14.05 vs. 19.01).

Tables 5–6 ↗ · Different systems and training data; not a controlled architecture ablation.

UNIAUDIO 2.0 / TRAINING

Build task relationships into the training sequence.

“Auditory sentences” join related text and audio segments so context can express a task.

AUDITORY SENTENCE · conceptual sequence
Segment 1Related audio + descriptioncontext
Segment 2Related audio + descriptionrelationship
ContinuationPredict the next text or audio spantask
WARM-UPPerception + generationInitialize specialized layers
PRETRAINText + audioBuild shared capability
MID-TRAINTask compositionsPractice using context

Text-to-audio series

DiffSound and Make-An-Audio explore how text can specify environmental sound, through discrete diffusion and prompt-enhanced latent diffusion.

DiffSound

An early route from language to environmental sound

Diffsound explores generating environmental sound from text. Discrete diffusion predicts and repeatedly refines acoustic tokens jointly, addressing the directional bias, error accumulation and sequential cost of autoregressive decoding.

It establishes an early part of my research trajectory: language becomes a condition for creating sound. The method is evaluated against an autoregressive baseline for both quality and speed.

Explore the idea Problem · method · evidence · significance
RESEARCH EXPLAINED4 slides · explore at your pace

DIFFSOUND / QUESTION

Generate the scene a sentence describes.

A sound description can combine events and backgrounds that must remain coherent over time.

“A dog barks while rain falls in the background.”
Foreground eventDog barkingAcoustic backgroundRainRelationshipSimultaneous

The target is a structured acoustic scene, not a spoken reading of the prompt.

DIFFSOUND / METHOD

Refine all acoustic tokens together.

Discrete diffusion revisits the whole spectrogram-token grid at each denoising step.

Autoregressive baseline●●●●····Predict the next tokenDiffSound······●●●···●●●●●●Corrupted tokensJoint refinementDecoded spectrogram

Schematic token states. The codec decoder and vocoder then render the waveform.

DIFFSOUND / EVIDENCE

Joint refinement improves relevance and fidelity.

Human ratings against the paper’s autoregressive token-decoder baseline.

Text relevance ↑
AR
2.747
DiffSound
3.833
Overall MOS ↑
AR
2.786
DiffSound
3.56

Reported generation speed: 5× the AR baseline in the paper’s inference setup.

DIFFSOUND / PLACE IN FIELD

An early discrete-diffusion route to text-to-audio.

The work made text-conditioned sound generation reproducible through code, models, and examples.

EXPLICIT TECHNICAL CONTEXT
Make-An-Audio · related work

Identifies DiffSound as an early text-to-audio method using discrete diffusion over VQ-VAE codes.

Read the discussion ↗
DIFFSOUNDDiscrete codesIterative token refinement
MAKE-AN-AUDIOContinuous latentsPrompt-enhanced diffusion

Make-An-Audio

Improve text-to-audio through better supervision and latent modeling

Make-An-Audio addresses scarce text–audio pairs and difficult waveform modeling through pseudo prompt enhancement, spectrogram latent diffusion and CLAP conditioning.

It extends the text-to-audio thread with complementary work on supervision and representation.

Explore the idea Problem · method · evidence · significance
RESEARCH EXPLAINED4 slides · explore at your pace

MAKE-AN-AUDIO / BOTTLENECK

Text-to-audio has a data problem and a modeling problem.

Paired captions are scarce, while raw waveforms are expensive generation targets.

SUPERVISION
Use more uncaptioned audio

Construct pseudo prompts and new concept combinations.

GENERATION
Model compressed spectrograms

Generate continuous latents conditioned on language.

Make-An-Audio addresses both constraints in one text-to-audio system.

MAKE-AN-AUDIO / METHOD

Improve the supervision before improving the generator.

Prompt enhancement expands training pairs; latent diffusion turns prompts into sound.

TRAINING DATAUncaptioned audio→Expert distillation→Dynamic reprogramming→Pseudo prompts
GENERATIONText / CLAP→Latent diffusion→Spectrogram decoder→Vocoder

The top lane constructs compositional supervision; the bottom lane describes inference.

MAKE-AN-AUDIO / EVIDENCE

Quality and prompt alignment improve together.

Reported AudioCaps comparison; DiffSound is the reference system.

FID ↓ · audio distribution
DiffSound
7.17
Make-An-Audio
4.61
CLAP ↑ · text–audio alignment
DiffSound
0.42
Make-An-Audio
0.482

CLAP-conditioned variant. This system comparison changes multiple components; it does not isolate prompt enhancement alone.

MAKE-AN-AUDIO / EXTENSIONS

The generator also becomes an editing tool.

The paper explores ways to condition the same audio-generation approach beyond a caption alone.

TEXT + EXISTING AUDIOModify a soundAdd noise, then denoise toward a new prompt.
MASKED AUDIOFill missing regions
KnownFillKnown
Inpainting uses additional fine-tuning.
IMAGE / VIDEODescribe, then synthesizeConvert visual information into a text condition.

Audio Tokenizer series

HiFi-Codec, ALMTokenizer, and ReasoningCodec develop the audio interface: compact coding, semantic-rich compression, then separate roles for reasoning and reconstruction.

HiFi-Codec

Make audio tokens easier to model—and codecs easier to reproduce

HiFi-Codec treats codec design as part of the generation problem: too many codebooks increase the generator’s prediction burden. Group-residual vector quantization targets high-fidelity reconstruction with four codebooks.

AcademiCodec complements the method with open training code and pretrained codec models, making the representation layer easier for other researchers to reproduce and extend.

Explore the idea Problem · method · evidence · significance
RESEARCH EXPLAINED4 slides · explore at your pace

HIFI-CODEC / QUESTION

Every extra codebook adds predictions for the generator.

Codec quality and generation cost are linked through the token interface.

8 CODEBOOKS / FRAME
8 token predictions
4 CODEBOOKS / FRAME
4 token predictions

HiFi-Codec targets reconstruction quality with fewer codebooks.

Motivation ↗ · Illustration at a fixed frame rate; not an inference-speed measurement.

HIFI-CODEC / METHOD

Split feature channels; refine each group’s residual.

Group-residual vector quantization gives each group its own quantization path.

Encoded featuresGroup AGroup BQ1Q2Q3Q4Quantize residual within each groupCombine → decodeSplit channels, not time

Shown: two groups × two residual stages = four codebooks.

HIFI-CODEC / EVIDENCE

Four codebooks retain quality in a matched-rate comparison.

24 kHz audio · 240× downsampling · 100 frames per second.

CodecCodebooksPESQ ↑STOI ↑
EnCodec · authors’ reproduction83.620.94
HiFi-Codec43.630.95

800 → 400 tokens/second at this frame rate, with comparable reported reconstruction scores.

Derived token rates. A codec-system comparison, not an isolated GRVQ ablation.

HIFI-CODEC / OPEN RESEARCH

Release the training tools as well as the codec.

AcademiCodec makes the representation layer available for reproduction and downstream modeling.

ACADEMICODEC
Train · encode · reconstruct

HiFi-Codec, EnCodec, and SoundStream implementations with pretrained models.

Explore the toolkit ↗
RESEARCH THREAD
A smaller audio interface

HiFi-Codec studies codebook efficiency. ALMTokenizer adds compression across time and semantic supervision.

Continue to ALMTokenizer ↗

ALMTokenizer

Compress the sequence without discarding what the model needs

ALMTokenizer asks what makes a codec useful to an audio language model, beyond reconstructing a waveform. Learnable queries compress information across frames, while semantic objectives encourage tokens to retain information useful for understanding. The result is a low-bitrate interface evaluated both as a codec and inside a shared audio language-model backbone.

This ties compression to downstream usefulness. Its query-based compression later becomes part of HeartCodec, giving a direct path from representation research to music generation.

Explore the idea Problem · method · evidence · significance
RESEARCH EXPLAINED4 slides · explore at your pace

ALMTOKENIZER / QUESTION

A good codec must also produce tokens a model can use.

Reconstruction quality alone does not tell us whether a token sequence supports understanding or generation.

CompactFewer time steps
SemanticMeaning survives compression
ReconstructableAcoustic detail survives decoding

Optimize the interface for all three requirements, across speech, sound, and music.

ALMTOKENIZER / METHOD

Learnable queries compress across neighboring frames.

The encoder gathers temporal context before quantization rather than compressing each frame independently.

Audio patches●●●●●●●●QueryContextual encoderKeep queryQuantize → compact audio tokensSemantic priors + masked reconstruction + token predictionTrain for meaning, fidelity, and predictability together

Schematic compression window. Token prediction is part of the training objective, not an extra inference stage.

ALMTOKENIZER / EVIDENCE

Changing the tokenizer changes downstream language modeling.

LibriSpeech test-clean · tokenizer comparison within the shared downstream setup.

TokenizerTTS WER ↓ASR WER ↓
DAC24.558.4 ± 1.2
Mimi16.023.1 ± 1.5
GLM-4-Voice9.916.3 ± 1.5
ALMTokenizer11.719.6 ± 1.8

Large gains over reconstruction-oriented tokens; the speech-specialized GLM-4-Voice remains stronger here.

Table 4 ↗ · WER: word error rate. ALMTokenizer also covers sound and music.

ALMTOKENIZER / APPLICATION

Query compression becomes HeartCodec’s token interface.

HeartMuLa is a concrete application of the earlier representation research.

ALMTOKENIZERQuery compressionSummarize context into fewer time steps
HEARTCODECLow-rate music tokensCreate a manageable generation target
HEARTMULAMusic language modelPredict compact tokens over time
WHY IT MATTERS
Sequence length is an architectural decision

Compression determines how much audio fits into the generator’s context.

TRACE THE CONNECTION
Documented in HeartCodec

The report describes query-based compression and cites ALMTokenizer.

Read §2.1 ↗

ReasoningCodec

An audio interface for both understanding and generation

ReasoningCodec separates audio into reasoning tokens and reconstruction tokens, connecting text-aligned analysis with high-fidelity synthesis.

Presented here as the tokenizer contribution within UniAudio 2.0, it extends the representation line from compact coding toward a shared understanding-and-generation interface.

Explore the idea Problem · method · evidence · significance
RESEARCH EXPLAINED4 slides · explore at your pace

REASONINGCODEC / QUESTION

Understanding and reconstruction ask different things of a token.

ReasoningCodec is the representation contribution inside UniAudio 2.0.

MEANING
What does the audio tell us?

Words, events, musical structure, and descriptions.

SIGNAL
What must the decoder recover?

Voice, timbre, texture, and fine acoustic cues.

Factorize these roles while allowing the two representations to communicate.

REASONINGCODEC / METHOD

Two branches, with semantic conditioning between them.

The reasoning branch aligns with language; the reconstruction branch preserves semantic-rich acoustic information.

AudioReasoning branchReconstruction branchReasoning tokensReconstruction tokensFiLM conditioningText / planningFlow decoder → audio

FiLM scales and shifts reconstruction features using information from the reasoning branch.

REASONINGCODEC / EVIDENCE

The two streams contribute differently to understanding.

Branch ablation in the same downstream evaluation; lower ASR error and higher classification accuracy are better.

Available tokensASR WER ↓Emotion accuracy ↑Sound accuracy ↑
Reasoning only10.150.233.0
Reconstruction only16.342.148.7
Both9.056.463.3

The combined representation improves all three shown tasks over either stream alone.

Table 3 ↗ · Selected columns; music classification favors reasoning-only tokens (80 vs. 70).

REASONINGCODEC / FIDELITY

Semantic abstraction still needs a reconstructable audio stream.

Speech reconstruction at the paper’s matched token-rate setting.

Wideband PESQ ↑
DAC
2.1
ALMTokenizer
2
ReasoningCodec
2.36
PLACE IN THE CODEC SERIES
Change what the representation is for

HiFi-Codec · fewer codebooks

ALMTokenizer · compact, semantic-rich tokens

ReasoningCodec · complementary functional roles

Codec-system comparison; the ASR/classification ablation on the previous page separately tests branch utility.

Speech generative models

InstructTTS studies natural-language style control; SimpleSpeech and SimpleSpeech 2 develop speech synthesis in scalar latent space.

InstructTTS

Language controls the delivery, too

InstructTTS studies an early natural-language interface for expressive speech: specify both the words to speak and a free-form description of how they should sound. The work pairs speech with style descriptions, aligns language and speech-style representations, and conditions discrete diffusion on the resulting style embedding. Disentanglement addresses the risk that a style representation also captures unwanted content or speaker information.

The contribution is the connection between an intuitive user instruction and acoustic control. The fixed-text listening comparison makes that connection audible.

Explore the idea Problem · method · evidence · significance
RESEARCH EXPLAINED4 slides · explore at your pace

INSTRUCTTTS / INTERFACE

Describe how to speak, as well as what to say.

InstructTTS explores free-form natural-language style control in an early January 2023 work.

TWO INDEPENDENT USER INPUTS · illustrative
Content“I thought you would be here.”words
Style“Speak softly, with a calm, steady tone.”delivery

A style request can combine emotion, energy, and speaking manner.

INSTRUCTTTS / METHOD

Learn a language–style connection before generating speech.

Paired descriptions and speech teach the style representation; content remains a separate condition.

Style descriptionAligned style embeddingWords / phonemesContent encoderDiscrete diffusionSpeechReduce content and speaker leakage from the style representation

INSTRUCTTTS / EVIDENCE

Measure naturalness and style relevance separately.

Human ratings with 95% confidence intervals; the baseline uses the same style encoder.

SystemNaturalness MOS ↑Style relevance RMOS ↑
Adapted StyleSpeech baseline4.04 ± 0.083.85 ± 0.10
InstructTTS · Mel4.35 ± 0.074.22 ± 0.09
InstructTTS · Wave / GRVQ3.95 ± 0.054.32 ± 0.07

The strongest style relevance and strongest naturalness occur in different variants.

INSTRUCTTTS / LISTEN

Keep the words fixed; change the style request.

Listen for the requested change in energy and delivery.

学长今天还说他喜欢我呢。你不珍惜我,我就跟别人跑了。

Calm & steady

镇定从容,语气平和,语调稳定

Angry & forceful

语调高昂,声音宏亮,内心非常愤慨

Selected project samples: a qualitative illustration of control. See the paper for systematic evaluation.

SimpleSpeech

Simplify speech generation by choosing a better latent space

SimpleSpeech asks whether speech synthesis becomes simpler when the latent space itself is easier to model. SQ-Codec bounds and scalar-quantizes the representation; a Transformer diffusion model generates speech in that space. Together with plain-text input and no phoneme-level duration alignment, this reduces several sources of complexity in non-autoregressive synthesis.

The key result is a representation–generator pairing. Tokenizer-swap experiments test the latent design through downstream speech generation, and SQ-Codec later supports the reconstruction thread in the thesis and HeartCodec.

Explore the idea Problem · method · evidence · significance
RESEARCH EXPLAINED4 slides · explore at your pace

SIMPLESPEECH / IDEA

Make the speech latent easier to generate.

SQ-Codec bounds and quantizes each latent coordinate, creating the target space for Transformer diffusion.

Bound with tanhValue in [−1, 1]Round to a fixed scalar gridCompact latent target−101

Illustrative grid. With S = 9, a bounded value 0.71 becomes round(0.71 × 9) / 9 ≈ 0.667.

SIMPLESPEECH / SYSTEM

Generate the whole latent sequence without phoneme durations.

Text, reference-voice features, and sentence duration condition a non-autoregressive Transformer.

Plain textReference voiceSentence duration
STARTGaussian noiseOne full latent sequence
REFINETransformer diffusionConditioned iterative denoising
RECONSTRUCTSQ-Codec decoderRender the waveform

Training transcripts are obtained with ASR. Sentence-level duration sets output length; word-level alignment is learned implicitly.

SIMPLESPEECH / EVIDENCE

The latent choice affects the generated speech.

Tokenizer-swap ablation: compare VAE and SQ-Codec targets in the synthesis setup.

TTS word error rate ↓
VAE
5.1%
SQ-Codec
3.3%
Predicted quality · DNS-MOS ↑
VAE
3.49
SQ-Codec
3.9

A representation should be evaluated through the generator that uses it.

Table 5 ↗ · DNS-MOS is an automatic estimate, not a human MOS.

SIMPLESPEECH / CONTINUATION

The scalar space becomes a reusable reconstruction target.

The method connects a speech synthesis experiment to later audio decoding work.

SIMPLESPEECHSQ-Codec + diffusionGenerate speech in scalar space.
SIMPLESPEECH 2Flow-based generationDevelop the generator and conditioning.
HEARTCODECMusic reconstructionDecode compact music tokens through SQ latents.

SimpleSpeech 2

Turn scalar-latent synthesis into a simpler, faster TTS system

SimpleSpeech 2 develops flow-based synthesis in SQ-Codec’s scalar latent space. Its Transformer uses Time-MoE to make computation depend on the diffusion timestep, alongside voice conditioning that separates speaker identity from recording attributes. Plain-text conditioning and sentence-level duration avoid phoneme-duration annotations. The paper evaluates the individual design choices as well as the complete system’s quality and generation speed.

It strengthens the link between representation and synthesis efficiency. Flow-based scalar generation also provides the methodological basis for the thesis’s reconstruction work.

Explore the idea Problem · method · evidence · significance
RESEARCH EXPLAINED4 slides · explore at your pace

SIMPLESPEECH 2 / ADVANCE

A faster generator in the same scalar-latent direction.

SimpleSpeech 2 develops flow matching, timestep-aware experts, and more robust voice conditioning.

SIMPLESPEECHScalar diffusion

Establish the bounded latent target.

SIMPLESPEECH 2Flow-based Transformer

Learn the movement from noise toward speech.

100 → 25Reported diffusion sampling steps
1.60 → 0.25Reported real-time factor

System comparison: training data and architecture also change.

SIMPLESPEECH 2 / METHOD

Let the denoising time help choose the expert.

Time-MoE combines the latent state and diffusion timestep to route computation inside the flow-based Transformer.

Noisy speech latentDiffusion timestepRouterExpert 1Expert 2Expert 3Expert 4Latent updateTimestep-aware feed-forward computation

SIMPLESPEECH 2 / ABLATION

Both Time-MoE and flow matching improve the tested setup.

Architecture/formulation ablations provide a closer test than comparing complete TTS systems.

SettingWER ↓DNS-MOS ↑
DDPM formulation9.03.82
Without Time-MoE8.23.80
SimpleSpeech 27.53.85

The latent space is only part of the answer; how the generator uses it also matters.

SIMPLESPEECH 2 / QUALITY

Faster generation accompanies better reported naturalness.

English zero-shot TTS system comparison, with 95% confidence intervals.

SystemNaturalness MOS ↑Real-time factor ↓
SimpleSpeech3.97 ± 0.211.60
SimpleSpeech 24.28 ± 0.120.25

RTF 0.25 means 4 seconds of audio per second of compute in this setup.

This is a system comparison, not a claim that flow matching alone causes the entire speed gain.

Omni models

My participation in LongCat-Omni and my first-author work Omni-AutoThink extend this research toward multimodal interaction and adaptive reasoning.

LongCat-Flash-Omni

Bring multimodal understanding into real-time interaction

I participated in LongCat-Flash-Omni, an open omni-modal model for real-time audio-visual interaction. The team combines multimodal perception, a sparse MoE backbone and speech reconstruction, with progressive training across modalities.

This extends my experience from audio models into a large multimodal system where perception and spoken response must work together.

Explore the idea Problem · method · evidence · significance
RESEARCH EXPLAINED4 slides · explore at your pace

LONGCAT-FLASH-OMNI / SYSTEM

Connect what the model sees, hears, and says.

I participated in the LongCat-Flash-Omni team project.

VIDEOFrame 1Frame 2Frame 3
AUDIOChunk 1Chunk 2Chunk 3

Synchronized input chunks preserve the relationship between visual events and speech.

LONGCAT-FLASH-OMNI / SERVING

Overlap perception and generation to reduce response delay.

Incremental processing starts before the user’s turn has ended.

Audio / video inputStream and encode
Language modelPrefill → speculative decode
Endpoint detection600–700 ms
Response delivery<100 ms

The reported <100 ms is after endpoint detection, not total end-to-end latency. Timeline is schematic.

LONGCAT-FLASH-OMNI / EVIDENCE

Evaluate joint audio–visual understanding.

Team-reported system results; selected cross-modality benchmarks.

Instruct modelOmniBench ↑WorldSense ↑DailyOmni ↑
Qwen3-Omni58.4152.0169.33
Gemini 2.5 Pro · budget 12866.8063.9680.61
LongCat-Flash-Omni61.3860.8982.38

OmniBench uses the team’s corrected scoring version. Model sizes and training differ.

LONGCAT-FLASH-OMNI / CONTRIBUTION

Participation in a large-scale omni-modal system.

This team project connects my audio research with multimodal interaction.

560B / 27BTotal / activated model parameters
OPEN PROJECT
Inspect the system

Technical report, released model, and serving implementation.

Official repository ↗

Related first-author work: Omni-AutoThink → adaptive reasoning.

Omni-AutoThink

Learn when multimodal reasoning is worth the effort

My first-author paper Omni-AutoThink studies how an omni model can decide when to reason. Adaptive SFT and Adaptive GRPO train the model to choose between a direct answer and a reasoning response across text, audio and visual inputs.

The work adds a decision-making dimension to multimodal capability: allocate reasoning where it helps. Its benchmark measures accuracy and thinking behavior across difficulty levels.

Explore the idea Problem · method · evidence · significance
RESEARCH EXPLAINED4 slides · explore at your pace

OMNI-AUTOTHINK / QUESTION

An omni model should learn when reasoning is useful.

My first-author work studies adaptive reasoning across text, audio, and vision.

Multimodal question
DIRECT ANSWERWhen the evidence is sufficient

Avoid unnecessary reasoning.

THINK, THEN ANSWERWhen inference is needed

Spend reasoning effort on the harder case.

The choice of response mode becomes learned behavior.

OMNI-AUTOTHINK / LEARNING

Reward correct direct answers without rewarding blind shortcuts.

Adaptive SFT initializes both modes; Adaptive GRPO samples them and learns from correctness.

CorrectIncorrectNo-think+2−1Think+10

Adaptive sampling/rejection changes which rollouts contribute to the update.

OMNI-AUTOTHINK / BEHAVIOR

Thinking increases with audio-question difficulty.

Text–audio benchmark · five calibrated difficulty levels.

Responses using thinking mode
L1 · easiest
17%
L2
37%
L3
60%
L4
69%
L5 · hardest
71%

Overall: 73% accuracy with 47% thinking rate.

Table 3 ↗ · Thinking rate measures mode selection, not token savings or latency.

OMNI-AUTOTHINK / ABLATION

SFT and RL together unlock adaptation across modalities.

Same base model · reported training-stage ablation · Pass@1 accuracy.

TrainingTextAudio + textAudio + vision + text
Base35%64%48%
+ SFT49%64%50%
+ RL only60%66%50%
+ SFT + RL66%73%69%

The accompanying benchmark evaluates four modality groups over five difficulty levels.

RESEARCH IN PRACTICE

From methods to a music system

HeartMuLa

Research methods become a working system

HeartMuLa is a concrete application of methods from my earlier research. ALMTokenizer’s query-based compression contributes to HeartCodec’s token interface; UniAudio’s Global–Local architecture supports hierarchical music modeling; and SQ-Codec from the SimpleSpeech line provides a latent space for flow-based reconstruction.

The case shows cumulative research: representation, modeling and reconstruction become complementary components of a team-built open music system.

Explore the idea Problem · method · evidence · significance
RESEARCH EXPLAINED4 slides · explore at your pace

HEARTMULA / APPLICATION

Earlier methods meet inside a song-generation system.

HeartMuLa is a concrete application of my representation and generation research.

ALMTokenizerquery compressionHeartCodec tokens
UniAudioGlobal–Local modelingHierarchical music generator
SimpleSpeechSQ-Codec latent spaceFlow-based reconstruction

HEARTMULA / SYSTEM

Compact token generation and waveform reconstruction work together.

Lyrics, tags, and reference audio condition the music language model.

Structured lyricsStyle tagsReference audio
SEQUENCE MODELGlobal–LocalPredict music tokens
DETOKENIZERFlow matchingRecover SQ-Codec latents
WAVEFORMSQ-Codec decoderRender the song

The representation connects long-range song modeling to detailed acoustic reconstruction.

HEARTMULA / EVIDENCE

A better decoder improves the same generated tokens.

Controlled decoder-stage comparison.

HeartCodec stageSongEval average ↑Phoneme error rate ↓
Pretrain + finetune4.310.1092
+ Reflow4.340.1258
+ SQ-Codec finetuning4.410.1005

Reflow shows a tradeoff; subsequent SQ-Codec finetuning improves both reported measures.

HEARTMULA / EXPLORE

Hear the system; trace the methods behind it.

Listen for sustained musical structure, vocal clarity, and acoustic detail.

SELECTED PUBLIC SAMPLE
Generated song

Illustration of the system; quantitative evidence is on the previous page.

OPEN SYSTEM
From paper to usable code

Code and models ↗

Project and listening examples ↗

CURRENT DIRECTION

Toward omni-modal interaction

My recent work extends audio modeling into multimodal systems, adaptive reasoning, and ongoing research on full-duplex interaction.

Full-duplex omni-modal interaction

I am currently working on full-duplex interaction models across modalities. This direction brings together audio generation, multimodal understanding and interaction over time.

The research goal: models that can keep perceiving while responding, and adjust as the interaction evolves.

PHD RESEARCH

Toward unified audio foundation models

My thesis studies three connected questions: how to represent audio for language modeling, how to unify audio tasks, and how to reconstruct high-fidelity sound from compact tokens.

It brings together the UniAudio series, LLM-Codec, ALMTokenizer, and scalar-latent generation. Earlier work in sound perception and language-controlled generation forms a broader research trajectory.

Read the research overview ↗

Blog

ESSAY · JUNE 2026 · EN / 中文

Audio Generative Modeling in Latent Space

Discrete tokens or continuous latents? A perspective on representation, modelability, and the paths toward audio generation.

↗

BACKGROUND

Education & Internships

Curriculum vitae ↗

Education

The Chinese University of Hong Kong

PhD in Systems Engineering & Engineering Management

Peking University

Master's degree

Shanghai University

Bachelor's degree

Research Internships

Microsoft Research Asia

Research Intern · Speech Group · Xu Tan

Tencent AI Lab

Research Intern · Speech Group

Honors & Awards

  • 2024
    IEEE SPS Young Author Best Paper AwardDiffsound
  • 2024
    ISCA Best Student Paper AwardSimpleSpeech · Interspeech
  • 2025
    ICLR Notable ReviewerNeurIPS Top Reviewer
  • 2023
    Outstanding GraduatePeking University
  • 2021
    DCASE Challenge Task 5 · 1st placeFew-shot bioacoustic event detection

Academic Service

Senior Program Committee (SPC) · AAAI

Reviewer for ICML, NeurIPS, ICLR, COLM, IJCAI, ACM MM, ICASSP, Interspeech, IEEE TASLP, and IEEE Signal Processing Letters.

LET'S CONNECT

Interested in audio intelligence?

I'm happy to discuss research ideas and collaborations.

dcyang@se.cuhk.edu.hk ↗