CMU · Language Technologies Institute

Zhuoyan Terry Tao

Graduate researcher in speech processing and audio machine learning at CMU WavLab. Drawn to problems where signal meets meaning. ✌︎

Terry at Kualoa Ranch, Oahu
Kualoa Ranch · Oahu, Hawaii
01

About

Speech Processing · Audio ML · NLP

I'm Zhuoyan Tao, Terry to friends and colleagues. I'm a master's student in CMU's Intelligent Information Systems program at the Language Technologies Institute, advised by Prof. Shinji Watanabe at WavLab.

My work sits at the boundary of signal processing, machine learning, and linguistics, and keeps circling one question: how meaning is carried in speech beyond just the words.

Before CMU I finished dual B.S. degrees at USC, and spent summer 2026 at Apple building simulation tooling for satellite connectivity. Outside of research, I enjoy exploring new places ☀︎

Education Carnegie Mellon University M.S. Intelligent Information Systems Language Technologies Institute Expected Dec 2027 University of Southern California B.S. Computer Science B.S. Applied and Computational Mathematics Magna Cum Laude · May 2026
Advisors Prof. Shinji Watanabe CMU WavLab · 2025–Present Prof. Shrikanth Narayanan USC SAIL · 2023–2026
Honors Phi Beta Kappa Elected 2025 Trustee Scholar Full-tuition merit scholarship, USC Grand Challenges Scholar NAE · USC Viterbi, 2026
Toolkit Speech & ML PyTorch · torchaudio · Hugging Face Transformers · ESPnet · VERSA Models Whisper · MMS · WavLM · HuBERT Systems Python · C++ · Go · Bash · Docker · Linux · SLURM · Git Compute Distributed training (DDP) · profiling & vectorization · NCSA Delta, PSC Bridges-2 GPU clusters
Publishes as Zhuoyan Tao
02

Research

CMU WavLab

Perceptual Speech Quality

Non-intrusive models that predict how a listener will rate speech from partial audio. The open question is how much acoustic context a quality judgment needs, and how to stay accurate while that context is still arriving.

USC SAIL Lab

Multilingual & Code-Switched Speech

Spoken language diarization for conversational, code-switching speech. I study what pre-trained speech encoders already encode about language identity, and when a lightweight probe is enough to recover it.

CMU WavLab

Prosody, Stance & Pragmatic Intent

Speech-to-speech translation that carries pragmatic intent across languages, so the output keeps what the speaker meant and not only what they said.

Open Source · ESPnet

Reproducible Speech Tooling

Merged contributions to ESPnet (16k+ stars), including a ≈50× speedup for speaker-embedding extraction and new TTS recipes. Benchmarks and evaluation code ship with the research so other groups can build on it.

03

Publications

Author names in order; my name in bold.

Peer-reviewed

001

ANCHOR: Autoregressive Non-intrusive Chunk-Ordered Refinement for Joint Multi-Resolution Speech Quality Modeling

Z. Tao, J. Shi, H. Shim, S. Watanabe

Extends ARECHO (NeurIPS 2025 Spotlight) to predict speech quality before an utterance finishes. Dual-resolution query tokens with chunk-first decoding in a shared Transformer decoder cut PLCMOS MAE by 48% on 2-second prefixes and locate a 4 to 6 second perceptual context horizon.

Interspeech 2026 · Oral First author arXiv:2606.10233 ↗︎

Under review

002

Off-Target LID Posterior Mass Predicts Probe Gains in Spoken Language Diarization: A Pre-Registered Validation

Z. Tao, A. Kommineni, S. Narayanan

Introduces off-target LID posterior mass as a measure that predicts probe gains on pre-trained speech encoders. Pre-registered on unseen language pairs and validated across Whisper, MMS, and classical LID on DISPLACE and MUCS Hindi-English code-switching data.

EMNLP · Under review First author
003

TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue

F. Jiang, R. Sanabria, S. Deshmukh, …, Z. Tao et al.

A multi-domain benchmark for turn-taking dynamics in spoken dialogue, measuring how well models anticipate when a speaker yields the floor across conversational settings.

IEEE SLT 2026 · Under review arXiv:2608.25218 ↗︎

In progress

004

Pragmatic Intent Preservation in Speech-to-Speech Translation

HuBERT- and prosody-based speech-to-speech translation systems that preserve stance, emotion, and dialog function across English and Spanish.

Interspeech 2026 S2ST Challenge