CMU · Language Technologies Institute
Zhuoyan Terry Tao

Graduate researcher in speech processing and audio machine learning at CMU WavLab. Drawn to problems where signal meets meaning. ✌︎

Terry at Kualoa Ranch, Oahu
Kualoa Ranch · Oahu, Hawaii
01

About Me

Speech Processing · Audio ML · NLP

I'm Zhuoyan Tao, Terry to friends and colleagues. I'm a master's student in CMU's Intelligent Information Systems program at the Language Technologies Institute, advised by Prof. Shinji Watanabe at WavLab.

My research spans speech quality evaluation, spoken language diarization, and speech-to-speech translation. It sits at the boundary of signal processing, machine learning, and linguistics, and keeps circling one question: how meaning is carried in speech beyond just the words.

Before CMU I finished dual B.S. degrees in Computer Science and Applied Mathematics at USC, and I spent summer 2026 at Apple building simulation tooling for satellite connectivity. I also contribute to ESPnet. Outside of research, I enjoy exploring new places ☀︎

Education CMU · M.S. Intelligent Information Systems Language Technologies Institute, expected Dec 2027
Previously USC · B.S. CS & Applied Math Magna Cum Laude, May 2026
Honors Phi Beta Kappa · Trustee Scholar Full-tuition merit scholarship, USC · NAE Grand Challenges Scholar
Advisor Prof. Shinji Watanabe CMU WavLab
Publishes as Zhuoyan Tao
02

Research

CMU WavLab · Interspeech 2026 Oral
ANCHOR: Speech Quality from Partial Audio
Extends ARECHO to predict speech quality before an utterance finishes. Dual-resolution query tokens with chunk-first decoding in a shared Transformer decoder cut PLCMOS MAE by 48% on 2-second prefixes and locate a 4 to 6 second perceptual context horizon.
USC SAIL Lab · EMNLP under review
Diagnosing Spoken Language Diarization
Introduces off-target LID posterior mass as a measure that predicts probe gains on pre-trained speech encoders. Pre-registered on unseen language pairs and validated across Whisper, MMS, and classical LID on DISPLACE and MUCS Hindi-English code-switching data.
CMU WavLab · Interspeech 2026 S2ST Challenge
Pragmatic Intent in Speech Translation
HuBERT- and prosody-based speech-to-speech translation systems that preserve stance, emotion, and dialog function across English and Spanish, so the translated speech keeps what the speaker meant, not only what they said.
CMU WavLab · IEEE SLT 2026 under review
TurnBench: Turn-Taking in Spoken Dialogue
A multi-domain benchmark for turn-taking dynamics in spoken dialogue, measuring how well models anticipate when a speaker yields the floor across conversational settings.
03

Publications

001
ANCHOR: Autoregressive Non-intrusive Chunk-Ordered Refinement for Joint Multi-Resolution Speech Quality Modeling
Z. Tao, J. Shi, H. Shim, S. Watanabe
Interspeech 2026 · Oral First author arXiv:2606.10233 ↗︎
002
Off-Target LID Posterior Mass Predicts Probe Gains in Spoken Language Diarization: A Pre-Registered Validation
Z. Tao, A. Kommineni, S. Narayanan
EMNLP · Under review First author
003
TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue
F. Jiang, R. Sanabria, S. Deshmukh, …, Z. Tao et al.
IEEE SLT 2026 · Under review arXiv:2608.25218 ↗︎
04

Experience

05

Links & Contact