Audio & multimodal AI · RL & post-training
Utkarsh Tyagi
I’m a Machine Learning Research Engineer at Scale AI in San Francisco. My research focuses on audio and multimodal AI, with an emphasis on post-training, reinforcement learning, and speech-to-speech models.
I work on steerable speech dialogue, rubric-based learning, and evaluations of how models reason across audio, images, and video.
I’m particularly interested in designing better learning signals and building models that remain reliable across natural, multi-turn interactions.
Previously, I completed my M.S. in Computer Science at the University of Maryland, where I worked with GAMMA Lab, Prof. Dinesh Manocha, and Prof. Ramani Duraiswami. I’ve also worked at Atlassian and Samsung.
- Speech & audio
- Multimodal reasoning
- RL & post-training
San Francisco, California
Updates
2021–2026New paper: Humanity’s Sixth Sense, a benchmark for intuitive visual reasoning in multimodal models.
POW3R will appear at NeurIPS 2026. SteerDuplex will appear at the NeurIPS 2026 Workshop RTCA.
Beyond Seeing will appear at EMNLP 2026. Our new Rubric Dropout paper is available.
Audio MultiChallenge received the ACL 2026 Best Resource Paper Award, shared with my coauthors.
Audio Hallucination Attacks will appear at Interspeech 2026. Our new Rubric-Guided Self-Distillation paper is available.
SciPredict will appear at ICML 2026, and Audio MultiChallenge at ACL 2026.
Granted U.S. Patent 12,505,292, on scheduling abstractive summaries for multi-party communication channels.
MULTIVOX and EGOILLUSION will appear at EMNLP 2025.
Joined Scale AI as a Machine Learning Research Engineer, working on audio and multimodal AI.
Completed my M.S. in Computer Science at the University of Maryland, College Park.
MMAU was selected for a Spotlight at ICLR 2025.
Awarded a $7,000 NeuroPAC Fellowship at UMD/UMIACS for research on EEG-based brain–computer interaction using LLMs.
One paper accepted to Interspeech 2024, two to ACL 2024, and two to NAACL 2024.
CompA accepted to ICLR 2024.
Attended EMNLP in person in Singapore.
Two papers accepted to EMNLP 2023.
Started my M.S. in Computer Science at the University of Maryland, College Park.
One paper accepted to ICCV 2023.
One paper accepted to Interspeech 2023.
One paper accepted to ACL 2023 and one to SIGIR 2023.
Our work on implicit hate speech in online conversations was accepted to the De-Factify 2 workshop at AAAI 2023.
Filed two patents with Atlassian on incident management in multi-party communication channels. Promoted to Software Engineer 2 and began collaborating with GAMMA Lab at the University of Maryland.
Joined MIDAS Lab at IIIT Delhi.
Joined Atlassian as a Software Engineer.