Audio & multimodal AI · RL & post-training

Utkarsh Tyagi

I’m a Machine Learning Research Engineer at Scale AI in San Francisco. My research focuses on audio and multimodal AI, with an emphasis on post-training, reinforcement learning, and speech-to-speech models.

I work on steerable speech dialogue, rubric-based learning, and evaluations of how models reason across audio, images, and video.

I’m particularly interested in designing better learning signals and building models that remain reliable across natural, multi-turn interactions.

Previously, I completed my M.S. in Computer Science at the University of Maryland, where I worked with GAMMA Lab, Prof. Dinesh Manocha, and Prof. Ramani Duraiswami. I’ve also worked at Atlassian and Samsung.

  • Speech & audio
  • Multimodal reasoning
  • RL & post-training
Utkarsh Tyagi

San Francisco, California

Updates

2021–2026
  • New paper: Humanity’s Sixth Sense, a benchmark for intuitive visual reasoning in multimodal models.

  • POW3R will appear at NeurIPS 2026. SteerDuplex will appear at the NeurIPS 2026 Workshop RTCA.

  • Beyond Seeing will appear at EMNLP 2026. Our new Rubric Dropout paper is available.

  • Audio MultiChallenge received the ACL 2026 Best Resource Paper Award, shared with my coauthors.

  • Audio Hallucination Attacks will appear at Interspeech 2026. Our new Rubric-Guided Self-Distillation paper is available.

  • SciPredict will appear at ICML 2026, and Audio MultiChallenge at ACL 2026.

  • Granted U.S. Patent 12,505,292, on scheduling abstractive summaries for multi-party communication channels.

  • MULTIVOX and EGOILLUSION will appear at EMNLP 2025.

  • Joined Scale AI as a Machine Learning Research Engineer, working on audio and multimodal AI.

  • Completed my M.S. in Computer Science at the University of Maryland, College Park.

  • MMAU was selected for a Spotlight at ICLR 2025.

  • MMAU and VDGD will appear at ICLR 2025.

  • Awarded a $7,000 NeuroPAC Fellowship at UMD/UMIACS for research on EEG-based brain–computer interaction using LLMs.

  • One paper accepted to Interspeech 2024, two to ACL 2024, and two to NAACL 2024.

  • CompA accepted to ICLR 2024.

  • Attended EMNLP in person in Singapore.

  • Two papers accepted to EMNLP 2023.

  • Started my M.S. in Computer Science at the University of Maryland, College Park.

  • One paper accepted to ICCV 2023.

  • One paper accepted to Interspeech 2023.

  • One paper accepted to ACL 2023 and one to SIGIR 2023.

  • Our work on implicit hate speech in online conversations was accepted to the De-Factify 2 workshop at AAAI 2023.

  • Filed two patents with Atlassian on incident management in multi-party communication channels. Promoted to Software Engineer 2 and began collaborating with GAMMA Lab at the University of Maryland.

  • Joined MIDAS Lab at IIIT Delhi.

  • Joined Atlassian as a Software Engineer.