ICLR 2027 Workshop on Machine Learning for Audio

Previous Workshops

Learning to Listen: ICML 2026 Workshop on Machine Learning for Audio

AI Heard That! ICML 2025 Workshop on Machine Learning for Audio

For questions, email mlforaudioworkshop@gmail.com

Workshop Description

Machine learning research for audio applications has seen heightened interest in recent years. Key drivers include the emergence of audio language models and multimodal/foundation models that can understand and generate speech, music, and audio events alongside rapidly increasing demand for low-latency, production-ready voice agents and real-time transcription.

These advances broaden the capabilities of audio systems, but also raise new challenges in data scale and evaluation, real-time and on-device constraints, and responsible use. These emerging challenges, alongside the success of previous iterations of the Machine Learning for Audio workshop at various venues, have inspired us to bring this workshop to ICLR for the first time in 2027. We believe ICLR is especially well suited to this moment, as a growing number of open problems in audio center on questions of representation learning. Neural audio codecs and discrete tokenizers now determine what generative and audio-language models can express; self-supervised encoders shape performance across speech, music, and environmental sound; and aligning audio with text and vision representations is key to building capable multimodal systems. The workshop will serve as a dedicated forum for sharing tools and benchmarks, forming new collaborations, and holding timely discussion of the ethical implications of generative audio and audio foundation models.

The Machine Learning for Audio workshop at ICLR 2027 will cover a range of tasks and challenges involving audio data, including speech, music, and environmental sound. Topics include, but are not limited to: audio representation learning (e.g. self-supervised encoders, neural audio codecs, and discrete tokenizers), audio-language and multimodal models, generative modeling/synthesis (e.g. text-to-speech, raw music and ambient sound generation), denoising and enhancement, data augmentation and audio datasets, acoustic event classification, transcription, source separation, efficient and on-device audio models, and evaluation of audio systems.

Call for Papers

We will solicit two types of submissions through OpenReview: original extended abstracts (up to 4 pages), and a Tiny Papers track (up to 2 pages) for late-breaking results, simple but unpublished ideas, modest self-contained theoretical results, follow-up experiments or re-analyses of published work, and fresh perspectives on existing research that align with the workshop topic. The workshop will be non-archival, and we will explicitly encourage the submission of novel and ongoing work rather than papers previously published at other machine learning venues.

Submissions will be reviewed by the organizers and an additional set of program committee members, with each reviewer assigned no more than three papers. Reviewers and organizers will not evaluate submissions from their own organization, and conflicts of interest will be managed through OpenReview. We will follow the ICLR 2027 Policies on Large Language Model Usage for authors and reviewers; AI may not serve as a primary author or reviewer, and, consistent with ICLR requirements, AI-generated papers will not be accepted in the Tiny Papers track. Authors will be encouraged to indicate at submission time whether they would like to present a live demo, and accepted demos will be showcased alongside their posters during the poster session.

Submission Portal TBA

Timeline

Data Release

Recognizing the scarcity of free, publicly available audio data, Modulate and Hume AI will contribute several audio datasets alongside the workshop, all of large scale for their respective domains, for use in contributed work. These datasets, accessible via Google Drive, will include acted speech (professionally acted scripts), spontaneous speech (streamer content), mimicked speech (short-form emotive recordings), and mimicked non-verbal speech. The organizers hope this will allow researchers from smaller research groups and academia to work with and validate findings on large, generalizable datasets. In previous iterations, multiple submissions utilized versions of provided data in their work, and corresponding white papers were posted.

Further details on available data described here.

Tentative Schedule

We plan for the workshop to be an 8-hour event. Below is an approximate timetable of the workshop schedule, subject to change. We have been careful to facilitate ample time for informal discussion during the coffee break, poster & demo session, and open conversation session, as well as time for audience participation during the panel discussion and Q&A sections following invited talks.

Time Activity Description
8:30 Invited Speakers 1 & 2 Two 25-minute talks by invited speakers and Q&A.
9:30 Contributed Talks 1–3 Three 15-minute contributed talks by selected submissions and Q&A.
10:30 Coffee Break
11:00 Invited Speakers 3 & 4 Two 25-minute talks by invited speakers and Q&A.
12:00 Lunch
1:00 Poster & Demo Session Poster session alongside live demos from selected submissions.
2:00 Invited Speakers 5 & 6 Two 25-minute talks by invited speakers and Q&A.
3:00 Contributed Talks 4–6 Three 15-minute contributed talks by selected submissions and Q&A.
4:00 Panel Discussion Panel of invited speakers, with a moderator facilitating discussion, including questions from the audience.
4:30 Wrap-up & Discussion A few minutes of closing remarks followed by informal conversation among workshop attendees.

Virtual Access

The workshop will be held in person, and all confirmed invited speakers will attend in person. For those unable to attend, talks and the panel will be recorded (with speaker consent) and shared via the ICLR virtual platform, accepted papers will be public on OpenReview, and posters, slides, and demo materials will remain available on the workshop website. Invited talk titles and accepted papers will be posted in advance so attendees can plan around content. Contributors unable to attend due to visa issues or other exceptional circumstances may share pre-recorded videos, and we will display their posters on their behalf.

Invited Speakers

We have curated a list of invited speakers from a wide variety of fields within the audio domain, listed below along with brief biographies. All confirmed invited speakers will be attending in-person.

Gašper Beguš is an Associate Professor of Linguistics at UC Berkeley, where he directs the Berkeley Speech and Computation Lab and develops interpretable deep learning models that learn spoken language from raw audio, including models of how infants acquire speech. He is also the Linguistics Lead at Project CETI, where he applies these methods to sperm whale communication, and he previously was an Assistant Professor at the University of Washington after receiving his PhD from Harvard.

Arushi Goel is a Senior Research Scientist at NVIDIA, where she works on audio-language models and is a co-author of the Audio Flamingo series of models for audio understanding and reasoning. She received her PhD from the University of Edinburgh, where her research focused on computer vision and vision-language learning.

Emily Mower Provost is a Professor and Associate Chair of EECS at University of Michigan, where she leads the CHAI Lab. She develops human-centered ML for modeling emotion and behavior from multimodal signals, with applications in mental health and assistive technology.

Roger (Xinyu) Ren is a Senior Applied Scientist at Amazon AGI Foundations where he owns native intelligence in full-duplex audio LLMs, enabling real-time voice models built-in reasoning and tool use. He previously led the speech-reasoning work that made Amazon Nova Omni state of the art on the MMAU benchmark and on-the-fly adapter-based speech recognition updates for Alexa. He received his master’s in Materials Engineering from McGill University, with research focused on computational fluid dynamics (CFD) and material design, and holds a second master’s in Data Science from the University of Rochester.

Julius O. Smith III is Professor Emeritus of Music and, by courtesy, Electrical Engineering at Stanford University’s Center for Computer Research in Music and Acoustics (CCRMA), and is widely known for pioneering digital waveguide synthesis and for his series of online books on audio signal processing. He is a Fellow of the Audio Engineering Society and the Acoustical Society of America, and previously led sound, music, and signal processing software development at NeXT Computer.

Sriram Srinivasan is Director of Engineering for Remote Presence Audio at Meta, leading teams building real-time audio communication technology across Instagram, WhatsApp, Messenger, and Facebook, including the Beryl echo removal system and the MLow low-bitrate codec. Previously, he led audio technology teams for Microsoft Teams and Skype, where he co-organized the first Deep Noise Suppression and Deep AEC challenges.

Organizers

Alice Baird is the VP of Research (Speech Intelligence) at Hume AI, NY, USA, where she currently works on modeling expressive human behaviors from speech. She earned her Ph.D. at the University of Augsburg, in 2022. Her work on emotion understanding from auditory, physiological, and multimodal data has been published extensively in leading journals and conferences in her field. She has co-organized several machine learning competitions including the ICML Expressive Vocalizations Workshop and NeurIPS and ICML Workshops on Machine Learning for Audio.

Chris Donahue is an assistant professor at Carnegie Mellon University and a research scientist at Google DeepMind. His research focuses on developing and responsibly deploying generative AI for music and creativity, thereby unlocking and augmenting human creative potential. His work involves (1) improving machine learning methods for controllable generative modeling for music, audio, and other sequential data, and (2) deploying real-world interactive systems that allow a broader audience—inclusive of non-musicians—to harness generative music AI through intuitive controls. He co-organized the Workshops on Machine Learning for Audio at ICML 2025 and 2026.

Brian Kulis is a professor at Boston University and former Amazon scholar who worked on Alexa. His research focuses broadly on machine learning, recently focused on applications to audio problems such as detection and generation. He has won two best paper awards at ICML and a best paper award at CVPR. He has previously organized two workshops at ICCV (in 2011 & 2013), two workshops at NeurIPS (in 2011 & 2023), and four workshops at ICML (in 2019, 2022, 2025, and 2026). He is regularly an area or senior area chair at major AI conferences, and has organized tutorials at ICML and ECCV.

David Liu is a PhD student in the Department of Computer Science at Boston University. His research interests are deep learning for audio, where he has recently been working on state-space models. He received his bachelor’s degree in computer science, data science, and mathematics from the University of Wisconsin - Madison in 2023. He co-organized the Workshops on Machine Learning for Audio at ICML 2025 and 2026.

Minje Kim is an Associate Professor in the Siebel School of Computing and Data Science at UIUC. Before that, he was at Indiana University and was an Amazon Scholar. As a researcher, he has focused on developing machine learning models for audio signal processing applications. He is a recipient of various awards, including the NSF Career Award, the Facilitating Learning Excellence Award by the Illinois Student Council, and the IEEE SPS Best Paper Award. He currently serves as Chair of the IEEE SPS Audio and Acoustic Signal Processing TC and as Senior Area Editor for IEEE TASPro and SPL. He has organized various academic events, such as IEEE WASPAA 2023 (General Co-Chair), and has co-organized various workshops, including HSCMA 2024, GenDA 2025, LRAC 2026, ICASSP 2023 Special Session on Neural Speech and Audio Coding, and Short Courses at ICASSP 2024.

Rachel Manzelli is a Machine Learning Tech Lead at Modulate, where she guides and contributes to the team building and maintaining models supporting Modulate’s products. She works on discriminative, generative and ensemble models to conduct nuanced audio-native conversational analysis, as well as real-time voice conversion. She co-organized the Machine Learning for Audio Synthesis workshop at ICML 2022 and the Workshop on Machine Learning for Audio at NeurIPS 2023, ICML 2025 and ICML 2026. She earned her bachelor’s degree in computer engineering from Boston University in 2019. During her undergraduate career, she conducted research in the areas of structured music generation and MIR.

Shrikanth (Shri) Narayanan is a University Professor and holder of the Niki and Max Nikias Chair in Engineering at the University of Southern California (USC). Shri is a Fellow of the National Academy of Inventors, ASA, ACM, IEEE, ISCA, APS, AAAS, AIMBE and the Association for the Advancement of Affective Computing (AAAC). Shri is a member of the US National Academy of Engineering, European Academy of Sciences and Arts and a 2022 Guggenheim Fellow.

Joe Wang is a Principal Applied Scientist at Amazon AGI where he specializes in realtime conversational models and responsible AI. Since 2025 he has been the RAI tech lead for all audio and speech generation for Amazon AGI foundation models including Nova Sonic and Nova 2 Sonic. He holds a Ph.D. in Electrical and Computer Engineering from Boston University.