{"652409":{"#nid":"652409","#data":{"type":"event","title":"Center for Signal and Information Processing Seminar","body":[{"value":"\u003Cp\u003E\u003Cstrong\u003ETime\u0026nbsp;and Date:\u003C\/strong\u003E\u0026nbsp;Nov 5th (Fri.) - EST 15:00 to to 16:00 (PDT 12:00 to 13:00)\u003C\/p\u003E\r\n\r\n\u003Cp\u003E\u003Cstrong\u003EBlueJeans link:\u003C\/strong\u003E (\u003Ca href=\u0022http:\/\/bluejeans.com\/4658604304\u0022\u003Ebluejeans.com\/4658604304\u003C\/a\u003E)\u003C\/p\u003E\r\n\r\n\u003Cp\u003E\u003Cstrong\u003ESpeaker:\u003C\/strong\u003E\u0026nbsp;Shinji Watanabe,\u0026nbsp;Carnegie Mellon University (CMU)\u003C\/p\u003E\r\n\r\n\u003Cp\u003E\u003Cstrong\u003ETitle:\u0026nbsp;\u003C\/strong\u003EMulti-Speaker Conversation Recognition based on End-to-End Neural Networks\u003C\/p\u003E\r\n\r\n\u003Cp\u003E\u003Cstrong\u003EAbstract:\u003C\/strong\u003E\u0026nbsp;Recently, the end-to-end automatic speech recognition (ASR) paradigm has attracted great research interest as an alternative to the conventional hybrid framework of deep neural networks and hidden Markov models. This talk introduces extensions of the basic end-to-end architecture to tackle major problems faced by current ASR technologies in adverse environments including distant-talk and multi-speaker conditions. First, we propose a unified architecture to encompass microphone-array signal processing such as a state-of-the-art neural beamformer and dereverberation within the end-to-end framework. This architecture allows speech enhancement and ASR components to be jointly optimized to improve the ASR objective and leads to an end-to-end framework that works well in the distant-talk scenario. Next, we extend the framework to deal with multi-speaker ASR, where the system directly decodes multiple label sequences from a single speech sequence by unifying source separation and speech recognition functions in an end-to-end manner. Finally, we will introduce our open source activities, called ESPnet (\u003Ca href=\u0022https:\/\/github.com\/espnet\/espnet\u0022\u003Ehttps:\/\/github.com\/espnet\/espnet\u003C\/a\u003E), which can reproduce various speech processing experiments including the above example.\u003C\/p\u003E\r\n\r\n\u003Cp\u003E\u003Cstrong\u003EBio:\u003C\/strong\u003E\u0026nbsp;Shinji Watanabe is an Associate Professor at Carnegie Mellon University, Pittsburgh, PA. He received his B.S., M.S., and Ph.D. (Dr. Eng.) degrees from Waseda University, Tokyo, Japan. He was a research scientist at NTT Communication Science Laboratories, Kyoto, Japan, from 2001 to 2011, a visiting scholar in Georgia institute of technology, Atlanta, GA in 2009, and a senior principal research scientist at Mitsubishi Electric Research Laboratories (MERL), Cambridge, MA USA from 2012 to 2017. Prior to the move to Carnegie Mellon University, he was an associate research professor at Johns Hopkins University, Baltimore, MD USA from 2017 to 2020. His research interests include automatic speech recognition, speech enhancement, spoken language understanding, and machine learning for speech and language processing. He has been published more than 300 papers in peer-reviewed journals and conferences and received several awards, including the best paper award from the IEEE ASRU in 2019. He served as an Associate Editor of the IEEE Transactions on Audio Speech and Language Processing. He was\/has been a member of several technical committees, including the APSIPA Speech, Language, and Audio Technical Committee (SLA), IEEE Signal Processing Society Speech and Language Technical Committee (SLTC), and Machine Learning for Signal Processing Technical Committee (MLSP).\u003C\/p\u003E\r\n","summary":null,"format":"limited_html"}],"field_subtitle":"","field_summary":[{"value":"\u003Cp\u003EShinji Watanabe of Carnegie-Mellon University will deliver the November 5 CSIP Seminar, which is entitled \u0026quot;Multi-Speaker Conversation Recognition based on End-to-End Neural Networks.\u0026quot;\u003C\/p\u003E\r\n","format":"limited_html"}],"field_summary_sentence":[{"value":"Shinji Watanabe of Carnegie-Mellon University will deliver the November 5 CSIP Seminar, which is entitled \u0022Multi-Speaker Conversation Recognition based on End-to-End Neural Networks.\u0022"}],"uid":"27241","created_gmt":"2021-11-03 15:53:14","changed_gmt":"2021-11-03 15:53:32","author":"Jackie Nemeth","boilerplate_text":"","field_publication":"","field_article_url":"","field_event_time":{"event_time_start":"2021-11-05T16:00:00-04:00","event_time_end":"2021-11-05T17:00:00-04:00","event_time_end_last":"2021-11-05T17:00:00-04:00","gmt_time_start":"2021-11-05 20:00:00","gmt_time_end":"2021-11-05 21:00:00","gmt_time_end_last":"2021-11-05 21:00:00","rrule":null,"timezone":"America\/New_York"},"extras":[],"groups":[{"id":"1255","name":"School of Electrical and Computer Engineering"}],"categories":[],"keywords":[],"core_research_areas":[],"news_room_topics":[],"event_categories":[],"invited_audience":[{"id":"78761","name":"Faculty\/Staff"},{"id":"78771","name":"Public"},{"id":"78751","name":"Undergraduate students"}],"affiliations":[],"classification":[],"areas_of_expertise":[],"news_and_recent_appearances":[],"phone":[],"contact":[{"value":"\u003Cp\u003E\u003Ca href=\u0022mailto:raquel.plaskett@ece.gatech.edu\u0022\u003ERaquel Plaskett\u003C\/a\u003E\u003C\/p\u003E\r\n\r\n\u003Cp\u003E\u003Ca href=\u0022mailto:huckiyang@gatech.edu\u0022\u003EHuck Yang\u003C\/a\u003E\u003C\/p\u003E\r\n\r\n\u003Cp\u003ECenter for Signal and Information Processing\u003C\/p\u003E\r\n\r\n\u003Cp\u003ESchool of Electrical and Computer Engineering\u003C\/p\u003E\r\n","format":"limited_html"}],"email":[],"slides":[],"orientation":[],"userdata":""}}}