BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//Computer Science and Engineering - ECPv6.13.0//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-ORIGINAL-URL:https://homecse.iitd.ac.in
X-WR-CALDESC:Events for Computer Science and Engineering
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:Asia/Kolkata
BEGIN:STANDARD
TZOFFSETFROM:+0530
TZOFFSETTO:+0530
TZNAME:IST
DTSTART:20260101T000000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;TZID=Asia/Kolkata:20260413T120000
DTEND;TZID=Asia/Kolkata:20260413T130000
DTSTAMP:20261010T130701
CREATED:20260407T073939Z
LAST-MODIFIED:20260407T073939Z
UID:2483-1776081600-1776085200@homecse.iitd.ac.in
SUMMARY:Advancing Safe Multimodal Intelligence by Dr. Pritam Sarkar
DESCRIPTION:Venue: Bharti-501/ MS Teams \nAbstract: Building artificial intelligence with meaningful real-world impact requires models that can understand and interact in both the virtual and physical world. This talk outlines four key capabilities necessary to achieve this: foundational world knowledge\, alignment with human values and expectations\, reasoning ability\, and the capacity to self-improve. The talk begins with an overview of our past contributions toward these capabilities\, with a particular focus on building foundational world knowledge in multimodal models and improving their alignment with human values and expectations. The next part of the talk delves into a fundamental challenge in current alignment approaches: the alignment tax. While existing methods aim to make models safer and more reliable\, they often degrade general capabilities—reducing response diversity and making models overly cautious or less useful. To address this\, we introduce Refined Regularized Preference Optimization (RRPO)\, a fine-grained alignment method. By penalizing only specific error tokens rather than entire responses\, RRPO mitigates harmful behaviors while simultaneously improving performance across diverse vision tasks. Finally\, the talk concludes with an overview of ongoing work on enabling stronger visual reasoning and outlines a research vision for developing machines that can adapt and improve over time\, fostering long-term usefulness and human–AI trust. \n\n \n \nBio: Pritam Sarkar is a Distinguished Postdoctoral Fellow at the Vector Institute and a Postdoctoral Research Fellow at the University of British Columbia\, where he works on multimodal AI\, especially with video\, image\, audio\, and language. He completed his PhD in September 2025 at Queen’s University\, Canada and during this time\, he was an intern at Google\, USA. His research has been recognized at leading venues including NeurIPS\, ICLR\, and AAAI\, with multiple Oral and Spotlight presentations. He received the IEEE Research Excellence Award in 2023 for his work on self-supervised learning. He actively serves the research community as an Area Chair and a Reviewer for leading conferences and journals such as NeurIPS\, CVPR\, and PAMI\, and is a strong proponent of open-source research. He is interested in developing safe and generalizable multimodal intelligence through algorithms that learn effectively with minimal human supervision. Find more: https://pritamsarkar.com/  \nHead Shot: https://pritamsarkar.com/assets/my_images/pp_square.jpg 
URL:https://homecse.iitd.ac.in/event/advancing-safe-multimodal-intelligence-by-dr-pritam-sarkar/
LOCATION:Bharti 501\, IIT Campus\, Hauz Khas\, New Delhi
END:VEVENT
END:VCALENDAR