BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//Computer Science and Engineering - ECPv6.13.0//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-ORIGINAL-URL:https://homecse.iitd.ac.in
X-WR-CALDESC:Events for Computer Science and Engineering
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:Asia/Kolkata
BEGIN:STANDARD
TZOFFSETFROM:+0530
TZOFFSETTO:+0530
TZNAME:IST
DTSTART:20250101T000000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;TZID=Asia/Kolkata:20250605T100000
DTEND;TZID=Asia/Kolkata:20250605T110000
DTSTAMP:20260924T024013
CREATED:20250602T072215Z
LAST-MODIFIED:20250602T094001Z
UID:1633-1749117600-1749121200@homecse.iitd.ac.in
SUMMARY:Multimodal Learning in 3D environments: Perception\, and Simulation
DESCRIPTION:Speaker: Dr. Arun Balajee Vasudevan is currently a Research Scientist at Amazon \nAbstract: \nAutonomous robots have several potential applications such as virtual assistants\, VR/AR\, gaming\, self-driving technologies\, city planning and others. To achieve autonomy\, a robot needs to see and hear the environment\, before it can converse or navigate favorably to perform a human-desired task. Precisely\, scene understanding begins with building 3D geometry\, decoding semantics\, understanding surround objects/humans\, and planning and actions. My research talk addresses these fundamental challenges independently under two broad themes: understanding geometry and multimodal perception & data-driven simulation and navigation. \nUnder geometry and perception\, I introduce the usage of several multimodalities such as the user’s gaze\, visual sensors such as cameras or range sensors (e.g. Kinect)\, and speech/language instructions from human referrals for robot perception tasks. Further\, I delve deep into the investigation of the audio sensing modality for the task using binaural sound microphones. Secondly\, regarding the geometry\, my talk addresses one of my ongoing works about the construction of digital twins of the real world with 4D reconstruction of dynamic scenes from ground visuals of a robot. \nUnder the theme of Data-driven simulation and robot navigation. Following perception\, robots must navigate and take meaningful actions in the world. This involves broadly two aspects: wayfinding and motion planning. Earlier works address wayfinding based on directional instructions\, overlooking human aspects. I briefly talk about a new paradigm that integrates principles from cognitive science with learning-based methods to tackle the challenge of language-based wayfinding for robots in real-world outdoor environments. The second aspect is motion planning for which I propose the learning of driver behavior models for MPC-based planners to build data-driven simulators. \nLong term\, I envision to bridge the above two themes to build multimodal digital twins simulators of the real-world. This potentially helps in training and testing of planners\, VR/AR setups\, gaming\, and others. Lastly\, I also cover my future research plan in the talk. \nShort Bio: \nArun Balajee Vasudevan is currently a Research Scientist at Amazon. Previously\, he was a postdoctoral researcher at Carnegie Mellon University under Prof. Deva Ramanan till March 2025. His core research interest is in Computer Vision and Multimodal Learning. He has works in multimodal (vision\, language and sounds) perception and navigation\, 3D/4D reconstruction\, Motion Planning and improving Foundational models. He published papers predominantly in vision and machine learning conferences/journals such as CVPR\, ECCV\, ICML\, IJCV\, TPAMI\, and others. He defended his PhD under Prof. Luc Van Gool at ETH Zurich. He received his MSc in Computer Science from EPFL in 2016 and an undergraduate degree in Electrical Engineering from the Indian Institute of Technology Jodhpur in the year 2014.
URL:https://homecse.iitd.ac.in/event/multimodal-learning-in-3d-environments-perception-and-simulation/
CATEGORIES:Seminars
END:VEVENT
END:VCALENDAR