BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//Computer Science and Engineering - ECPv6.13.0//NONSGML v1.0//EN
CALSCALE:GREGORIAN
METHOD:PUBLISH
X-WR-CALNAME:Computer Science and Engineering
X-ORIGINAL-URL:https://homecse.iitd.ac.in
X-WR-CALDESC:Events for Computer Science and Engineering
REFRESH-INTERVAL;VALUE=DURATION:PT1H
X-Robots-Tag:noindex
X-PUBLISHED-TTL:PT1H
BEGIN:VTIMEZONE
TZID:Asia/Kolkata
BEGIN:STANDARD
TZOFFSETFROM:+0530
TZOFFSETTO:+0530
TZNAME:IST
DTSTART:20240101T000000
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTART;TZID=Asia/Kolkata:20241219T120000
DTEND;TZID=Asia/Kolkata:20241219T130000
DTSTAMP:20260924T090447
CREATED:20241126T043850Z
LAST-MODIFIED:20241212T122814Z
UID:212-1734609600-1734613200@homecse.iitd.ac.in
SUMMARY:Global Search and Discovery with Differential Policy Optimization
DESCRIPTION:Chandrajit Bajaj\, UT Austin \nReinforcement learning (RL) with continuous state and action spaces is arguably one the most challenging problems within the field of machine learning.  Most current learning methods focus on integral identities such as value (Q) functions to derive an optimal strategy for the learning agent. In this talk we present the dual form of the original RL formulation to propose the first differential RL framework that can handle settings with limited training samples and short-length episodes. Our approach introduces Differential Policy Optimization (DPO)\, a pointwise and stage-wise iteration method that optimizes policies encoded by local-movement operators. We prove a pointwise convergence estimate for DPO and provide a regret bound comparable with the best current theoretical derivation. Such pointwise estimate ensures that the learned policy matches the optimal path uniformly across different steps. We then apply DPO to a class of practical RL problems with continuous state and action spaces\,  e.g. shape and material optimization and discovery of new molecules with targeted dynamics. \nThis is joint work with Garvit Bansal\, Minh Nguyen.
URL:https://homecse.iitd.ac.in/event/prospecting-for-global-optimizers-with-physics-agents/
LOCATION:Bharti 501\, IIT Campus\, Hauz Khas\, New Delhi
CATEGORIES:Seminars
END:VEVENT
END:VCALENDAR