<node id="662098">
  <nid>662098</nid>
  <type>event</type>
  <uid>
    <user id="27707"><![CDATA[27707]]></user>
  </uid>
  <created>1665677633</created>
  <changed>1665677633</changed>
  <title><![CDATA[PhD Defense by Nirbhay Modhe]]></title>
  <body><![CDATA[<p><strong>Title</strong>: Leveraging Value-awareness for Online and Offline Model-based Reinforcement Learning</p>

<p>Date: Thursday, October 27th, 2022</p>

<p>Time: 9:00 AM - 11:00 AM Eastern Time</p>

<p>Location (virtual): <a href="https://bluejeans.com/264974579/4014">https://bluejeans.com/264974579/4014</a></p>

<p>&nbsp;</p>

<p><strong>Nirbhay Modhe</strong></p>

<p>Ph.D. Candidate</p>

<p>School of Interactive Computing</p>

<p>College of Computing</p>

<p>Georgia Institute of Technology</p>

<p>&nbsp;</p>

<p><strong>Committee</strong></p>

<p>Dr. Dhruv Batra (advisor), School of Interactive Computing, Georgia Institute of Technology</p>

<p>Dr. Zsolt Kira, School of Interactive Computing, Georgia Institute of Technology</p>

<p>Dr. Mark Riedl, School of Interactive Computing, Georgia Institute of Technology</p>

<p>Dr. Gaurav Sukhatme, University of Southern California</p>

<p>Dr. Ashwin Kalyan, Allen Institute for AI (AI2)</p>

<p>&nbsp;</p>

<p><strong>Summary</strong></p>

<p>Model-based Reinforcement Learning (RL) lies at the intersection of planning and learning for sequential decision making. Value-awareness in model learning has recently emerged as a means to imbue task or reward information into the objective of model learning, in order for the model to leverage specificity of a task. While finding success in theory as being superior to maximum likelihood estimation in the context of (online) model-based RL, value-awareness has remained impractical for most non-trivial tasks.</p>

<p>&nbsp;</p>

<p>This thesis aims to bridge the gap in theory and practice by applying the principle of value-awareness to two settings -- the online RL setting and offline RL setting. First, within online RL, this thesis revisits value-aware model learning from the perspective of minimizing performance difference, obtaining a novel value-aware model learning objective as a direct upper bound of it. Then, this thesis investigates and remedies the issue of stale value estimates that has so far been holding back the practicality of value-aware model learning. Using the proposed remedy, performance improvements are presented over maximum-likelihood based baselines and existing value-aware objectives, in several continuous control tasks, while also enabling existing value-aware objectives to become performant.</p>

<p>&nbsp;</p>

<p>In the offline RL setting, this thesis takes a step back from model learning and applies value-awareness towards better data augmentation. Such data augmentation, when applied to model-based offline RL algorithms, allows for leveraging unseen states with low epistemic uncertainty that have previously not been reachable within the assumptions and limitations of model-based offline RL. Value-aware state augmentations are found to enable better performance on offline RL benchmarks compared to existing baselines and non-value-aware alternatives.</p>
]]></body>
  <field_summary_sentence>
    <item>
      <value><![CDATA[Leveraging Value-awareness for Online and Offline Model-based Reinforcement Learning]]></value>
    </item>
  </field_summary_sentence>
  <field_summary>
    <item>
      <value><![CDATA[]]></value>
    </item>
  </field_summary>
  <field_time>
    <item>
      <value><![CDATA[2022-10-27T10:00:00-04:00]]></value>
      <value2><![CDATA[2022-10-27T12:00:00-04:00]]></value2>
      <rrule><![CDATA[]]></rrule>
      <timezone><![CDATA[America/New_York]]></timezone>
    </item>
  </field_time>
  <field_fee>
    <item>
      <value><![CDATA[]]></value>
    </item>
  </field_fee>
  <field_extras>
      </field_extras>
  <field_audience>
          <item>
        <value><![CDATA[Faculty/Staff]]></value>
      </item>
          <item>
        <value><![CDATA[Public]]></value>
      </item>
          <item>
        <value><![CDATA[Undergraduate students]]></value>
      </item>
      </field_audience>
  <field_media>
      </field_media>
  <field_contact>
    <item>
      <value><![CDATA[]]></value>
    </item>
  </field_contact>
  <field_location>
    <item>
      <value><![CDATA[]]></value>
    </item>
  </field_location>
  <field_sidebar>
    <item>
      <value><![CDATA[]]></value>
    </item>
  </field_sidebar>
  <field_phone>
    <item>
      <value><![CDATA[]]></value>
    </item>
  </field_phone>
  <field_url>
    <item>
      <url><![CDATA[https://bluejeans.com/264974579/4014]]></url>
      <title><![CDATA[BlueJeans]]></title>
            <attributes><![CDATA[]]></attributes>
    </item>
  </field_url>
  <field_email>
    <item>
      <email><![CDATA[]]></email>
    </item>
  </field_email>
  <field_boilerplate>
    <item>
      <nid><![CDATA[]]></nid>
    </item>
  </field_boilerplate>
  <links_related>
      </links_related>
  <files>
      </files>
  <og_groups>
          <item>221981</item>
      </og_groups>
  <og_groups_both>
          <item><![CDATA[Graduate Studies]]></item>
      </og_groups_both>
  <field_categories>
          <item>
        <tid>1788</tid>
        <value><![CDATA[Other/Miscellaneous]]></value>
      </item>
      </field_categories>
  <field_keywords>
          <item>
        <tid>100811</tid>
        <value><![CDATA[Phd Defense]]></value>
      </item>
      </field_keywords>
  <userdata><![CDATA[]]></userdata>
</node>
