Commit 6358d9f3 authored by Guilhem Saurel's avatar Guilhem Saurel
Browse files

amit_icra_22: fix details

parent cd5b7dc8
Pipeline #17636 passed with stage
in 8 seconds
......@@ -2,7 +2,7 @@
title: "Value learning from trajectory optimization and Sobolev descent: A step toward reinforcement learning with superlinear convergence properties"
subtitle: IEEE ICRA - International Conference on Robotics and Automation, 2022
author:
- Amit Parag ^1, 2^
- Amit Parag ^1,2^
- Sébastien Kleff ^1,3^
- Léo Saci ^1^
- <a href="https://gepettoweb.laas.fr/index.php/Members/NicolasMansard">Nicolas Mansard</a> ^1,2^
......@@ -13,8 +13,8 @@ org:
- ^3^ New York University, USA
hal: https://hal.archives-ouvertes.fr/hal-03356261
peertube: https://peertube.laas.fr/videos/embed/9a3c5258-e5b7-49a5-a153-02e804a06f65
sourcecode: https://github.com/amitparag/Kuka-arm-DvP
peertube: https://peertube.laas.fr/videos/embed/00e812a0-7321-4f05-8ab8-3250f2a49deb
code: https://github.com/amitparag/Kuka-arm-DvP
...
## Abstract
......@@ -22,7 +22,7 @@ sourcecode: https://github.com/amitparag/Kuka-arm-DvP
The recent successes in deep reinforcement learning largely rely on the capabilities of generating masses of data, which in turn implies the use of a simulator.
In particular, current progress in multi body dynamic simulators are underpinning the implementation of reinforcement learning for end-to-end control of robotic systems.
Yet simulators are mostly considered as black boxes while we have the knowledge to make them produce a richer information.
In this paper, we are proposing to use the derivatives of the simulator to help with the convergence of the learning.
In this paper, we are proposing to use the derivatives of the simulator to help with the convergence of the learning.
For that, we combine model-based trajectory optimization to produce informative trials using 1st- and 2nd-order simulation derivatives.
These locally-optimal runs give fair estimates of the value function and its derivatives, that we use to accelerate the convergence of the critics using Sobolev learning.
We empirically demonstrate that the algorithm leads to a faster and more accurate estimation of the value function.
......
Supports Markdown
0% or .
You are about to add 0 people to the discussion. Proceed with caution.
Finish editing this message first!
Please register or to comment