Self-Supervised Adversarial Imitation Learning

Juarez Monteiro; Nathan Gavenski; Felipe Meneguzzi; Rodrigo C. Barros

doi:10.48550/arXiv.2304.10914

Self-Supervised Adversarial Imitation Learning

Juarez Monteiro, Nathan Gavenski, Felipe Meneguzzi, Rodrigo C. Barros

Computing Science

Research output: Working paper › Preprint

Abstract

Behavioural cloning is an imitation learning technique that teaches an agent how to behave via expert demonstrations. Recent approaches use self-supervision of fully-observable unlabelled snapshots of the states to decode state pairs into actions. However, the iterative learning scheme employed by these techniques is prone to get trapped into bad local minima. Previous work uses goal-aware strategies to solve this issue. However, this requires manual intervention to verify whether an agent has reached its goal. We address this limitation by incorporating a discriminator into the original framework, offering two key advantages and directly solving a learning problem previous work had. First, it disposes of the manual intervention requirement. Second, it helps in learning by guiding function approximation based on the state transition of the expert's trajectories. Third, the discriminator solves a learning issue commonly present in the policy model, which is to sometimes perform a `no action' within the environment until the agent finally halts.

Original language	English
Publisher	ArXiv
Number of pages	8
DOIs	https://doi.org/10.48550/arXiv.2304.10914
Publication status	Published - 21 Apr 2023

Bibliographical note

This paper has been accepted in the International Joint Conference on Neural Networks (IJCNN) 2023
This work was supported by UK Research and Innovation [grant number EP/S023356/1], in the UKRI Centre for Doctoral Training in Safe and Trusted Artificial Intelligence (www.safeandtrustedai.org) and made possible via King’s
Computational Research, Engineering and Technology Environment (CREATE) [27].

Keywords

cs.LG
cs.AI

Access to Document

10.48550/arXiv.2304.10914Licence: CC BY

Self-Supervised Adversarial Imitation LearningSubmitted manuscript, 396 KB

Cite this

@techreport{2940bcc42d724bb1914d0bbb7a573b8f,

title = "Self-Supervised Adversarial Imitation Learning",

abstract = "Behavioural cloning is an imitation learning technique that teaches an agent how to behave via expert demonstrations. Recent approaches use self-supervision of fully-observable unlabelled snapshots of the states to decode state pairs into actions. However, the iterative learning scheme employed by these techniques is prone to get trapped into bad local minima. Previous work uses goal-aware strategies to solve this issue. However, this requires manual intervention to verify whether an agent has reached its goal. We address this limitation by incorporating a discriminator into the original framework, offering two key advantages and directly solving a learning problem previous work had. First, it disposes of the manual intervention requirement. Second, it helps in learning by guiding function approximation based on the state transition of the expert's trajectories. Third, the discriminator solves a learning issue commonly present in the policy model, which is to sometimes perform a `no action' within the environment until the agent finally halts.",

keywords = "cs.LG, cs.AI",

author = "Juarez Monteiro and Nathan Gavenski and Felipe Meneguzzi and Barros, {Rodrigo C.}",

note = "This paper has been accepted in the International Joint Conference on Neural Networks (IJCNN) 2023 This work was supported by UK Research and Innovation [grant number EP/S023356/1], in the UKRI Centre for Doctoral Training in Safe and Trusted Artificial Intelligence (www.safeandtrustedai.org) and made possible via King{\textquoteright}s Computational Research, Engineering and Technology Environment (CREATE) [27].",

year = "2023",

month = apr,

day = "21",

doi = "10.48550/arXiv.2304.10914",

language = "English",

publisher = "ArXiv",

type = "WorkingPaper",

institution = "ArXiv",

}

TY - UNPB

T1 - Self-Supervised Adversarial Imitation Learning

AU - Monteiro, Juarez

AU - Gavenski, Nathan

AU - Meneguzzi, Felipe

AU - Barros, Rodrigo C.

N1 - This paper has been accepted in the International Joint Conference on Neural Networks (IJCNN) 2023 This work was supported by UK Research and Innovation [grant number EP/S023356/1], in the UKRI Centre for Doctoral Training in Safe and Trusted Artificial Intelligence (www.safeandtrustedai.org) and made possible via King’s Computational Research, Engineering and Technology Environment (CREATE) [27].

PY - 2023/4/21

Y1 - 2023/4/21

N2 - Behavioural cloning is an imitation learning technique that teaches an agent how to behave via expert demonstrations. Recent approaches use self-supervision of fully-observable unlabelled snapshots of the states to decode state pairs into actions. However, the iterative learning scheme employed by these techniques is prone to get trapped into bad local minima. Previous work uses goal-aware strategies to solve this issue. However, this requires manual intervention to verify whether an agent has reached its goal. We address this limitation by incorporating a discriminator into the original framework, offering two key advantages and directly solving a learning problem previous work had. First, it disposes of the manual intervention requirement. Second, it helps in learning by guiding function approximation based on the state transition of the expert's trajectories. Third, the discriminator solves a learning issue commonly present in the policy model, which is to sometimes perform a `no action' within the environment until the agent finally halts.

AB - Behavioural cloning is an imitation learning technique that teaches an agent how to behave via expert demonstrations. Recent approaches use self-supervision of fully-observable unlabelled snapshots of the states to decode state pairs into actions. However, the iterative learning scheme employed by these techniques is prone to get trapped into bad local minima. Previous work uses goal-aware strategies to solve this issue. However, this requires manual intervention to verify whether an agent has reached its goal. We address this limitation by incorporating a discriminator into the original framework, offering two key advantages and directly solving a learning problem previous work had. First, it disposes of the manual intervention requirement. Second, it helps in learning by guiding function approximation based on the state transition of the expert's trajectories. Third, the discriminator solves a learning issue commonly present in the policy model, which is to sometimes perform a `no action' within the environment until the agent finally halts.

KW - cs.LG

KW - cs.AI

U2 - 10.48550/arXiv.2304.10914

DO - 10.48550/arXiv.2304.10914

M3 - Preprint

BT - Self-Supervised Adversarial Imitation Learning

PB - ArXiv

ER -

Self-Supervised Adversarial Imitation Learning

Abstract

Bibliographical note

Keywords

Access to Document

Fingerprint

Cite this