TY - GEN
T1 - Self-Supervised Adversarial Imitation Learning
AU - Monteiro, Juarez
AU - Gavenski, Nathan
AU - Meneguzzi, Felipe
AU - Barros, Rodrigo C.
N1 - This work was supported by UK Research and Innovation [grant number EP/S023356/1], in the UKRI Centre for Doctoral Training in Safe and Trusted Artificial Intelligence (www.safeandtrustedai.org) and made possible via King’s Computational Research, Engineering and Technology Environment (CREATE) [27].
PY - 2023/8/2
Y1 - 2023/8/2
N2 - Behavioural cloning is an imitation learning technique that teaches an agent how to behave via expert demonstrations. Recent approaches use self-supervision of fully-observable unlabelled snapshots of the states to decode state pairs into actions. However, the iterative learning scheme employed by these techniques is prone to get trapped into bad local minima. Previous work uses goal-aware strategies to solve this issue. However, this requires manual intervention to verify whether an agent has reached its goal. We address this limitation by incorporating a discriminator into the original framework, offering two key advantages and directly solving a learning problem previous work had. First, it disposes of the manual intervention requirement. Second, it helps in learning by guiding function approximation based on the state transition of the expert's trajectories. Third, the discriminator solves a learning issue commonly present in the policy model, which is to sometimes perform a 'no action' within the environment until the agent finally halts.
AB - Behavioural cloning is an imitation learning technique that teaches an agent how to behave via expert demonstrations. Recent approaches use self-supervision of fully-observable unlabelled snapshots of the states to decode state pairs into actions. However, the iterative learning scheme employed by these techniques is prone to get trapped into bad local minima. Previous work uses goal-aware strategies to solve this issue. However, this requires manual intervention to verify whether an agent has reached its goal. We address this limitation by incorporating a discriminator into the original framework, offering two key advantages and directly solving a learning problem previous work had. First, it disposes of the manual intervention requirement. Second, it helps in learning by guiding function approximation based on the state transition of the expert's trajectories. Third, the discriminator solves a learning issue commonly present in the policy model, which is to sometimes perform a 'no action' within the environment until the agent finally halts.
KW - Adversarial Learning
KW - Imitation Learning
KW - Learning from Observation
KW - Self-Supervised Learning
UR - https://www.scopus.com/pages/publications/85169569844
U2 - 10.1109/IJCNN54540.2023.10191197
DO - 10.1109/IJCNN54540.2023.10191197
M3 - Published conference contribution
AN - SCOPUS:85169569844
SN - 978-1-6654-8868-6
T3 - International Joint Conference on Neural Networks (IJCNN)
BT - 2023 International Joint Conference on Neural Networks (IJCNN)
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2023 International Joint Conference on Neural Networks, IJCNN 2023
Y2 - 18 June 2023 through 23 June 2023
ER -