Optimal Stroke Learning with Policy Gradient Approach for Robotic Table   Tennis

Yapeng Gao; Jonas Tebbe; Andreas Zell

arXiv:2109.03100·cs.RO·October 11, 2022

Optimal Stroke Learning with Policy Gradient Approach for Robotic Table Tennis

Yapeng Gao, Jonas Tebbe, Andreas Zell

PDF

1 Video

TL;DR

This paper introduces a policy gradient method based on TD3 for robotic table tennis, achieving high success rates in real scenarios with efficient training and domain transfer from simulation.

Contribution

A novel policy gradient approach with TD3 backbone for learning robotic table tennis strokes, enabling effective simulation-to-reality transfer with minimal training.

Findings

01

Achieved 98% success rate in real scenarios

02

Reduced training time to approximately 1.5 hours

03

Outperformed existing RL methods in simulation

Abstract

Learning to play table tennis is a challenging task for robots, as a wide variety of strokes required. Recent advances have shown that deep Reinforcement Learning (RL) is able to successfully learn the optimal actions in a simulated environment. However, the applicability of RL in real scenarios remains limited due to the high exploration effort. In this work, we propose a realistic simulation environment in which multiple models are built for the dynamics of the ball and the kinematics of the robot. Instead of training an end-to-end RL model, a novel policy gradient approach with TD3 backbone is proposed to learn the racket strokes based on the predicted state of the ball at the hitting time. In the experiments, we show that the proposed approach significantly outperforms the existing RL methods in simulation. Furthermore, to cross the domain from simulation to reality, we adopt an…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

Man VS Machine: Who Plays Table Tennis Better? 🤖· youtube