Towards Natural Language-Driven Assembly Using Foundation Models

Omkar Joglekar; Tal Lancewicki; Shir Kozlovsky; Vladimir Tchuiev,; Zohar Feldman; Dotan Di Castro

arXiv:2406.16093·cs.RO·June 25, 2024

Towards Natural Language-Driven Assembly Using Foundation Models

Omkar Joglekar, Tal Lancewicki, Shir Kozlovsky, Vladimir Tchuiev,, Zohar Feldman, Dotan Di Castro

PDF

Open Access

TL;DR

This paper introduces a novel approach using Large Language Models to develop a global control policy for robots, enabling high-precision assembly tasks through dynamic skill switching and enhanced language understanding.

Contribution

It presents a framework that leverages LLMs for controlling robots with high precision by dynamically switching between specialized skills, addressing limitations of generalist policies in industrial tasks.

Findings

01

Effective transfer of control policies to high-precision skills

02

Enhanced language interpretation for robotic control

03

Successful dynamic context switching in assembly tasks

Abstract

Large Language Models (LLMs) and strong vision models have enabled rapid research and development in the field of Vision-Language-Action models that enable robotic control. The main objective of these methods is to develop a generalist policy that can control robots with various embodiments. However, in industrial robotic applications such as automated assembly and disassembly, some tasks, such as insertion, demand greater accuracy and involve intricate factors like contact engagement, friction handling, and refined motor skills. Implementing these skills using a generalist policy is challenging because these policies might integrate further sensory data, including force or torque measurements, for enhanced precision. In our method, we present a global control policy based on LLMs that can transfer the control policy to a finite set of skills that are specifically trained to perform…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsModular Robots and Swarm Intelligence

MethodsSparse Evolutionary Training