MANGO: A Benchmark for Evaluating Mapping and Navigation Abilities of   Large Language Models

Peng Ding; Jiading Fang; Peng Li; Kangrui Wang; Xiaochen; Zhou; Mo Yu; Jing Li; Matthew R. Walter; Hongyuan Mei

arXiv:2403.19913·cs.CL·August 9, 2024·2 cites

MANGO: A Benchmark for Evaluating Mapping and Navigation Abilities of Large Language Models

Peng Ding, Jiading Fang, Peng Li, Kangrui Wang, Xiaochen, Zhou, Mo Yu, Jing Li, Matthew R. Walter, Hongyuan Mei

PDF

Open Access 1 Repo

TL;DR

MANGO is a new benchmark designed to evaluate large language models' abilities in text-based mapping and navigation tasks using maze question-answering, revealing current models' limitations and guiding future improvements.

Contribution

This paper introduces MANGO, a comprehensive benchmark for assessing and advancing the mapping and navigation skills of large language models in text-based environments.

Findings

01

GPT-4 performs poorly on maze navigation questions

02

Mapping and navigation skills are crucial for downstream tasks like textgame playing

03

MANGO provides a platform for future research and improvement

Abstract

Large language models such as ChatGPT and GPT-4 have recently achieved astonishing performance on a variety of natural language processing tasks. In this paper, we propose MANGO, a benchmark to evaluate their capabilities to perform text-based mapping and navigation. Our benchmark includes 53 mazes taken from a suite of textgames: each maze is paired with a walkthrough that visits every location but does not cover all possible paths. The task is question-answering: for each maze, a large language model reads the walkthrough and answers hundreds of mapping and navigation questions such as "How should you go to Attic from West of House?" and "Where are we if we go north and east from Cellar?". Although these questions are easy to humans, it turns out that even GPT-4, the best-to-date language model, performs poorly at answering them. Further, our experiments suggest that a strong mapping…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

oaklight/mango
noneOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNatural Language Processing Techniques · Topic Modeling · Speech and dialogue systems

MethodsAttention Is All You Need · Linear Layer · Layer Normalization · Byte Pair Encoding · Multi-Head Attention · Softmax · Dense Connections · Label Smoothing · Adam · Absolute Position Encodings