Is artificial intelligence still intelligence? LLMs generalize to novel   adjective-noun pairs, but don't mimic the full human distribution

Hayley Ross; Kathryn Davidson; Najoung Kim

arXiv:2410.17482·cs.CL·October 24, 2024

Is artificial intelligence still intelligence? LLMs generalize to novel adjective-noun pairs, but don't mimic the full human distribution

Hayley Ross, Kathryn Davidson, Najoung Kim

PDF

Open Access 1 Repo

TL;DR

This paper investigates the ability of large language models to generalize to novel adjective-noun pairs, showing they can mimic human judgments in context but still have limitations in out-of-context inferences.

Contribution

The study introduces methods to evaluate LLMs' understanding of adjective-noun combinations and reveals their partial success in mimicking human-like judgments.

Findings

01

Largest models generalize well in context

02

LLMs show human-like judgment distribution in 75% of cases

03

Room for improvement in out-of-context inference accuracy

Abstract

Inferences from adjective-noun combinations like "Is artificial intelligence still intelligence?" provide a good test bed for LLMs' understanding of meaning and compositional generalization capability, since there are many combinations which are novel to both humans and LLMs but nevertheless elicit convergent human judgments. We study a range of LLMs and find that the largest models we tested are able to draw human-like inferences when the inference is determined by context and can generalize to unseen adjective-noun combinations. We also propose three methods to evaluate LLMs on these inferences out of context, where there is a distribution of human-like answers rather than a single correct answer. We find that LLMs show a human-like distribution on at most 75\% of our dataset, which is promising but still leaves room for improvement.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

rossh2/artificial-intelligence
noneOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNatural Language Processing Techniques · Topic Modeling