Can Subcategorisation Probabilities Help a Statistical Parser?

John Carroll; Guido Minnen (University of Sussex); Ted Briscoe; (Cambridge University)

arXiv:cmp-lg/9806013·cmp-lg·May 23, 2007·56 cites

Can Subcategorisation Probabilities Help a Statistical Parser?

John Carroll, Guido Minnen (University of Sussex), Ted Briscoe, (Cambridge University)

PDF

Open Access

TL;DR

This paper investigates whether incorporating subcategorisation frequency data into a statistical parser improves its accuracy, demonstrating significant gains with large-scale lexical frequency information.

Contribution

It provides empirical evidence that subcategorisation probabilities derived from large corpora can enhance the performance of statistical parsers.

Findings

01

Subcategorisation frequencies improve parser accuracy

02

Large-scale lexical data benefits parsing performance

03

Empirical validation with ten million words

Abstract

Research into the automatic acquisition of lexical information from corpora is starting to produce large-scale computational lexicons containing data on the relative frequencies of subcategorisation alternatives for individual verbal predicates. However, the empirical question of whether this type of frequency information can in practice improve the accuracy of a statistical parser has not yet been answered. In this paper we describe an experiment with a wide-coverage statistical grammar and parser for English and subcategorisation frequencies acquired from ten million words of text which shows that this information can significantly improve parse accuracy.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNatural Language Processing Techniques · Topic Modeling · Speech and dialogue systems