Robust Probabilistic Predictive Syntactic Processing

Roark, Brian

Computer Science > Computation and Language

arXiv:cs/0105019 (cs)

[Submitted on 9 May 2001]

Title:Robust Probabilistic Predictive Syntactic Processing

Authors:Brian Roark

View PDF

Abstract: This thesis presents a broad-coverage probabilistic top-down parser, and its application to the problem of language modeling for speech recognition. The parser builds fully connected derivations incrementally, in a single pass from left-to-right across the string. We argue that the parsing approach that we have adopted is well-motivated from a psycholinguistic perspective, as a model that captures probabilistic dependencies between lexical items, as part of the process of building connected syntactic structures. The basic parser and conditional probability models are presented, and empirical results are provided for its parsing accuracy on both newspaper text and spontaneous telephone conversations. Modifications to the probability model are presented that lead to improved performance. A new language model which uses the output of the parser is then defined. Perplexity and word error rate reduction are demonstrated over trigram models, even when the trigram is trained on significantly more data. Interpolation on a word-by-word basis with a trigram model yields additional improvements.

Comments:	Ph.D. Thesis, Brown University, Advisor: Mark Johnson. 140 pages, 40 figures, 27 tables
Subjects:	Computation and Language (cs.CL)
ACM classes:	I.2.7
Cite as:	arXiv:cs/0105019 [cs.CL]
	(or arXiv:cs/0105019v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.cs/0105019

Submission history

From: Brian Roark [view email]
[v1] Wed, 9 May 2001 17:01:10 UTC (163 KB)

Computer Science > Computation and Language

Title:Robust Probabilistic Predictive Syntactic Processing

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Robust Probabilistic Predictive Syntactic Processing

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators