z-logo
open-access-imgOpen Access
Unsupervised Learning of the Morphology of a Natural Language
Author(s) -
John Goldsmith
Publication year - 2001
Publication title -
computational linguistics
Language(s) - English
Resource type - Journals
SCImago Journal Rank - 0.314
H-Index - 98
eISSN - 1530-9312
pISSN - 0891-2017
DOI - 10.1162/089120101750300490
Subject(s) - computer science , generative grammar , heuristics , natural language processing , artificial intelligence , minimum description length , grammar , metric (unit) , segmentation , set (abstract data type) , natural language , generative model , probabilistic logic , linguistics , philosophy , operations management , economics , programming language , operating system
This study reports the results of using minimum description length (MDL) analysis to model unsupervised learning of the morphological segmentation of European languages, using corpora ranging in size from 5,000 words to 500,000 words. We develop a set of heuristics that rapidly develop a probabilistic morphological grammar, and use MDL as our primary tool to determine whether the modifications proposed by the heuristics will be adopted or not. The resulting grammar matches well the analysis that would be developed by a human morphologist.In the final section, we discuss the relationship of this style of MDL grammatical analysis to the notion of evaluation metric in early generative grammar.

The content you want is available to Zendy users.

Already have an account? Click here to sign in.
Having issues? You can contact us here
Accelerating Research

Address

John Eccles House
Robert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom