z-logo
open-access-imgOpen Access
A Resource-Light Approach to Morpho-Syntactic Tagging Anna Feldman* and Jirka Hana (*Montclair State University, Charles University) Amsterdam: Rodopi (Language and computers: Studies in practical linguistics, volume 70), 2010, xiv+185 pp; hardbound, ISBN 978-90-420-2768-8, €40.00
Author(s) -
Christian Monson
Publication year - 2011
Publication title -
computational linguistics
Language(s) - English
Resource type - Journals
SCImago Journal Rank - 0.314
H-Index - 98
eISSN - 1530-9312
pISSN - 0891-2017
DOI - 10.1162/coli_r_00042
Subject(s) - computer science , state (computer science) , volume (thermodynamics) , morpho , linguistics , natural language processing , artificial intelligence , programming language , philosophy , physics , quantum mechanics , optics
Anna Feldman and Jirka Hana had a problem. Wanting to extract Russian verb frames, they lacked a tool for the necessary first step: morphological analysis of Russian words, disambiguated for context. To avoid the significant overhead of building a contextual-ized morphological analyzer from scratch, Feldman and Hana wondered if an analyzer that was already available for Czech would perform adequately on Russian. This book is the culmination of five years' research on projecting to a target language a contextualized morphological analyzer that was built for a separate source language, when both source and target belong to the same language family (Slavic, Romance, etc.). The authors succeed at building competitive morphological analysis systems for the target languages they consider (Russian, Catalan, and Portuguese), while expending a minimum of effort to construct specialized resources for these targets. At the culmination of their book, in Chapter 7, Feldman and Hana report a 6% absolute improvement, 79.7% vs. 73.5% labeling accuracy, when using a Czech morphological analyzer projected to Russian as opposed to training a statistical analyzer directly on a small sample (1,758 words) of hand-annotated Russian. Unfortunately missing is a formal demonstration that hand-labeling 1,758 words with morphological analyses requires an equivalent human effort to projecting an analyzer from one language to another. The authors' final morphological projection incorporates a variety of improvements that require human intervention: from a hand-built morphological guesser on the target language side, to hand-defined rules that identify cognates between source and target languages and that render the syntactic structure of the source language more similar to the target's structure. Nowhere do the authors report the person-hours required to build each of these components and the reader is left to trust that constructing the projected systems takes as little time as is implied. A word of warning to those with a linguistics background: The authors prefer the language of natural language processing (NLP) to standard linguistic terminology. As a prime example, the title of this book includes the phrase morpho-syntactic tagging, a term from NLP. Part-of-speech tagging, in languages with little inflectional morphology, such as English, involves assigning to each word one part-of-speech tag from a small set of 50 or so possible tags. For the more inflected Slavic and Romance languages considered in this book, the tag sets include as many as 4,000 tags, each marking a full suite of morpho-syntactic features, such as tense, case, or number. …

The content you want is available to Zendy users.

Already have an account? Click here to sign in.
Having issues? You can contact us here
Accelerating Research

Address

John Eccles House
Robert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom