NLP Components
This project provided several NLP tools such as a dependency parser, a semantic role labeler, a penn-to-dependency converter, a prop-to-dependency converter, and a morphological analyzer. All tools were written in Java and developed by the Computational Language and EducAtion Research (CLEAR) group at the University of Colorado at Boulder.
Word Sense Disambiguation and Efficient Annotation
Supervised machine learning was widely used in natural language processing and, based on the extensive OntoNotes sense tagged data, we had a state-of-the-art WSD system for English verbs that approached human accuracy. Check back to this cite soon for a link to a downloadable version.
However, porting this approach to other domains and other languages required additional annotated training data, which was expensive to obtain. How did one choose the data for annotation? Random sampling was a common approach but not the most efficient one. Various types of selective sampling could be used to achieve the same level of performance as random sampling but with less data. Active learning was one type of selective sampling, but in many situations it was not practical (e.g. a multi-annotator, double-annotation environment). Dmitry Dligach's dissertation focused on developing selective sampling algorithms that were similar in spirit to active learning but more practical. They utilized his state-of-the-art automatic word sense disambiguation system. He had also looked into evaluating various popular annotation practices such as single annotation, double annotation, and batch active learning.
VerbNet Class Disambiguator
Understanding verbs was central to deep semantic parsing, requiring the identification of not only a verb's meaning but also how it connected the participants in the sentence. Disambiguating verbs using a lexicon that had already been enriched with syntactic and semantic information would bring end systems a step closer to accurate knowledge representation and reasoning than a more traditional lexicon. VerbNet had already been shown to be a good resource for identifying deep semantics, having been used for semantic role labeling (Swier and Stevenson, 2004), the creation of conceptual graphs (Hensman and Dunion, 2004), and semantic parsing (Shi and Mihalcea, 2005). However, many verbs were members of multiple VerbNet classes, with each class membership corresponding roughly to different senses of the verbs. Therefore, application of VerbNet's semantic and syntactic information to specific text required first identifying the appropriate VerbNet class of each verb in the text.
At that time in development was the The VerbNet Class Disambiguator, which used a supervised machine learning approach to classify verb tokens with VerbNet classes. It had been trained and tested with 30 verbs to date. With this initial sample, it achieved 90% accuracy, which represented a 61% error reduction over the most-frequent-class baseline. Work was underway to increase its coverage to all the multiclass verbs in the Semlink corpus.