A class on the BabelNet APIs for querying BabelNet and Wikipedia and for performing multilingual Word Sense Disambiguation, and other useful APIs for Wiktionary and other resources.
Home Page and Blog of the Multilingual NLP course @ Sapienza University of Rome
Saturday, May 26, 2012
Lecture 16: knowledge-based WSD (18/05/12)
Knowledge-based Word Sense Disambiguation. The Lesk and Extended Lesk algorithm. Structural approaches: similarity measures and graph algorithms. Conceptual density. Structural Semantic Interconnections. Evaluation: precision, recall, F1, accuracy. Baselines. The Senseval and SemEval evaluation competitions. Applications of Word Sense Disambiguation. Issues: representation of word senses, domain WSD, the knowledge acquisition bottleneck.
Wednesday, May 16, 2012
Lecture 15: more on Word Sense Disambiguation (15/05/12)
Supervised Word Sense Disambiguation: pros and cons. Vector representation of context. Main supervised disambiguation paradigms: decision trees, neural networks, instance-based learning, Support Vector Machines. Unsupervised Word Sense Disambiguation: Word Sense Induction. Context-based clustering. Co-occurrence graphs: curvature clustering, HyperLex.
Friday, May 11, 2012
Lecture 14: project presentation and Word Sense Disambiguation (11/05/12)
The NLP projects are online (2012_project.pdf). Please start your discussion! Topics today: introduction to Word Sense Disambiguation (WSD). Motivation. The typical WSD framework. Lexical sample vs. all-words. WSD viewed as lexical substitution and cross-lingual lexical substitution. Knowledge resources. Representation of context: flat and structured representations. Main approaches to WSD: Supervised, unsupervised and knowledge-based WSD. Two important dimensions: supervision and knowledge.
Friday, May 4, 2012
Lecture 13: lexical semantics and lexical resources (04/05/12)
Lexemes, lexicon, lemmas and word forms. Word senses: monosemy vs. polysemy. Special kinds of polysemy. Computational sense representations: enumeration vs. generation. Graded word sense assignment. Encoding word senses: paper dictionaries, thesauri, machine-readable dictionary, computational lexicons. WordNet. Wordnets in other languages. BabelNet.
Monday, April 30, 2012
Seminar by Prof. Iryna Gurevych: How to UBY - a Large-Scale Unified Lexical-Semantic Resource (30/04/12)
Title: How to UBY - a Large-Scale Unified Lexical-Semantic Resource
Speaker: Iryna Gurevych
Download the presentation
Abstract: The talk will present UBY, a large-scale resource integration
project based on the Lexical Markup Framework (LMF, ISO 24613:2008).
Currently, nine lexicons in two languages (English and German) have been
integrated: WordNet, GermaNet, FrameNet, VerbNet, Wikipedia (DE/EN),
Wiktionary (DE/EN), and OmegaWiki. All resources have been mapped to the
LMF-based model and imported into an SQL-DB. The UBY-API, a common Java
software library, provides access to all data in the database. The nine
lexicons are densely interlinked using monolingual and cross-lingual sense
alignments. These sense alignments yield enriched sense representations
and increased coverage. A sense alignment framework has been developed for
automatically aligning any pair of resources mono- or cross-lingually. As
an example, the talk will report on the automatic alignment of WordNet and
Wiktionary. Further information on UBY and UBY-API is available at:
http://www.ukp.tu-darmstadt. de/data/lexical-resources/uby/
Bio: Iryna Gurevych leads the UKP Lab in the Department of Computer
Science of the Technische Universität Darmstadt (UKP-TUDA) and at the
Institute for Educational Research and Educational Information (UKP-DIPF)
in Frankfurt. She holds an endowed Lichtenberg-Chair "Ubiquitous Knowledge
Processing" of the Volkswagen Foundation. Her research in NLP primarily concerns applied lexical semantic algorithms, such as computing semantic relatedness of words or paraphrase recognition, and their use to enhance the performance of NLP tasks, such as information retrieval, question answering, or summarization
Speaker: Iryna Gurevych
Abstract: The talk will present UBY, a large-scale resource integration
project based on the Lexical Markup Framework (LMF, ISO 24613:2008).
Currently, nine lexicons in two languages (English and German) have been
integrated: WordNet, GermaNet, FrameNet, VerbNet, Wikipedia (DE/EN),
Wiktionary (DE/EN), and OmegaWiki. All resources have been mapped to the
LMF-based model and imported into an SQL-DB. The UBY-API, a common Java
software library, provides access to all data in the database. The nine
lexicons are densely interlinked using monolingual and cross-lingual sense
alignments. These sense alignments yield enriched sense representations
and increased coverage. A sense alignment framework has been developed for
automatically aligning any pair of resources mono- or cross-lingually. As
an example, the talk will report on the automatic alignment of WordNet and
Wiktionary. Further information on UBY and UBY-API is available at:
http://www.ukp.tu-darmstadt.
Bio: Iryna Gurevych leads the UKP Lab in the Department of Computer
Science of the Technische Universität Darmstadt (UKP-TUDA) and at the
Institute for Educational Research and Educational Information (UKP-DIPF)
in Frankfurt. She holds an endowed Lichtenberg-Chair "Ubiquitous Knowledge
Processing" of the Volkswagen Foundation. Her research in NLP primarily concerns applied lexical semantic algorithms, such as computing semantic relatedness of words or paraphrase recognition, and their use to enhance the performance of NLP tasks, such as information retrieval, question answering, or summarization
Sunday, April 29, 2012
Mid-term exam (27/04/2012)
Mid-term exam: morphology and regular expressions, language modeling, probabilistic part-of-speech tagging, probabilistic syntactic parsing.
Lecture 12: introduction to semantics (20/04/12)
More exercises in preparation for the mid-term exam. Introduction to computational semantics. Syntax-driven semantic analysis. Semantic attachments. First-Order Logic. Lambda notation and lambda calculus for semantic representation.
Wednesday, April 18, 2012
Lecture 11: exercises (17/04/12)
Exercises about regular expressions for morphological analysis, n-gram models and stochastic part-of-speech tagging.
Subscribe to:
Posts (Atom)



