Affiliation:
1. Bar Ilan University, Israel. omer.goldman@gmail.com
2. Bar Ilan University, Israel. reut.tsarfaty@biu.ac.il
Abstract
Abstract
Morphological tasks use large multi-lingual datasets that organize words into inflection tables, which then serve as training and evaluation data for various tasks. However, a closer inspection of these data reveals profound cross-linguistic inconsistencies, which arise from the lack of a clear linguistic and operational definition of what is a word, and which severely impair the universality of the derived tasks. To overcome this deficiency, we propose to view morphology as a clause-level phenomenon, rather than word-level. It is anchored in a fixed yet inclusive set of features, that encapsulates all functions realized in a saturated clause. We deliver MightyMorph, a novel dataset for clause-level morphology covering 4 typologically different languages: English, German, Turkish, and Hebrew. We use this dataset to derive 3 clause-level morphological tasks: inflection, reinflection and analysis. Our experiments show that the clause-level tasks are substantially harder than the respective word-level tasks, while having comparable complexity across languages. Furthermore, redefining morphology to the clause-level provides a neat interface with contextualized language models (LMs) and allows assessing the morphological knowledge encoded in these models and their usability for morphological tasks. Taken together, this work opens up new horizons in the study of computational morphology, leaving ample space for studying neural morphology cross-linguistically.
Subject
Artificial Intelligence,Computer Science Applications,Linguistics and Language,Human-Computer Interaction,Communication
Reference57 articles.
1. Subjectivity and sentiment analysis of Modern Standard Arabic;Abdul-Mageed,2011
2. Parsing by chunks;Abney,1991
3. Parts and wholes: Patterns of relatedness in complex morphological systems and why they matter;Ackerman;Analogy in Grammar: Form and Acquisition,2009
4. Paradigms and periphrastic expression: A study in realization-based lexicalism;Ackerman;Projecting Morphology,2004
5. Representations and architectures in neural sentiment analysis for morphologically rich languages: A case study from Modern Hebrew;Amram,2018