Adaptor Grammars for the Linguist: Word Segmentation Experiments for Very Low-Resource Languages - Laboratoire d'Informatique pour la Mécanique et les Sciences de l'Ingénieur Accéder directement au contenu
Communication Dans Un Congrès Année : 2018

Adaptor Grammars for the Linguist: Word Segmentation Experiments for Very Low-Resource Languages

Résumé

Computational Language Documentation attempts to make the most recent research in speech and language technologies available to linguists working on language preservation and documentation. In this paper, we pursue two main goals along these lines. The first is to improve upon a strong baseline for the unsupervised word discovery task on two very low-resource Bantu languages, taking advantage of the expertise of linguists on these particular languages. The second consists in exploring the Adaptor Grammar framework as a decision and prediction tool for linguists studying a new language. We experiment 162 grammar configurations for each language and show that using Adaptor Grammars for word segmentation enables us to test hypotheses about a language. Specializing a generic grammar with language specific knowledge leads to great improvements for the word discovery task, ultimately achieving a leap of about 30% token F-score from the results of a strong baseline.
Fichier principal
Vignette du fichier
W18-5804.pdf (637.89 Ko) Télécharger le fichier
Origine : Fichiers produits par l'(les) auteur(s)
Loading...

Dates et versions

hal-01910757 , version 1 (01-11-2018)

Identifiants

Citer

Pierre Godard, Laurent Besacier, François Yvon, Martine Adda-Decker, Gilles Adda, et al.. Adaptor Grammars for the Linguist: Word Segmentation Experiments for Very Low-Resource Languages. Workshop on Computational Research in Phonetics, Phonology, and Morphology, Oct 2018, Bruxelles, Belgium. pp.32 - 42, ⟨10.18653/v1/P17⟩. ⟨hal-01910757⟩
189 Consultations
235 Téléchargements

Altmetric

Partager

Gmail Facebook X LinkedIn More