Skip to main navigation Skip to search Skip to main content

Discriminating between similar languages using a combination of typed and untyped character n-grams and words

  • H. Gómez-Adorno
  • , I. Markov
  • , J. Baptista
  • , G. Sidorov
  • , D. Pinto

Research output: Chapter in Book / Report / Conference proceedingConference contributionAcademicpeer-review

Abstract

© 2017 Association for Computational LinguisticsThis paper presents the CIC UALG's system that took part in the Discriminating between Similar Languages (DSL) shared task, held at the VarDial 2017 Workshop. This year's task aims at identifying 14 languages across 6 language groups using a corpus of excerpts of journalistic texts. Two classification approaches were compared: a single-step (all languages) approach and a two-step (language group and then languages within the group) approach. Features exploited include lexical features (unigrams of words) and character n-grams. Besides traditional (untyped) character n-grams, we introduce typed character n-grams in the DSL task. Experiments were carried out with different feature representation methods (binary and raw term frequency), frequency threshold values, and machine-learning algorithms - Support Vector Machines (SVM) and Multinomial Naive Bayes (MNB). Our best run in the DSL task achieved 91.46% accuracy.
Original languageEnglish
Title of host publicationVarDial 2017 - 4th Workshop on NLP for Similar Languages, Varieties and Dialects, Proceedings
PublisherAssociation for Computational Linguistics (ACL)
Pages137-145
ISBN (Electronic)9781945626432
Publication statusPublished - 2017
Externally publishedYes
Event4th Workshop on NLP for Similar Languages, Varieties and Dialects, VarDial 2017 - Valencia, Spain
Duration: 3 Apr 2017 → …

Conference

Conference4th Workshop on NLP for Similar Languages, Varieties and Dialects, VarDial 2017
Country/TerritorySpain
CityValencia
Period3/04/17 → …

Funding

This work was partially supported by the Mexican Government (Conacyt projects 240844 and 20161958, SIP-IPN 20151406, 20161947, 20161958, 20151589, 20162204, and 20162064, SNI, COFAA-IPN) and by the Portuguese Government, through Fundac¸ão para a Ciência e a Tecnologia (FCT) with reference UID/CEC/50021/2013. This work was partially supported by the Mexican Government (Conacyt projects 240844 and 20161958, SIP-IPN 20151406, 20161947, 20161958, 20151589, 20162204, and 20162064, SNI, COFAA-IPN) and by the Portuguese Government, through Funda??o para a Ci?ncia e a Tecnologia (FCT) with reference UID/CEC/50021/2013.

FundersFunder number
Ci?ncia e a Tecnologia
Fundac¸ão para a Ciência e a Tecnologia
Mexican Government
Fundação para a Ciência e a TecnologiaUID/CEC/50021/2013
Consejo Nacional de Ciencia y TecnologíaSIP-IPN 20151406, 20151589, 20162204, 20162064, 20161947, 20161958, 240844
Sistema Nacional de Investigadores

    Fingerprint

    Dive into the research topics of 'Discriminating between similar languages using a combination of typed and untyped character n-grams and words'. Together they form a unique fingerprint.

    Cite this