Abstract
© 2017 Association for Computational LinguisticsThis paper presents the CIC UALG's system that took part in the Discriminating between Similar Languages (DSL) shared task, held at the VarDial 2017 Workshop. This year's task aims at identifying 14 languages across 6 language groups using a corpus of excerpts of journalistic texts. Two classification approaches were compared: a single-step (all languages) approach and a two-step (language group and then languages within the group) approach. Features exploited include lexical features (unigrams of words) and character n-grams. Besides traditional (untyped) character n-grams, we introduce typed character n-grams in the DSL task. Experiments were carried out with different feature representation methods (binary and raw term frequency), frequency threshold values, and machine-learning algorithms - Support Vector Machines (SVM) and Multinomial Naive Bayes (MNB). Our best run in the DSL task achieved 91.46% accuracy.
| Original language | English |
|---|---|
| Title of host publication | VarDial 2017 - 4th Workshop on NLP for Similar Languages, Varieties and Dialects, Proceedings |
| Publisher | Association for Computational Linguistics (ACL) |
| Pages | 137-145 |
| ISBN (Electronic) | 9781945626432 |
| Publication status | Published - 2017 |
| Externally published | Yes |
| Event | 4th Workshop on NLP for Similar Languages, Varieties and Dialects, VarDial 2017 - Valencia, Spain Duration: 3 Apr 2017 → … |
Conference
| Conference | 4th Workshop on NLP for Similar Languages, Varieties and Dialects, VarDial 2017 |
|---|---|
| Country/Territory | Spain |
| City | Valencia |
| Period | 3/04/17 → … |
Funding
This work was partially supported by the Mexican Government (Conacyt projects 240844 and 20161958, SIP-IPN 20151406, 20161947, 20161958, 20151589, 20162204, and 20162064, SNI, COFAA-IPN) and by the Portuguese Government, through Fundac¸ão para a Ciência e a Tecnologia (FCT) with reference UID/CEC/50021/2013. This work was partially supported by the Mexican Government (Conacyt projects 240844 and 20161958, SIP-IPN 20151406, 20161947, 20161958, 20151589, 20162204, and 20162064, SNI, COFAA-IPN) and by the Portuguese Government, through Funda??o para a Ci?ncia e a Tecnologia (FCT) with reference UID/CEC/50021/2013.
| Funders | Funder number |
|---|---|
| Ci?ncia e a Tecnologia | |
| Fundac¸ão para a Ciência e a Tecnologia | |
| Mexican Government | |
| Fundação para a Ciência e a Tecnologia | UID/CEC/50021/2013 |
| Consejo Nacional de Ciencia y Tecnología | SIP-IPN 20151406, 20151589, 20162204, 20162064, 20161947, 20161958, 240844 |
| Sistema Nacional de Investigadores |
Fingerprint
Dive into the research topics of 'Discriminating between similar languages using a combination of typed and untyped character n-grams and words'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver