TY - GEN
T1 - News Article Text Classification in Indonesian Language
AU - Wongso, Rini
AU - Luwinda, Ferdinand Ariandy
AU - Trisnajaya, Brandon Christian
AU - Rusli, Olivia
AU - Rudy, null
PY - 2017
Y1 - 2017
N2 - This research intends to find the appropriate algorithm to automatically classify a news article in Indonesian Language. We obtain our dataset which is taken by using a web crawling method from www.cnnindonesia.com. First of all, the document will first undergo some Text Preprocessing method in the form of Lemmatization and Stopwords Removal. The reason we are doing the Text Preprocessing step before anything else is to minimize the noise in the document. Next, we apply Feature Selection onto the document to further separate important words and less important words inside the document. After applying Feature Selection, the document will be classified by the classifier. We are comparing the TF-IDF and SVD algorithm for feature selection, while also comparing the Multinomial Naïve Bayes, Multivariate Bernoulli Naïve Bayes, and Support Vector Machine for the Classifiers. Based on the test results, the combination of TF-IDF and Multinomial Naïve Bayes Classifier gives the highest result compared to the other algorithms, which precision is 0.9841519 and its recall is 0.9840000. The result outperform the previous similar study that classify news article in Indonesian language which obtained 85% of accuracy.
AB - This research intends to find the appropriate algorithm to automatically classify a news article in Indonesian Language. We obtain our dataset which is taken by using a web crawling method from www.cnnindonesia.com. First of all, the document will first undergo some Text Preprocessing method in the form of Lemmatization and Stopwords Removal. The reason we are doing the Text Preprocessing step before anything else is to minimize the noise in the document. Next, we apply Feature Selection onto the document to further separate important words and less important words inside the document. After applying Feature Selection, the document will be classified by the classifier. We are comparing the TF-IDF and SVD algorithm for feature selection, while also comparing the Multinomial Naïve Bayes, Multivariate Bernoulli Naïve Bayes, and Support Vector Machine for the Classifiers. Based on the test results, the combination of TF-IDF and Multinomial Naïve Bayes Classifier gives the highest result compared to the other algorithms, which precision is 0.9841519 and its recall is 0.9840000. The result outperform the previous similar study that classify news article in Indonesian language which obtained 85% of accuracy.
UR - https://www.scopus.com/pages/publications/85040012495
U2 - 10.1016/j.procs.2017.10.039
DO - 10.1016/j.procs.2017.10.039
M3 - Conference contribution
SN - 9781510849914
T3 - Procedia Computer Science
SP - 137
EP - 143
BT - 2nd International Conference on Computer Science and Computational Intelligence, ICCSCI 2017
A2 - Budiharto, W.
A2 - Suryani, D.
A2 - Wulandhari, L.A.
A2 - Chowanda, A.
A2 - Gunawan, A.A.S.
A2 - Hanafiah, N.
A2 - Ham, H.
A2 - Meiliana, null
PB - Elsevier B.V.
T2 - 2nd International Conference on Computer Science and Computational Intelligence, ICCSCI 2017
Y2 - 13 October 2017 through 14 October 2017
ER -