Spelling Correction with the Dictionary Method for the Turkish Language Using Word Embeddings

Aydoğan, Murat; Karcı, Ali

Spelling Correction with the Dictionary Method for the Turkish Language Using Word Embeddings

dc.contributor.author	Aydoğan, Murat
dc.contributor.author	Karcı, Ali
dc.date.accessioned	2021-03-24T08:34:33Z
dc.date.available	2021-03-24T08:34:33Z
dc.date.issued	2020
dc.department	İnönü Üniversitesi	en_US
dc.description.abstract	Abstract: Today, a massive amount of data is being produced, which is referred to as “big data.” A significant part of big data is composed of text data, which has made text processing all the more important. However, when text processing studies are examined, it can be seen that while there are many world language-oriented studies, especially the English language, there has been an insufficient level of studies published specific to the Turkish language. Therefore, Turkish was chosen as the target language for the study. A Turkish corpus of approximately 10.5 billion words was created, consisting of unlabeled data containing no spelling errors. Word vectors were trained using the Word2Vec method on this corpus. Based on this corpus, a new method was proposed called the “dictionary method,” with a dictionary created covering almost all known Turkish words. Then, text classification was applied to a multi-class Turkish dataset. This dataset contains 10 classes and approximately 1.5 million samples. Vector values of the token words in this dataset were transferred from the dictionary by transfer learning. However, words not found in the created dictionary were considered as incorrect; then, using LSTM (Long Short-Term Memory), which is a deep neural network (DNN) architecture, the proposed method attempts to predict correct or similar words as replacement words. Following this process, it was seen that the accuracy rate improved by 8.68%. Turkish dataset that is created, corpus and dictionary will be shared with researchers in order to contribute to Turkish text processing studies.	en_US
dc.description.abstract	Öz: Günümüzde oldukça büyük miktarda veri üretilmektedir. Üretilen bu büyük verinin çok önemli bir kısmı ise text verilerinden oluşmaktadır. Bu durum, text processing çalışmalarının daha da önem kazanmasını sağlamıştır. Ancak yapılan çalışmalar incelendiğinde başta İngilizce olmak üzere birçok dünya dili odaklı çalışmalar yapılırken Türkçe diline özgü çalışmaların yeterli sayıda olmadığı görülmüştür. Bu nedenle bu çalışmada hedef dil olarak Türkçe seçilmiştir. Etiketsiz verilerden oluşan ve yazım yanlışı bulunmayan yaklaşık 10.5 milyar kelimeden oluşan etiketsiz ve büyük Türkçe bir derlem üretilmiştir. Word2Vec metodu kullanılarak bu derlem üzerinde kelime vektörleri eğitilmiştir. Bu derlemi temel alarak “Sözlük Metodu” adı verilen yeni bir yöntem önerilmiştir, üretilen derlem içindeki kelimeler ile hemen hemen tüm Türkçe kelimeleri kapsayan bir sözlük oluşturulmuştur. Daha sonra çok sınıflı Türkçe bir dataset üzerinde metin sınıflandırma işlemi uygulanmıştır. Bu veriseti içerisindeki token kelimelerin vektörel değerleri sözlükten transfer öğrenme ile aktarılmıştır. Ancak sözlükte bulunmayan kelimelerin hatalı kelimeler olduğu düşünülerek bir derin sinir ağı mimarisi olan LSTM (Uzun Kısa Süreli Bellek) yöntemi ile bu kelimelerin yerine doğru veya yakın anlamlı kelimeler tahmin edilmeye çalışılmıştır. Bu işlemin ardından metin sınıflandırma uygulamasının doğruluk oranında %8.68 oranında gelişme olduğu görülmüştür. Üretilen Türkçe veriseti, derlem ve sözlük Türkçe metin işleme çalışmalarına katkı sağlamak amacıyla araştırmacılarla paylaşılacaktır.	en_US
dc.identifier.citation	AYDOĞAN M,KARCI A (2020). Spelling Correction with the Dictionary Method for the Turkish Language Using Word Embeddings. Avrupa Bilim ve Teknoloji Dergisi, 0(Ejosat Özel Sayı 2020 (ARACONF)), 57 - 63. Doi: 10.31590/ejosat.araconf8	en_US
dc.identifier.doi	10.31590/ejosat.araconf8	en_US
dc.identifier.endpage	63	en_US
dc.identifier.issn	2148-2683
dc.identifier.issue	Ejosat Özel Sayı 2020 (ARACONF)	en_US
dc.identifier.startpage	57	en_US
dc.identifier.trdizinid	364706	en_US
dc.identifier.uri	https://doi.org/10.31590/ejosat.araconf8
dc.identifier.uri	https://hdl.handle.net/11616/19717
dc.identifier.uri	https://search.trdizin.gov.tr/yayin/detay/364706
dc.indekslendigikaynak	TR-Dizin	en_US
dc.language.iso	en	en_US
dc.relation.ispartof	Avrupa Bilim ve Teknoloji Dergisi	en_US
dc.relation.publicationcategory	Makale - Ulusal Hakemli Dergi - Kurum Öğretim Elemanı	en_US
dc.rights	info:eu-repo/semantics/openAccess	en_US
dc.title	Spelling Correction with the Dictionary Method for the Turkish Language Using Word Embeddings	en_US
dc.title.alternative	Kelime Gömmelerini Kullanarak Türkçe Dili İçin Sözlük Metodu ile Yazım Düzeltme	en_US
dc.type	Article	en_US

Dosyalar

Orijinal paket

Listeleniyor 1 - 1 / 1

İsim:: Makale Dosyası.pdf
Boyut:: 1.01 MB
Biçim:: Adobe Portable Document Format
Açıklama:: Makale Doyası

İndir

Lisans paketi

Listeleniyor 1 - 1 / 1

İsim:: license.txt
Boyut:: 1.71 KB
Biçim:: Item-specific license agreed upon to submission
Açıklama:

İndir

Koleksiyon

TR-Dizin İndeksli Yayınlar Koleksiyonu