91×ÔÅÄ

Language Technologies and Corpus Linguistics Seminar Begins at KTMU


  • 2026-09-22

91×ÔÅÄ (KTMU) is hosting an international seminar addressing the use of natural language processing technologies and corpus linguistics in language and translation research. The seminar, organized in cooperation with the Department of Translations at the Faculty of Humanities of KTMU and Saarland University in Germany, will continue from September 21 to 24, 2026.

The seminar, titled “Natural Language Processing Technologies and Corpus Linguistics” and organized within the framework of the Erasmus+ KA171 Inter-Institutional Agreement, began on September 21 at the Kasım Tynystanov Hall of the Faculty of Humanities. During the four-day program, participants are receiving theoretical and practical knowledge on natural language processing, corpus creation, and the use of digital research tools.

Digital Methods in Language Research to Be Examined

The seminar focuses on current methods at the intersection of computer science, artificial intelligence, and linguistics. Topics include metadata and linguistic annotation, syntactic analysis, Universal Dependencies (UD), corpus querying using regular expressions, AI-assisted corpus analysis, discourse analysis, qualitative data analysis, and parallel corpora.

In the seminar sessions conducted by Dr. Marie-Pauline Krielke and Dr. Andrea Wurm from Saarland University, participants will carry out practical activities on selecting corpus tools appropriate to their research questions, creating queries, performing syntactic analysis, and making use of different digital platforms.

The program also includes topics such as examining translation errors through student translator corpora, evaluating errors arising from content and linguistic knowledge, translation history, and bilingualism research. The creation of parallel corpora from historical documents in French and German, scanning the documents and converting them into text with the help of artificial intelligence, aligning the texts, and manually annotating them are among the practical activities included in the program.

Digital Resources for the Kyrgyz Language Are Being Developed

At the opening of the training, Prof. Dr. Aida Kasieva, Head of the Department of Translations, gave a presentation titled “Corpus Linguistics of the Kyrgyz Language.” The presentation focused on the development of electronic corpora for Kyrgyz and the use of language data in scientific research.

In the program, participants are introduced to methods of selecting, creating, and using corpora appropriate to their own research questions. The program also examines how parallel corpus studies can be conducted using different language pairs, such as Kyrgyz-English, Kyrgyz-Turkish, Russian-Turkish, and Turkish-English.

Intr. Dr. Leyla Babatürk, Vice Dean of the Faculty of Humanities at KTMU, emphasized the importance of developing high-quality and verified digital resources for Kyrgyz. Babatürk stated that the first large electronic corpus of Kyrgyz with morphological annotation began to be created in 2019 using literary texts, and that the corpus reached approximately 4 million words with the addition of newspaper archives, proverbs, and the 10-volume collected works of Chyngyz Aitmatov.

Stating that artificial intelligence systems are fed by language data, Babatürk noted that reliable digital language resources are needed to reduce the errors produced by these systems in Kyrgyz.

Prof. Dr. Aida Kasiyeva also stated that the seminar aims to contribute to the study, preservation, and development of Kyrgyz through contemporary technologies. Drawing attention to the fact that Kyrgyz is among the languages with limited digital resources, Kasiyeva emphasized the importance of efforts to strengthen the language’s presence in the digital environment.

Practical Program from Artificial Intelligence to Parallel Corpora

On the second day of the seminar, the focus is on metadata, linguistic annotation, corpus queries using regular expressions, and AI-assisted corpus analysis. On the third day, participants will work on Universal Dependencies, syntactic analysis, and discourse analysis using Sketch Engine. On the final day, qualitative data analysis using MAXQDA and TEI/XML, parallel corpora, and related digital applications will be covered.

Attention Drawn to the Development of the Kyrgyz Language in the Digital Environment

The fact that the training is being held during the week that includes September 23, the Day of the State Language of the Kyrgyz Republic, also brings efforts aimed at preserving and developing Kyrgyz in the digital environment to the agenda.

This aspect of the seminar is in line with the approach of the “National Spirit – Global Summit” Doctrine, published by the Presidency of the Kyrgyz Republic, concerning the preservation and development of the state language and its strengthening through the opportunities provided by modern science and technology.

Recalling that Kyrgyz was granted the status of a state language on September 23, 1989, Aida Kasiyeva referred to the importance of State Language Day in terms of preserving and developing the mother tongue and expanding its use in different areas of social life.

The four-day international seminar will conclude on September 24. Participants will receive certificates at the end of the program.

    Share on social media: