About the project

PROJECT FOR CREATING AN ACADEMIC KAZAKH CORPUS

Research project topic: Digital humanities: Developing an academic Kazakh language corpus

Project number: AP23488585

Implementation period: 2024-2026

Funding organisation: Ministry of Science and Higher Education of the Republic of Kazakhstan

Concept of the research project

The project's primary objective is to enhance the scientific and practical capabilities of the Kazakh language in academic texts, including monographs, articles, reports, theses, abstracts, and dissertations, with the aim of establishing the concept of academic Kazakh and selecting vocabulary for academic texts in Kazakh.

The project's research will address several issues, including increasing the number and diversity of Kazakh language corpora, promoting the Kazakh language internationally through digital corpora, enhancing the scientific potential and functional capabilities of the national language, and accelerating the digitisation of the humanities.

This interdisciplinary study combines humanities and computer research to present digital humanities as a set of methodological approaches that can be used in corpus linguistics.

The project aims to address a significant empirical gap in modern research infrastructure for studying the current state of scientific Kazakh by introducing an academic written corpus of modern Kazakh, comprising at least 5,000,000 words, of which at least 50,000 are annotated.

Academic Kazakh Corpus

The Academic Kazakh Corpus is a corpus of written scientific texts. The current volume of the Academic Written Kazakh Corpus (AWKC) consists of more than 25 million tokens (Table 1).

Table 1 - Statistics of the Academic Kazakh Corpus

IndicatorValue
Total text count2887
Total token count25 487 465
Total sentence count1 267 045

In collecting and selecting academic texts for the corpus, it was intended to maintain the condition of data diversity (variation). Therefore, the corpus included various types of academic texts, scientific articles, monographs, dissertations, abstracts, reports, annotations, textbooks, teaching aids, and theses. All texts were published between 2000 and 2024. The names of academic text types were given in full, and abbreviated names were given in Kazakh and English. The abbreviated English names were included in the data naming convention (Table 2).

Table 2 - Academic text names and abbreviations

In KazakhIn EnglishAbbreviation
монографияmonographmonogr
диссертацияdissertationdiss
выпускная работаgraduate papergradpr
статьяarticleartic
тезисthesisthesis
учебникtextbooktxtbk
учебное пособиеtextbook1txtbk1
абстрактabstractabst
докладconference paperconfp
аннотацияannotationannot

All texts included in the academic corpus are academic written works published in Kazakhstan. The corpus does not contain samples of spoken language. The academic written Kazakh corpus includes 12 subjects in the Humanities and Languages and Literary Sciences according to the Register of Higher and Postgraduate Educational Programs of the Ministry of Science and Higher Education of the Republic of Kazakhstan. The names of the subjects are given in full and abbreviated names in Kazakh and English. The abbreviated English names are included in the data naming convention (Table 3).

Table 3 - Subject names and abbreviations

In KazakhIn EnglishAbbreviation
Humanities
Археология және этнологияArcheology and ethnologyAE
ТарихHistoryHIST
ИсламтануIslamic studiesISS
Мұражай ісі және ескерткіштерді қорғауMuseum work and protection of monumentsMWPM
ШығыстануOriental studiesORTS
ФилологияPhilologyPHILG
ФилософияPhilosophyPHIL
ДінтануReligious studiesRELS
ТеологияTheologyTHEOL
ТүркітануTurkic studiesTURKS
Languages and Literature
ЛингвистикаLinguisticsLING
ӘдебиеттануLiteratureLITER

Certain criteria were established for collecting and selecting data for the academic corpus. These criteria were comprehensively thought out and established during the planning and design of the corpus. According to the first criterion, the data for the academic corpus should be collected from domestic scientific research publications. That is, scientific works published in Kazakh in foreign publications were not selected. The second criterion is that the data should be written works of native speakers of Kazakh. For example, translated versions of abstracts were not collected since, taking into account that the language of the texts of such abstracts may be a translation language that is not written in the native Kazakh language (if the article is written in Russian or a foreign language, the abstracts are translated into Kazakh), they were not selected for the corpus. Another criterion is that the journals in which the scientific articles were published must be on the List of Publications approved by the Science Committee of the Ministry of Science and Higher Education. Articles published in scientific, popular science and other journals that were not included in this list were not included in the selection. Similarly, monographs, textbooks and teaching aids were taken into account, published by decisions of the Educational and Methodological or Scientific Councils of higher education institutions and scientific research centres and by the publishing houses of higher education institutions in Kazakhstan. Monographs, textbooks, teaching aids, dissertations and annotations were collected from the open worldwide web. Difficulties were encountered in collecting and processing texts from the first three types. Since these books are processed by publishers in a certain format, difficulties were encountered in converting printed texts into a machine-readable format. The next criterion is to pay close attention to the correct collection of data and their electronic, machine-readable format. It is important to remember that the correct formatting of texts affects the accurate and correct results of other linguistic and technical work performed on the basis of the corpus, for example, the development of keyword lists. The data collection process and its stages were described step-by-step in the Methodological Instructions developed within the framework of the project. Another criterion for data collection was that all publications should be dated between 2000 and 2024. The time period covering approximately two decades was closely related to the goal of the research project: the academic corpus aims to reflect the state of the modern Kazakh scientific language. At the same time, it allows for various activities based on the academic corpus, for example, researchers conducting scientific research by selecting academic texts published in a certain period or time interval.

Along with the data collection, processing was carried out. The data was translated from pdf and word formats to txt format. Later, the data was converted to Unicode UTF-8 format. Some images, tables, and other images in the texts were removed and displayed with annotations. Bibliographies and lists of literature in academic texts were removed.

The corpus was created using the balanced or selective construction approach. During the design and planning of the corpus creation, the volume of data was determined and approved. The data selection frame for the corpus took into account the equal or similar volume of texts collected from each discipline. It was also taken into account that the corpus represents a certain time period (2000-2024). If these conditions are consistent with the balanced or selective construction approach, we aim to increase the volume of the corpus in the future. In this regard, the possibility of using a monitor corpus approach is envisaged; that is, there is a possibility of supplementing the corpus with data over time.

Publications on the topic of the research project

Research articles:

Г.Ә. Сәрсеке. Корпус лингвистикасы терминдерін қазақ тілінде қалыптастыру және біріздендіру // Bulletin of Toraigyrov University. Philological series. - 2024. - № 3. - P. 322-333. https://doi.org/10.48081/ZRHT9356

Г.Ә. Сәрсеке, Б.Е. Акмагамбетова, Ж.Қ. Сентірбек. Академиялық корпус және оны қазақ тілінде әзірлеудің өзектілігі мен маңыздылығы // Bulletin of the Sh. Ualikhanov State University. Philology series. - 2025. - №2. - P. 213-227. DOI: 10.59102/kufil/2025/iss2pp213-227

Ж.М. Қоңыратбаева. Қазақ тілінің академиялық корпусын әзірлеу: мақсаты, міндеті, маңызы (гуманитарлық ғылымдар бағытында) // Bulletin of Buketov Karaganda University. Series «Philology». - 2025. - 3(119). - P. 213-227. https://doi.org/10.31489/2025Ph3/88-98

Г.Ә. Сәрсеке. Сигнал зат есімдер: академиялық қазақ тілінің корпусы негізінде зерттеу // Bulletin of the Kazakh State University of Philology named after Abylai Khan. Series «Philological Sciences». - 2025. - No. 4 (79). - P. 346-359. https://doi.org/10.48371/PHILS.2025.4.79.024

Г.Ә. Сәрсеке. Академиялық дискурстағы деп аталатын синтаксистік құрылымды корпус негізінде талдау // Bulletin of Buketov Karaganda University. «Philology» series. - 2026. - 1(121). - P. 59-68. https://doi.org/10.31489/2026PHI1/59-68

Г.Ә. Сәрсеке, Б. Е. Акмагамбетова. Академиялық дискурс: есімдіктердің қолданыс ерекшеліктері // Bulletin of Toraigyrov University. Philological series. - 2026. - № 1. - P. 487-504. https://doi.org/10.48081/BGQF1769

Scientific reports:

Г.Ә. Сәрсеке, А.Е. Хопур, Ж.Қ. Сентірбек. Корпус лингвистикасын пән ретінде айқындау мәселесі / «Қазақ тіл біліміндегі ғылыми зерттеудің бағыттары, әдіснамасы мен әдістемесі». Materials of the Republican Scientific and Practical Conference dedicated to the 95th anniversary of the birth of Gabdolla Kaliyev, scientist, doctor of philological sciences, professor, academician of the Academy of Social Sciences of the Republic of Kazakhstan. - Almaty: Abai Kazakh National University of Education, «Ulagat» publishing house, 2024. - P. 85-87.

Sarseke G.A, A.E. Khopur, M.Sh. Kurmanali. Design and development of an academic written corpus of Kazakh: challenges and potential perspectives / Moving Central Asian Studies ever further: Orthodox vs Unorthodox approaches / European Society for Central Asian Studies (ESCAS) conference, June 12-15 2025, Tashkent, Samarkand, Uzbekistan.

Textbooks:

G.A. Sarseke. Corpus linguistics: textbook. Astana: «Bulatov A.Zh.» Publishing House, 2026. - 238 p.

Dictionaries:

Cəрсеке Г.Ə. Sarseke G.A. Академиялық кілт сөздер тізімдері: Академиялық жазбаша қазақ тілінің корпусы | Academic Key Word Lists: Academic Written Kazakh Corpus. Astana: «Булатов А.Ж.» ЖК, 2026. - 132 б. (p.)

Copyright certificates:

CERTIFICATE on entering information into the state register of rights to copyrighted objects

December 9, 2025 No. 65097. Сәрсеке Гүлнәр Әдебиетқызы. Академиялық кілт сөздер тізімдері: Академиялық жазбаша қазақ тілінің корпусы (АЖҚК). Сөздік / Academic Key Word Lists: Academic Written Kazakh Corpus (AWKC). - Dictionary. Объектіні жасаған күні: 02.12.2025

Scroll to Top