Armenian on the Map of the Digital World

Committed to preserving and advancing Armenia's national heritage, the Izmirlian Foundation is supporting innovative efforts to strengthen Armenian language technologies and secure the language's place in the AI era.

  • July 30, 2026
  • Mediamax
  • 3 photo(s)

If you ask ChatGPT a question in Armenian today, chances are you will get an answer. It can translate, proofread text, draft an email, or perhaps even write a poem. At first glance, it may seem that Armenian has already secured its place in the world of artificial intelligence. Yet this impression can be misleading.

Solely the number of Armenian-language websites, digitized books, or online dictionaries does not measure a language’s digital presence. A more important question is: to what extent do machines actually understand Armenian? Can they accurately analyze sentence structure and understand context, distinguish the meanings of words, process millions of texts, recognize printed text, work with dialects and colloquial speech, produce high-quality translations, or become voice assistants that truly “think” in Armenian? 

None of this is possible without large-scale, high-quality linguistic data. Without such an infrastructure, even the most advanced artificial intelligence systems will remain limited in what they can do with Armenian.

To address this gap, the Armenian Language Technologies Research Laboratory programme has been operating at Yerevan State University since March 2025. The programme is funded by the Higher Education and Science Committee of the Republic of Armenia as part of the country's state language policy priorities. The Laboratory is also supported by the Izmirlian Charitable Foundation through its project “Armenian as a Viable Language in the Technological World.”

The project's goal is not limited to creating new digital tools. It seeks to build an ecosystem that will ensure Armenian remains a viable language with equal representation in artificial intelligence systems while expanding the use of those systems in areas such as general education, public services, public administration and beyond. In other words, the project's overarching objective is to ensure the digital vitality and digital equality of the Armenian language.

According to Siranush Dvoyan, Chair of the Language Committee of the Republic of Armenia, the future of a language can no longer be discussed solely in terms of language literacy. If people spend a significant part of their lives in digital environments, the language itself must also remain viable in the face of the technological world's challenges and meet emerging needs through high-quality tools. To illustrate what this means in practice, Dvoyan recalls a simple example she used for years with her students.

“I would ask them to open any search engine and type in the word 'table. At first, the result was usually the same. Search engines provided only very limited information about the word-perhaps identifying it as a noun and noting that its plural form was tables, and little else.”

Then, as Dvoyan recalls, at one point, that changed.

During one of these exercises, her students showed her that the search engine was now providing a much richer linguistic description of the word table. Dvoyan says this came as a welcome surprise because once a language is equipped with the necessary digital resources-digital dictionaries, annotated corpora and curated linguistic data-it begins to become “visible” to machines, and its capabilities in the digital world expand significantly.

According to Dvoyan, the rapid development of artificial intelligence in recent years has also changed the priorities of language policy. While electronic resources were once enough to improve communication and quality of life, today AI systems require large-scale, machine-readable linguistic data to work effectively with Armenian-language knowledge.

Marat Yavrumyan, Head of the Armenian Language Technologies Research Laboratory, notes that a significant share of Armenian-language content - whether written or spoken - remains inaccessible to machines today.

“The key is to close the gap accumulated over previous years as quickly as possible, without compromising on quality,” says Yavrumyan.

At the same time, large-scale, high-quality Armenian-language data are not a cure-all. What is needed is a comprehensive, continuously updated infrastructure-an ecosystem in which text, audio and visual data are filtered, standardized, licensed, annotated and made available in formats suitable for training modern AI models.

“Our experience in Armenian speech recognition, speech synthesis, and language model evaluation shows that even powerful multilingual models often fail to distinguish Standard Armenian from dialects or colloquial speech. They also struggle with proper names, terminology, and longer contexts because Armenian is underrepresented or unevenly represented in the datasets used to train such models,” says Karen Avetisyan, Scientific Lead for Artificial Intelligence at the Laboratory.

It is therefore important not only to rely on existing models, but also to develop open datasets, reproducible benchmarks, Armenian-specific tokenization methods, and other practical tools whose performance is measured against clear technical criteria rather than general impressions.

The Laboratory is currently developing an AI evaluation system for Armenian. 

“If we cannot measure it, we cannot tell whether a system is actually getting better at Armenian,” says Yavrumyan.

To that end, the Laboratory is developing around twenty benchmarks that will enable Armenian's level of inclusion and performance across different AI systems and large language models to be evaluated alongside those of other languages using the same criteria.

Another key area of work is advancing Armenian language technologies and digital humanities by developing the resources and machine-processing infrastructure they require. The chosen approach is to develop Optical Character Recognition (OCR) technology together with the supporting infrastructure that will make it possible to transform the entire body of written heritage ever created in Armenian-from the press and literature to manuscripts, archival materials and more, whether printed or handwritten-into high-quality, machine-readable data that can be accessed instantly.

According to Marat Yavrumyan, teaching machines to read Armenian is only the first step. For machines not only to recognize words but also to understand them and their context, specialized linguistic resources and tools are essential.

The same thinking lies behind the Laboratory's plan to develop a national Ngram Viewer platform. The platform will enable users to track how frequently particular words and expressions have been used over time and how their meanings and prevalence have evolved.

“This platform will provide a new impetus for the development of digital humanities within the Armenian-language context. One of the simplest examples is that it will allow us to examine, through objective statistical analysis, what social and cultural realities in Armenia the language has historically reflected, and how it has portrayed social history across different periods of time,” Yavrumyan explains.

As part of its work to advance digital humanities, the Laboratory has been developing the Armenian National Treebank project since 2017. Within the project, five openly accessible, morphologically and syntactically annotated treebank corpora have been created for Eastern Armenian, Western Armenian, and Middle Armenian. Together, they remain the only resources of their kind worldwide.

The project also envisages the development of a new digital dictionary of the Armenian language.

According to Siranush Dvoyan, this is not intended to be just another online dictionary.

“The explanatory dictionaries that are widely used today for academic and educational purposes were compiled in the previous millennium. They remain fundamental scholarly works. But the language has changed. New words and concepts have emerged, while existing words have acquired new meanings and new uses. Language is a living phenomenon. If we want Armenian to have a full-fledged presence in the digital world, we need a different kind of dictionary."

As our conversation draws to a close, I ask what has been the most challenging part of the project.

“There simply aren’t enough hours in the day,” Dvoyan replies without hesitation. 

“The team is so deeply engaged, and there is so much to be done, that time always seems to be in short supply."

Artificial intelligence can do remarkable things, however, in Armenian, only to the extent that people have taught it to understand Armenian. The next time ChatGPT gives you an excellent answer in Armenian, perhaps the credit should not go to the chatbot alone. Much of that answer is the result of invisible, everyday human work. 

Astghik Hovhannesov

Photos: Emin Aristakesyan

The Izmirlian Foundation is committed to advancing the sustainable development of Armenia, guided by its mission of “helping to preserve a nation.” Over more than 35 years of activity in Armenia, the Foundation has launched and implemented a wide range of initiatives with long term systematic impact across education, social welfare, healthcare, tourism, and economic development, as well as in culture, the Armenian language, artificial intelligence, high technology, and innovation.