The development of the Large Lithuanian Language Speech Corpus LIEPA-3 has been completed in Lithuania. Researchers from Vilnius University, Vytautas Magnus University, and the Institute of the Lithuanian Language have collected and annotated 10,000 hours of Lithuanian speech recordings, equivalent to more than one year of continuous speech. It is the largest Lithuanian speech dataset ever created for artificial intelligence technologies.
Why is such a corpus needed?
Modern artificial intelligence systems – from voice assistants, automatic subtitles to chatbots – rely on large volumes of high-quality speech data to function effectively.
While major world languages already benefit from extensive speech resources, Lithuanian has until now lacked datasets of this scale.

Professor Dr Gražina Korvel, VU MIF
As VU Faculty of Mathematics and Informatics (VU MIF) Professor Dr Gražina Korvel, project lead of LIEPA-3, notes, although technology continues to advance rapidly every year, Lithuanian language still often does not work properly in these systems, or performs less effective than one would expect: “The reason is simple – artificial intelligence still lacks sufficient examples of Lithuanian speech from which it could learn to understand living, authentic language as it is used in everyday life.”
LIEPA-3 was therefore developed to provide this large-scale speech data foundation for modern AI technologies.
Authentic Lithuanian Speech Recorded
LIEPA-3 stands out not only for its scale but also for its diversity. The corpus includes spontaneous, read, and dialectal speech collected from radio programmes, telephone conversations, publicly available recordings, and texts recorded specifically for the project.

Gediminas Navickas, VU MIF
VU MIF lecturer Gediminas Navickas, a project expert and head of the VU part of LIEPA-3, emphasises: “A large part of the spontaneous speech corpus would not have been possible without the cooperation of our media partners. We are grateful to the Lithuanian public radio LRT, radio station „Žinių Radijas“, and Martynas Mažvydas National Library of Lithuania for granting access to recordings from their audio archives. This partnership enabled us to compile valuable Lithuanian speech material and significantly strengthened national language technology resources.”

Professor Dr Gailius Raškinis, VMU
The importance of data diversity is further underlined by Vytautas Magnus University (VMU) Professor Dr Gailius Raškinis, a project expert, who remarks: “When compiling the read speech component of the corpus, phonetic diversity was ensured through the use of computational algorithms that selected a range of texts as wide and varied as possible.”
To ensure the corpus reflects contemporary Lithuanian language as it is actually spoken – including variation in speakers voices, speaking styles, regional pronunciation features, age groups, speaking rates, recording devices, and acoustic environments – broad public participation was essential.

Professor Dr Daiva Vitkutė-Adžgauskienė, VMU
More than 7,000 Lithuanian residents contributed voice recordings, and VMU Professor Dr Daiva Vitkutė-Adžgauskienė, head of the VMU part of the project, points out: “We are grateful to companies Gooliver and Lucid Agreements, as well as their business partners, for their efforts in collecting read speech recordings across all municipalities of Lithuania and ensuring the representativeness of the recordings across all these dimensions. In total, more than 7,000 Lithuanian residents contributed their voice recordings to the LIEPA-3 corpus.”
A Dedicated Section for Lithuanian Dialects
From a linguistic perspective, Lithuania is highly diverse in dialects.
Institute of the Lithuanian Language (LKI) Professor Dr Habil. Danguolė Mikulėnienė, head of the LKI part of the project, observes: “It is easy to notice that local people speak differently, for example, in Alytus region than they do around Utena, Telšiai, or Mažeikiai. For this reason, it was important to supplement the LIEPA-3 corpus with speech recordings reflecting dialectal features.”

Professor Dr Habil. Danguolė Mikulėnienė, LKI
She further highlights that the systematically collected and annotated 100 hours of dialectal material document the state of Lithuanian regional varieties in the third decade of the twenty-first century and allow researchers to trace emerging regional forms as well as long-term linguistic developments: “These recordings enable linguists not only to observe the emergence of new dialectal – or perhaps simply regionally distinctive – forms of Lithuanian, but also to anticipate possible long-term development trends. The expanded range of Lithuanian speech represented through dialect recordings will benefit everyone concerned with the long-term sustainability of the Lithuanian language.”
Audio Recordings Alone Are Not Enough
For artificial intelligence systems to learn language, recordings alone are not sufficient. They must be transcribed into text and aligned with precise time boundaries for each phrase.

Data annotation process
All recordings in the LIEPA-3 corpus were annotated at the phrase level, while 500 hours were additionally annotated at the lexical-unit and phoneme levels, enabling the development and training of advanced Lithuanian speech recognition technologies.
Collaboration Between Computer Scientists and Philologists
Beyond the technical achievement, one of the most important outcomes of the project has been close collaboration between computer scientists and philologists, strengthening interdisciplinary research.
As Vilnius University Faculty of Philology Professor Dr Vytautas Kardelis states: “The work carried out, and the results achieved during the project demonstrate that a large speech corpus is needed not only for speech technologies. It is also extremely important for understanding contemporary Lithuanian. The volume of material alone is not what matters most. More important is that computer scientists and philologists involved in the project discovered how knowledge from their respective fields can be combined and applied not only to speech technologies but also to linguistic research.”

Professor Dr Vytautas Kardelis, VU FLF
He further reflects: “This is not only about the tools that colleagues in computer science can develop for language analysis. It is also about verifying long-standing linguistic hypotheses and developing new theoretical approaches. Such cooperation and the advancement of interdisciplinary research are, in my view, among the most important directions that linguistics should pursue.”
Freely Available to Everyone
The LIEPA-3 corpus has been published under an open licence and is freely available to researchers, universities, companies, and developers working on Lithuanian-language AI solutions.
It is hosted in the CLARIN-LT open language resources repository ( https://hdl.handle.net/20.500.11821/101 ) and is also available through Lithuania’s open data portal data.gov.lt
As Professor Dr Gražina Korvel concludes: “LIEPA-3 is not a final product but a foundation upon which Lithuanian-language artificial intelligence solutions will be built.”
Beyond its technological role, the corpus will also serve as a valuable resource for scientific research supporting studies in linguistics, artificial intelligence, and digital technologies, while also enabling analysis of how Lithuanian is used across different regions and generations.
12 June 2026