What is Noor Program and how does it bring conversational voice to under-represented languages?

Noor Program is yeshcube’s applied research line dedicated to developing voice synthesis and cultural adaptation for languages with a limited presence in today’s speech technologies.
The program works initially with Darija, Tarifit, Wolof, Bambara and Pulaar. Its aim is to build the capabilities these languages need in order to be used in conversational artificial intelligence systems with a voice that is intelligible, natural and appropriate to the cultural context of its speakers.
What it means for a language to be under-represented
A language can have a large community of speakers and, at the same time, very few computational resources.
Speech synthesis systems, known as TTS after text-to-speech, need datasets that pair texts with voice recordings. Those corpora have to contain enough phonetic diversity, consistent pronunciations and adequate samples of rhythm, intonation and linguistic variation.
Languages with a greater digital presence have large amounts of data and processing tools. In others, the corpora are small, their phonetic coverage is incomplete, or they were not created to train conversational systems. Recent research confirms that synthesis quality is still uneven across languages, and that general multilingual models do not deliver uniform results in intelligibility, naturalness and use outside the training domain.
This lack of linguistic infrastructure limits access to voice assistants, educational tools, guidance services and other applications in which conversation is the main interface.
Developing voice synthesis models
Noor’s first line of work focuses on building TTS capabilities for the selected languages.
The process covers the collection, generation and curation of voice data, the training of the models and their evaluation with native speakers. Testing has to look at several dimensions:
- Intelligibility of the words generated.
- Naturalness of pronunciation.
- Prosody and rhythm of speech.
- Stability across sentences of different lengths.
- Latency during a conversational interaction.
- Fit between register and context of use.
A synthetic voice can pronounce isolated words correctly and still be poorly suited to conversation. Intonation, pauses and rhythm shape comprehension and the perception of naturalness, especially when the system is conveying sensitive information or accompanying a long interaction.
Linguistic and cultural adaptation
The second line studies how to adapt conversational behavior to each language community.
Literally translating a system designed in another language can introduce awkward expressions, overly formal registers or inappropriate ways of asking questions. It can also lose cultural meanings carried by metaphors, forms of politeness and ways of expressing uncertainty, distress or trust.
Noor develops contextualisation frameworks to define vocabulary, idiomatic constructions, levels of formality and interaction criteria. This work requires native speakers to take part both in preparing the data and in evaluating the responses and voices generated.
Human validation is necessary because automated metrics can detect acoustic or linguistic errors, but on their own they cannot establish whether an interaction is culturally coherent.
Integration into conversational AI systems
The capabilities developed by Noor will be able to be incorporated into Somia and into other conversational solutions, whether our own or third parties’.
The program is initially aimed at contexts where language can become a barrier to access, such as guidance in reception centers, educational integration and services provided by public bodies to people with different mother tongues.
The scope of each integration will depend on how far the language has been developed, the quality reached by the voice model and the validation carried out for the specific context. Having a working synthesiser is a first technical layer. Incorporating it into a service also requires checking the comprehension, acceptance and cultural fit of the interaction.
Research built with language communities
Developing technologies for under-represented languages requires expertise in computational linguistics, speech processing and machine learning. It also requires working directly with the communities that use those languages.
Noor Program is developed within Allies and looks for the participation of specialist scientific teams, organizations with access to speaker communities, and technology partners able to contribute to creating and evaluating voice models.
The purpose is to widen the coverage of conversational artificial intelligence without treating languages as simple translation variants. Each one needs its own data, its own linguistic evaluation and an adaptation that respects how its community speaks and communicates.
Other programs


