yeshcube opens the Noor program to build conversational voice capabilities in under-represented languages

yeshcube is opening the Noor program, an applied research line focused on developing voice synthesis and cultural contextualisation capabilities for languages that lack coverage in today’s commercial TTS systems.
The program addresses a documented technical problem: the absence of working voice models in languages spoken by populations with high migratory mobility, which prevents conversational AI systems from being deployed in the very contexts where they could be most useful.
Technical diagnosis
Commercially available speech synthesis (TTS) systems concentrate their development on languages with large user bases and abundant training data. Languages such as Darija (Moroccan Arabic), Tarifit, Wolof, Bambara and Pulaar fall outside that coverage despite being spoken by millions of people.
That limitation has direct consequences for the deployment of conversational technologies. A system for emotional support or initial guidance that operates only in vehicular languages (Spanish, French, English, Modern Standard Arabic) does not reach users whose mother tongue is another. The barrier lies in the linguistic availability of the interaction layer. Access to the device does not explain it.
The problem has two distinct dimensions. The first is strictly technical: there are no TTS models trained for these languages to a quality sufficient for conversational interaction. The second is contextualisation: even with voice synthesis available, a conversational system requires cultural adaptation in expressions, registers, ways of approaching sensitive content and interaction protocols.
Scope of the program
The Noor program structures its work into two parallel research lines.
Developing voice synthesis models. Creating TTS voices in Darija, Wolof, Bambara, Tarifit and Pulaar with parameters optimized for conversational interaction: reduced latency, prosodic naturalness and register adapted for support contexts. This line includes generating and curating training datasets specific to each language.
Cultural contextualisation frameworks. Developing adaptation protocols that let conversational systems operate with cultural coherence. This covers identifying idiomatic expressions, adapting communicative structures for sensitive information, defining registers appropriate to the context, and validation with native speakers from each language community.
Application in the Somia ecosystem
The capabilities developed in the Noor program are integrated into the architecture of the Somia ecosystem. The voice models and contextualisation frameworks will make it possible to extend emotional support and guidance functionality to new languages while maintaining the ecosystem’s design, validation and certification standards.
Priority application contexts include initial guidance systems in reception centers, integration assistants in educational settings and support tools for public services attending to a linguistically diverse population.
The resulting solutions will follow yeshcube’s ERL validation methodology and, once sufficiently mature, will be able to obtain Somia Within certification for transfer.
Technical and scientific collaboration
The Noor program is developed within the Allies system. The project’s technical complexity and linguistic specificity require collaboration with institutions specializing in computational linguistics, natural language processing for low-resource languages, and organizations with access to native speaker communities for validation and data curation.
Lines of collaboration are open for scientific partners with capabilities in NLP and voice synthesis, development partners with complementary technologies, and transfer partners present in deployment contexts.
Does your organization have capabilities relevant to the Noor program?


