Current AI and the United Nations Development Programme (UNDP) have collaborated on a new brief titled Introducing the Data to AI Value Chain: Promoting Linguistic and Cultural Diversity for AI Diffusion. The brief examines why so many efforts to develop AI for local languages end before they produce tools people can use. It also outlines what governments and public institutions can do to move promising work into public use.
The analysis draws on two years of UNDP's Local Language Accelerator in 11 countries, UNDP's AI Landscape Assessments across 26 countries, and Current AI's funding for cultural preservation projects in Africa, Latin America, and the Middle East. The brief's central finding is that collecting language data alone does not lead to AI that communities can use, shape, and sustain. UNDP's Data to AI Value Chain maps what must happen from data collection through deployment and reuse. Governments and funders must support the entire process, while communities help determine what is built, retain authority over how their knowledge is used, and share in the value created.
The language gap keeps widening.
AI research and funding are concentrated in languages that already have large digital datasets and many models. That work produces more data and models for the same languages, widening the gap with languages that receive little investment. English gains up to 50,000 new models a year, while languages such as Congolese-Swahili, Wolaytta, Kwangali, and Kuanyama gain an average of only 4.15 new models a year.
In Kisumu, Kenya, a small business owner cannot access a loan application or customer service tool in Dholuo because no one has built one. The technology exists, but the dominant commercial model gives companies little reason to build it. AI services divide text into units called tokens, and a word that counts as one token in English may count as four or five in Dholuo. Because providers charge by token, every Dholuo request costs more.
Coordination cannot change how current models process Dholuo, but governments can address other barriers. They can lower development costs by sharing computing infrastructure, create demand by purchasing tools for public services, and keep those services running through operating budgets. Yet many grants cover data collection and end before the organizations they fund can prepare the data, train and test a model, deploy a service, and maintain it.
Preparing data for model training depends on annotators who label and transcribe it. Their work is essential, but they are often underpaid and invisible within the supply chains of the institutions that benefit from their labor.
Fair exchange is a precondition for scale. Annotators must be paid fairly, and communities that contribute recordings and cultural knowledge must retain authority over how those contributions are used and share in the value they create.
Without fair exchange, power remains with frontier model companies, cloud providers, and platform owners. They decide what gets built, how it is deployed, and who captures the value.
Multilingual AI is therefore a concrete test of AI sovereignty. Communities have agency when they can govern their knowledge and shape the systems built from it. Countries have sovereignty when they can choose, inspect, adapt, or replace the infrastructure they depend on.
From policy to practice: two new frameworks
The policy brief, co-published this month by Current AI and UNDP, introduces two new frameworks that can be read alongside each other.
UNDP's Data to AI Value Chain provides a system-wide view of what must happen from data collection through deployment and reuse. It shows governments and funders where support is needed and where progress can break down.
Current AI's Cultural Preservation Pipeline shows what this work requires when AI is used to preserve and carry cultural knowledge forward. It begins with community governance and keeps communities involved as their knowledge is collected, digitized, used to train models, tested for cultural accuracy, and turned into tools. The people who hold that knowledge decide what may be shared and what must remain protected.

At Current AI, we are putting this approach into practice through our pilot cultural preservation grant cohort, which supports four organizations across sub-Saharan Africa, the Arab region, and the Brazilian Amazon with $3.2 million. Portal sem Porteiras works with Indigenous communities in the Brazilian Amazon and Cerrado to build offline-first, locally hosted tools. Communities decide which tools to build and how to use them, while the data remains within the territory.
Build together without surrendering control.
No country can provide all the data, models, testing methods, and computing resources required to build and maintain AI across thousands of languages. Countries can share infrastructure and retain sovereignty when no single participant owns the whole system or can withdraw everyone else's access.
The AI Potluck is Current AI's model for that form of shared ownership. Governments, organizations, and builders contribute what they do best without handing the result to a single company or country.
Earlier this week, Current AI joined 59 other organizations in a five-year commitment intended to help an estimated 3.4 billion people use AI in their own language and voice. Reaching them will require every part of the Data to AI Value Chain, with communities and public institutions shaping what is built and retaining control over what they contribute.
This commitment will matter if communities retain authority over the knowledge they contribute and share in the value it creates, while countries can choose, inspect, adapt, or replace the infrastructure they depend on.

