Language Belongs to Its Speakers: Why African AI Must Be Built From the Ground Up

Tag: General news

Published On: September 18, 2026

A few days ago, Google published a piece by James Manyika, their SVP of Research, Labs, Technology and Society, called "AI for everyone in every language." It's about closing the gap between the languages AI understands well and the thousands of languages people actually speak.

Buried in that piece is a mention that means a lot to us at Digital Umuganda. Google names WAXAL, an open speech dataset covering 27 Sub-Saharan African languages spoken by more than 100 million people across 26 plus countries, and credits it to partners including Makerere University and Digital Umuganda.

I want to talk about why this matters, and why it's bigger than one dataset or one mention.

Most AI still doesn't understand most of Africa

Here's the uncomfortable truth. The internet, and by extension the data that trains most AI models, is built overwhelmingly around a small number of dominant languages. English, Mandarin, Spanish, French. If your language isn't well represented online, it's poorly represented in AI, or missing entirely.

That's not a small gap. It's hundreds of millions of people across this continent who cannot reliably speak to their phone, get health information, or use a digital service in the language they actually think and dream in. When AI fails to understand a language, it fails to understand the people who speak it.

Why the Google mention matters, and why it's not the point

For years, organizations like ours have argued that closing this gap requires community-based, grassroots data collection. Not scraping the web harder. Actually going into communities, working with speakers, capturing how language is really used, tone, code-switching, dialect, the whole texture of it.

Seeing Google build WAXAL this way, in partnership with African institutions rather than around them, is proof that this approach is now recognized by the biggest labs in the world as the right one. That validation matters. But I don't want to overstate our role in one dataset. What matters more is what this shift represents for the continent.

What this actually means for people

This isn't an abstract research problem. It's about whether a community health worker in rural Rwanda can get accurate guidance in Kinyarwanda during a disease outbreak. Whether a farmer in northern Ghana can ask a question about crop disease and get an answer in a language they speak at home, not just English or French. Whether a mother can navigate a health system without needing to find someone to translate for her.

The Manyika piece includes a good example of this working already. Viamo's "Ask Viamo Anything" service has been piloted in Rwanda, using Gemini to power a voice assistant that runs on basic feature phones, not smartphones. It's already answered more than 2 million questions. That detail matters because most of the continent doesn't access the internet through smartphones with reliable data. If AI only works on high-end devices, it simply doesn't reach most people.

Where Digital Umuganda fits

Our work has always started from the same place: African languages need African-led data collection, done properly, in partnership with the communities who speak those languages. We've built datasets, tools, and applications across Kinyarwanda and other African languages, with health, agriculture, and public services as the proving ground, because those are the areas where getting language right can genuinely change outcomes.

We're not claiming to have solved this. We're one part of a much bigger effort that needs many more hands.

The bigger point

Here's what I keep coming back to. It's not enough for African language data to simply feed AI models built elsewhere. The continent needs more African-led organizations doing this work, setting the standards, owning the datasets, and making sure the benefits of this technology come back to the communities that made it possible in the first place.

The infrastructure gap here is still massive. Most African languages remain under-resourced or entirely absent from AI systems. Closing that gap needs funders, researchers, governments, and technology partners willing to invest in African-led work, not just African data.

If you're working on language technology, health systems, or digital access in Africa and want to talk about partnership, I'd like to hear from you.