Mechanism Type 2: Community Data and Open Source AI Ecosystems
> 3.2. Community Data and Open Source AI Ecosystems
As AI systems become central to communication, translation and information access, the question of who builds them and whose data they reflect has become increasingly important. Community data and open source AI ecosystems are emerging as an alternative to large commercial or state-controlled models. These efforts aim to ensure that languages, cultures and communities that are usually underrepresented in mainstream datasets have a direct role in shaping the systems that will affect them.
One of the strongest examples is Masakhane, a pan African grassroots research community working on natural language processing for African languages. Its work is entirely open, community driven and based on the idea that people who speak a language should guide how it is represented in AI systems. Masakhane has produced datasets, translation models and research that would not exist otherwise. Link: https://www.masakhane.io
In Latin America, initiatives such as LatinX NLP (https://latinxnlp.org) and regional open data collectives have taken similar approaches, creating corpora and models that reflect the linguistic diversity of the region. These projects operate with strong volunteer energy and are filling gaps left by commercial AI models that largely prioritise English and a small set of global languages.
Several countries are exploring more structured approaches. Brazil has begun investing in sovereign AI models developed with public research institutions, aiming to ensure that national datasets and public sector use cases are not entirely dependent on foreign providers. Kenya, Rwanda and the GovStack community (https://www.govstack.global) are developing open, modular digital infrastructure components that can support public services, including AI-enabled systems, with shared governance and local oversight.
The open source AI ecosystem has also expanded rapidly. Projects like Hugging Face (https://huggingface.co) have become central hubs where models, datasets and evaluation tools are shared openly. Although not all models are community governed, the ecosystem allows civil society groups, researchers and local institutions to experiment, audit and adapt AI systems without relying exclusively on proprietary services.
These examples show a mechanism where data and model building are shaped by the communities that need them, rather than extracted from them. Community data ecosystems do not solve all the challenges of AI, and they often lack stable funding, but they offer a path toward systems that reflect linguistic diversity, regional realities and public interest needs. This is especially relevant for civil society groups working in contexts where dominant AI models fail or where local languages and knowledge systems are poorly represented.