A growing pushback against big tech companies, and greater awareness of the value of data, is spurring interest in data collectives and cooperatives, which give communities control over the collection, management, and distribution of their data. This alternative allows creators to benefit from data sets that may otherwise be ignored or misused.
A handful of tech companies dominate the generative artificial intelligence industry, with American and Chinese frontier models controlling the lion’s share of the market. Some countries are building their own large language models because their language and culture are not adequately represented in GPT, Gemini, Claude, or Qwen.
Communities that possess smaller or unusual data sets can gain from having control over them, Raffi Krikorian, chief technology officer at Mozilla Foundation, told Rest of World. The nonprofit last year set up Mozilla Data Collective to provide a platform for such data sets from communities, organizations, and individuals around the world.
“More people are starting to feel like they’re sitting on unique information, not generally available on the internet, and they want to turn the tables on what governance looks like for that data,” Krikorian said.
“The anti-Big AI, anti-Big Tech push is a convenient bedfellow. The big companies have built themselves up on the backs of all these people creating data, who think it’s time to set their own terms now,” he said.
A blueprint for responsible use #
Companies including Meta, Open AI, Google, and Anthropic have scraped nearly all available data from the internet to train their AI models, and have been accused of using copyrighted material without consent. Tech companies have said it qualifies as fair use, which allows the use of such material for research and other purposes. Some countries are trying to balance the need for good data with the need to compensate creators.
Data collectives, or cooperatives, offer a blueprint for the responsible use of data, and ensure the value goes to those generating the data, Astha Kapoor, co-founder and director of Aapti Institute, a tech research firm in India, told Rest of World.
“Beyond consent and compensation, collective action around data gives communities the opportunity to direct data towards issues they may care about,” she said. “Communities can negotiate the terms on which their data is used at every stage of the AI lifecycle [with] mechanisms for accountability and redressal, in case their terms are breached.”
Workers, producers, consumers, and others have been establishing cooperatives and other community-led associations to pool resources, share benefits, and address socioeconomic challenges for centuries. The United Nations marked 2025 as the year of cooperatives, positioning them as “essential solutions to today’s global problems,” kindling renewed interest in data collectives and cooperatives.
They cover a wide range: The Kerala Food Platform enables about 2,500 farmers in the southern Indian state to trace and market their produce, including rice, fish, fruits, and vegetables. Mexico-based PescaData helps small-scale fishers in Latin America and the Caribbean to manage and benefit from their catch records, while the Native BioData Consortium is a repository of the genetic and environmental data of Indigenous people.
Increasingly, data sets are being created for AI-related purposes, including in low-resource language communities, whose data, including voice data, is valuable for training small models and creating speech recognition tools. For these communities, a data collective is a more practical solution, Krikorian said.
“An OpenAI or Anthropic is not going to prioritize a language that’s only spoken by 1 million people,” he said. “But if it existed, a chatbot that can communicate in their language is hugely beneficial to that community.”
“Linguistic identity crisis” #
Long before the launch of ChatGPT, Meesum Alam realized that dozens of languages were dying in his native Pakistan. He belonged to the Baloch community but could not speak Balochi, the language of his forefathers, and experienced a “linguistic identity crisis,” he told Rest of World.
Alam began to document languages that were at risk of dying because few people spoke them. He began with Dawoodi, which had about 300 speakers, and Kalasha, with some 3,000 speakers, and collected voice data in 39 languages from communities and local organizations in Pakistan. He put the data sets, totaling about 700 hours, on Mozilla Data Collective, where they have been used by companies including Meta to build speech recognition tools, he said.
“These languages were never part of the digital world,” said Alam, a Ph.D. candidate in computational linguistics at Indiana University**. **“For the communities, being able to interact with AI in their own language is a first step into the AI world.”
The communities decided their data sets can only be used for research or non-commercial purposes. Meta and other big companies have to negotiate the terms of use with the communities, Alam said. Money is not a motivating factor, he said.
“They don’t trust the big tech companies because they know they can take the data and monetize it,” he said. “They want a fair deal for the entire community, and they can only get that through a data collective. It gives some power back to communities.”
Besides collectives, other frameworks in use include data trusts, where a trustee manages data on behalf of a group; data unions, which aggregate the data of individual members to negotiate collectively with buyers; data commons, such as Wikimedia and OpenStreetMap, which have different governance structures; and data donation schemes, such as the Personal Genome Project, where individuals contribute their data for public benefit.
“It’s a very emotional thing”
In the African continent, which has long been subject to extractive practices, the Nwulite Obodo Open Data License, launched in 2024, enables creators, communities, researchers, and others to share data sets without giving up the right to benefit from them. About 70 African data sets under NOODL are part of the Mozilla Data Collective, including speech data sets in more than 20 African languages, music, lullabies, and poetry.
These languages are not recognized in the mainstream linguistic framework, so having them in the data collective increases their visibility and accessibility, Emmanuel Ngue Um, a regional researcher at the Institute of African Digital Humanities, which has published about 40 data sets on Mozilla Data Collective, told Rest of World.
“Most importantly, they can require users to clarify the purpose of their access request,” he said. “This creates opportunities for cooperation and partnership between those with the technological ability to support language work in underserved communities, and the communities.”
Still, data collectives can face governance and scaling challenges, Kapoor said.
“Building data cooperatives solely to steward data is not feasible because sustainability becomes an issue, and the only viable pathway becomes monetization of the data the cooperative is meant to safeguard, which is problematic,” she said.
Earlier this month, the Mozilla Data Collective made three community-generated data sets available for paid commercial licensing, ahead of opening the compensation feature to all users so that communities creating the data are “valued, recognized and supported,” it said.
For Alam, who is on the hunt for low-resource languages in India and Bangladesh as well, the benefit is clear. AI adoption is growing quickly, and for many communities, having their language data sets easily accessible through a data collective is the only way they can participate. “Every other day, I get a text or a voice note from someone who is able to communicate with a chatbot in their own language for the first time,” he said. “It’s a very emotional thing for them when they can do that.”