– The estimated reading time is 7 min.
Author
Much of what AI knows about the world comes from data most people never see and from technical processes that determine what’s included or left out.
As AI becomes increasingly influential in daily life, a new platform aims to help communities participate in that foundation being built behind the scenes to ensure they’re portrayed accurately and fairly, especially in imagery. Called the Community Library Creator, the Microsoft research effort gives groups a structured way to help shape how they show up in AI-generated pictures.
“There isn’t a ground truth for how people should be represented,” says Anja Thieme, a principal researcher at Microsoft with a background in social psychology and human-computer interaction. “It needs to be collectively defined and negotiated.”
Communities are building their own libraries instead of relying on patchy data that exists online, using the new tool to help them define what good representation looks like and build the data around it in a way AI models and systems can learn from to better reflect their existence.
What is a community library?
A community library is a collection of images or videos created by a group of people with shared experiences — such as a disability, identity or lived perspective — working together through advocacy organizations to show how they want to be represented in AI. Each image is paired with descriptions that explain what matters about it, adding context that helps AI systems understand not just what something looks like, but what it represents from the community’s perspective.
The Community Library Creator software, developed by Thieme and her team alongside Microsoft’s Accessibility Team, provides the tools and structure to help guide that process. But the choices about what to include, emphasize or change come from the community itself. The result is a resource that can be used to help AI systems learn from more diverse lived experience, not just patterns in data scraped from the internet.
Why do we need community libraries?
How people are portrayed shapes how they’re understood in the world and what opportunities are available to them. And as AI-generated images become more common, those portrayals are increasingly set by these systems.
AI models learn patterns from large collections of images and text, and they can repeat whatever gaps or distortions exist in that data. The plethora of images online of mystical dwarves, for example, can lead AI systems to generate people with dwarfism having pointy ears or other fantasy features, Thieme says. And people with limb differences are depicted online primarily in medical or athletic settings, meaning AI doesn’t have the context needed to show them in other typical situations, like working or meeting for coffee with friends.
Instead of leaving the training to engineers or to whatever exists on the internet, these libraries give communities a way to actively shape both the data and the evaluation standards that direct AI development. That matters because representation isn’t just technical, Thieme says, but needs to be defined by the people it affects, drawing on lived experience others can’t fully replicate.
How do community libraries actually work?
Community libraries can be built through the structured process Microsoft researchers created for the platform, drawing on inclusive design principles to help make something abstract — representation — easier to define.
It starts simply: Community members choose a small set of images that feel meaningful and explain why using the Community Library Creator’s step-by-step method. Those reflections help them zero in on key themes — like family life, work or everyday routines — that represent their shared experiences.
From there, they curate a larger collection of images or videos organized around those themes, adding detailed descriptions that explain what’s important in each scene. A community of Black people with albinism, for example, might highlight elements like wearing hats for sun protection or sitting close to their work due to the vision disabilities common with their condition — details that help AI systems better understand their daily lives and what makes their community’s experiences unique. Each library aims to gather about 400 “real-world” images from community members, capturing a range of experiences with a focus on quality over quantity, Thieme says.
Those images and annotations become training material. Prompts generated from the library are used to create AI images, which community members review and rate based on how well they match their desired representation. Over time, those ratings create a feedback loop, giving AI systems a clearer sense of what “good” looks like, as defined by the community, Thieme says.
Who owns the data — and what happens once a community library is built?
A key part of the model is that the community itself, through the advocacy organization that built the library, owns it and decides if and how the data is shared, including whether to make it available to researchers and developers on a platform such as Hugging Face, or to place limits on its use.
That also means the community retains oversight. The images are collected with consent, and if someone wants their data removed later, the community can make that change. Enabling that kind of control isn’t typical in AI today, Thieme says, and required working through contracts and internal processes to ensure it was possible.
In a space where much of today’s AI data is scraped from immense amounts of information online, with little visibility into its origins and no way of tracking it, that approach puts control back with the people represented — not just in how they’re portrayed but in how their data is used.
“It’s putting people and their data and their rights first,” Thieme says, “and I think we need to see this a lot more in general in the AI space.”
What’s next with these libraries? How could they shape the future of AI?
For now, the Community Library Creator is a tool being used in a controlled way with specific advocacy organizations. Thieme and Cecily Morrison co-lead a team — including engineers and researchers in areas such as design, accessibility, machine learning and human-computer interaction — working on the engineering, safety and legal controls needed to expand to more communities. The libraries are a way to widen who gets to shape AI in the first place — not just the systems themselves, but the data and criteria used to define success behind them. As the work evolves, the libraries could involve more communities and help train and evaluate future AI models.
“We’re really trying to widen participation in AI and who gets to have a voice in shaping the future,” Thieme says. “We’re opening ways and providing tooling for others to have an opportunity to create the future of AI. That’s what, for me, is the most important part.”
Learn more about how communities are helping shape AI: Why better AI starts with the people it often misses
*Lead image: Photos from community libraries created by advocacy organizations working with Microsoft’s Community Library Creator tool. The platform helps communities define how they want to be represented and contribute training data for AI systems. Photos provided by Black Albinism, Kilimanjaro Blind Trust Africa (KBTA), Ottobock and Short Stature Society Kenya (SSSK). *
Susanna Ray writes about AI and technology, with stories that show its real world impact and examine how innovation is reshaping work, business and society. She previously reported for Bloomberg News and other major international news organizations in the U.S. and abroad, covering beats ranging from politics and government to business and aviation. Follow her work on Microsoft Source.