cd /news/artificial-intelligence/scraped-without-asking-indigenous-ar… · home topics artificial-intelligence article
[ARTICLE · art-82666] src=smarterarticles.co.uk ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Scraped Without Asking: Indigenous Archives and the Limits of AI Law

A June 2026 analysis by Kerri J. Malloy, assistant professor of Native American and Indigenous Studies at San José State University and a citizen of the Yurok and Karuk peoples, published in Governing, details how AI systems scrape Indigenous cultural materials from institutional archives without tribal authority, converting them into training data for outside purposes. At the 19th session of the Expert Mechanism on the Rights of Indigenous Peoples in Geneva from 13 to 17 July 2026, Indigenous advocates pressed for legal obligations requiring institutions to seek tribal consent before using such materials in AI, highlighting the case of 26 playable wax-cylinder recordings from 1890 held by the Library of Congress.

read29 min views1 publishedAug 1, 2026
Scraped Without Asking: Indigenous Archives and the Limits of AI Law
Image: Smarterarticles (auto-discovered)

Scraped Without Asking: Indigenous Archives and the Limits of AI Law #

In March 1890, an anthropologist named Jesse Walter Fewkes carried a wax-cylinder phonograph to Calais, Maine, and over three days recorded thirty-six cylinders of Passamaquoddy songs, creation stories, vocabulary and legend. The voices belonged mostly to two men, Peter Selmore and Newell Josephs. Fewkes was experimenting; he wanted to know whether the new machine could capture human speech in the field. The result is now understood to be the oldest surviving ethnographic field recording anywhere in the world. For more than a century those cylinders sat in institutional custody, first at the Peabody Museum of Archaeology and Ethnology at Harvard University and then, from 1970, inside the American Folklife Center at the Library of Congress, catalogued, preserved, and effectively unreachable by the community whose ancestors had sung into the horn. Of the original thirty-six, only twenty-six remain playable.

Nobody in 1890 asked Peter Selmore whether his voice could be digitised, indexed, transcribed by an algorithm, or fed into a statistical model that might one day predict the next word in a Passamaquoddy sentence. The question would have been unintelligible. The technologies that make it urgent did not exist, and the legal framework that might have answered it did not exist either. That gap between how the material was gathered and what can now be done with it is the precise location of a fight that reached the United Nations in the summer of 2026.

At the nineteenth session of the Expert Mechanism on the Rights of Indigenous Peoples, held at the Palais des Nations in Geneva from 13 to 17 July 2026 with a dedicated panel on artificial intelligence and the rights of Indigenous Peoples on its agenda, Indigenous advocates from several countries pressed a demand that a decade ago would have sounded speculative. They argued that artificial-intelligence systems are now being trained on, and deployed against, the vast holdings of Indigenous cultural material sitting in universities, museums, national libraries and government archives, and that the institutions holding those materials have no clear obligation to treat tribal authority over the knowledge as anything more than a courtesy. The oral histories, the language recordings, the ceremonial records, the photographs, the governance documents, much of it gathered under conditions that would fail any serious contemporary standard of informed consent, are being converted into training data and searchable outputs that serve outside purposes. The communities from which the knowledge originated are frequently the last to know.

The question at the centre of the Geneva session was not whether this is happening. It plainly is. The question was whether anyone is legally, ethically or procedurally required to stop it, or to ask first.

How an Archive Becomes a Model #

To understand what changed, it helps to be precise about the mechanism. A June 2026 analysis published in Governing by Kerri J. Malloy, an assistant professor of Native American and Indigenous Studies at San José State University and a citizen of the Yurok and Karuk peoples, laid out the sequence in unsentimental terms. AI systems scrape materials held in institutional archives and digital repositories without reference to tribal authority. Those materials are converted into training data or into searchable, generative outputs. The outputs serve the purposes of the party running the system, which is almost never the community whose knowledge was ingested. Malloy, whose scholarship centres on genocide, transitional justice and the mechanics of redress, frames this not as an accident of technology but as the latest expression of a much older pattern, in which knowledge is separated from the people who hold it and put to work elsewhere.

The point worth dwelling on is that the scraping requires no malice, or even any awareness that the material is Indigenous. A digitised photographic collection, a corpus of transcribed oral histories, a set of language recordings released under an open licence by a well-meaning library: to a web crawler assembling a training set, these are simply text, image and audio. The metadata that would flag a recording as ceremonial, restricted, seasonally sensitive, or governed by protocols about who may hear it and when, is either absent or discarded during ingestion. The model learns the patterns and forgets the provenance. What comes out the other side is a system that can generate plausible imitations of a cultural form, answer questions about restricted knowledge, or reconstruct fragments of a language, without any accountability toward the community.

This is the harm that data-sovereignty scholars have described for years under the heading of data colonialism, a term meant to make the analogy explicit: just as historical colonialism appropriated land and labour, the contemporary extraction of data and knowledge appropriates the raw material of culture and computation. The analogy is not rhetorical excess. The material sitting in these archives arrived there through the same institutions, and often the same expeditions, that removed ancestral remains and ceremonial objects. The wax cylinders and the funerary belongings travelled together. That the recordings can now be reanimated by machine learning does not sever them from that history. It extends it.

What Geneva Was Actually Arguing About #

The Expert Mechanism on the Rights of Indigenous Peoples is not a court, and it cannot compel anyone to do anything. Established by the Human Rights Council in 2007 under resolution 6/36, it is a body of seven independent experts that provides the Council with advice and expertise and assists states in achieving the goals of the United Nations Declaration on the Rights of Indigenous Peoples. Its authority is persuasive rather than coercive. But the instrument it exists to interpret carries more weight than its soft-law status suggests, because a great many of its provisions are now treated as reflecting customary international norms.

That instrument, the Declaration adopted by the General Assembly in 2007, contains in Article 31 a passage that reads with uncanny prescience given what has happened since. Indigenous peoples, it states, have the right to maintain, control, protect and develop their cultural heritage, traditional knowledge and traditional cultural expressions, as well as the manifestations of their sciences, technologies and cultures, including human and genetic resources, seeds, medicines, knowledge of the properties of fauna and flora, oral traditions, literatures, designs, sport and traditional games, and visual and performing arts. They also have, it continues, the right to maintain, control, protect and develop their intellectual property over such cultural heritage, traditional knowledge and traditional cultural expressions.

Read in 2007, Article 31 was understood mostly as a shield against biopiracy and the commercial appropriation of designs and medicines. Read in 2026, the phrase “maintain, control, protect and develop” runs directly into the architecture of machine learning. If a community has the right to control its traditional cultural expressions, and an AI company ingests a digitised archive of those expressions to train a commercial model, the community's control has been overridden without its involvement. The Declaration also insists, repeatedly, on the principle of free, prior and informed consent, the requirement that Indigenous peoples be consulted and give agreement before measures affecting them are taken. The Expert Mechanism has previously produced a dedicated study on the repatriation of ceremonial objects, human remains and intangible cultural heritage, explicitly bringing the intangible, the songs and stories and knowledge, within the frame of restitution. The 2026 advocates were extending that logic one step further, into the training corpus.

The gap the Geneva delegates were pointing at is the gap between principle and obligation. The Declaration says communities have the right to control. It does not say that a university digitising its holdings must obtain consent before a third party scrapes them, nor that a museum must embed enforceable restrictions in the metadata it publishes, nor that an AI developer must check provenance before ingestion. Those operational duties do not yet exist in most jurisdictions. The advocates wanted them written down.

The Principles Built Before the Machines Arrived #

What makes the current moment unusual is that Indigenous communities did not wait for the AI industry to arrive before building governance frameworks. The intellectual scaffolding was largely in place, developed through the 1990s and 2000s in the context of research ethics and health data, and it maps onto the machine-learning problem with remarkable directness.

The oldest of these frameworks is OCAP, the First Nations principles of Ownership, Control, Access and Possession, established in 1998 and now administered by the First Nations Information Governance Centre in Canada. OCAP holds that a First Nation collectively owns its information in the same way an individual owns personal information; that it may assert control over data at every stage of the research cycle; that it must be able to access information about itself regardless of who physically holds it; and that possession, the physical custody of data, is the mechanism that makes ownership more than symbolic. The last principle is the sharpest when applied to AI. Possession says that if you want to protect knowledge, you keep it where you can defend it. A model trained on a copy you no longer control is the negation of possession.

The more recent and internationally influential framework is the set of CARE Principles for Indigenous Data Governance, drafted at a workshop in Gaborone, Botswana, in November 2018 and published through the Global Indigenous Data Alliance. CARE stands for Collective Benefit, Authority to Control, Responsibility and Ethics, and it was written deliberately as a counterweight to the open-data movement's FAIR principles, which hold that data should be Findable, Accessible, Interoperable and Reusable. The tension between the two acronyms is the entire argument in miniature. FAIR optimises for sharing and reuse; it says nothing about power, history or consent. CARE was built to reinsert those considerations, to say that the ease with which data can be shared is not the same as the right to share it, and that governance must account for the power differentials that shaped how Indigenous data came to sit in institutional hands. The Authority to Control principle is unambiguous when applied to a training set: the authority to decide whether a corpus becomes model weights rests with the community, not the archive.

These frameworks share a feature that distinguishes them from most Western data-protection law. They treat knowledge as collective and relational rather than as individual property with a fixed author and an expiry date. Conventional intellectual-property regimes are built around individual ownership, novelty and a term that eventually lapses into the public domain. Indigenous knowledge is frequently held communally, transmitted orally across generations, and governed by protocols that have nothing to do with authorship and everything to do with relationship, season, initiation and place. When a song enters the public domain under copyright law, the community's protocols governing who may sing it do not lapse. The two systems are not merely different in detail; they are built on incompatible premises about what knowledge is and who it belongs to. The scraping of an archive collapses that incompatibility in favour of the system that ignores protocol.

The Labels That Travel With the Knowledge #

If the principles are the theory, a handful of practical tools have emerged to make them operational, and their design reveals how hard the problem actually is. The most widely adopted is the system of Traditional Knowledge and Biocultural Labels developed by Local Contexts, an organisation co-founded by the legal scholar Jane Anderson and the digital-humanities scholar Kim Christen, and now co-directed by Christen, Anderson, Māui Hudson of Whakatōhea and James Francis of the Penobscot Nation. The Labels are not licences in the copyright sense. They are metadata, attached to cultural material, that carry the community's own statements about provenance, protocol and permission: who the cultural authority is, what traditional protocols govern access, and what uses the community regards as acceptable. A Provenance Label identifies the group that holds authority over the material. A Protocol Label communicates the customary rules attached to it. A Permission Label states what the community has approved. The point is to make Indigenous authority legible inside the metadata of a digitised collection, so that a curator, a researcher, or in principle an automated system, encounters the community's terms at the moment of access rather than never.

The companion tool comes from the same hand: Mukurtu, a free, open-source content-management system first built in 2007 by Christen and the developer Craig Dietrich for the Warumungu Aboriginal community in central Australia, and now maintained at Washington State University. Mukurtu was designed around a premise that most archives find alien: that access should be differential rather than uniform. A single item in a Mukurtu archive can be visible to the general public in one form, to community members in another, to a particular family or ceremonial group in a third, and to no one outside a restricted circle at all. Where a conventional digital repository asks how to maximise open access, Mukurtu asks who is allowed to see what, and encodes the answer.

The Passamaquoddy cylinders became the demonstration case for all of this. When the Library of Congress launched its Ancestral Voices project, engineers at its National Audiovisual Conservation Center used an Archéophone playback machine and digital restoration systems to extract sound from the 1890 wax that had been physically unplayable, and then, crucially, handed curatorial control back toward the Tribe. Passamaquoddy speakers transcribed and translated the recordings; elders reviewed them; the material was described using Mukurtu and tagged with Local Contexts Traditional Knowledge Labels, so that the digitised voices now travel with the community's own statements of authority and protocol attached. It is the closest thing to a model of how digital repatriation can work when an institution chooses to share power rather than merely access.

But the Passamaquoddy case also exposes the limit of the whole apparatus, and it is a limit the Geneva advocates understood well. Labels and differential access work only inside systems that agree to honour them. A Traditional Knowledge Label is metadata; a web crawler assembling a training corpus is under no obligation to read it, and a large language model does not preserve it. The moment a labelled recording is copied outside the governed platform, whether by an open-data release, a partner institution's mirror, or a scraper that ignores robots directives, the protocol evaporates. The tools that Indigenous communities built to assert authority were designed for a world of human curators making deliberate choices about individual items. They were not designed for a world of automated ingestion at web scale, where the default is to take everything and ask nothing.

When Revitalisation and Extraction Use the Same Tool #

The uncomfortable truth threaded through the whole debate is that the technology now driving the extraction is the same technology offering some communities their best hope of linguistic survival. This is not a case where the harm and the benefit are cleanly separable. They run through the identical set of tools.

An analysis published in July 2026 by researchers at the University of Melbourne made the double edge explicit. The same AI systems that can support Indigenous language revitalisation, generating learning materials, reconstructing grammatical patterns from fragmentary historical records, filling gaps in documentation left by generations of suppression, can, if they are built and controlled by outside institutions, deepen the very patterns of extraction they appear to remedy. The mechanism is concentration of authority. A language model that becomes the authoritative interface to a language, trained on the community's own recordings but owned and operated by a university or a company, does not restore control. It relocates it. The community becomes a user of a system built from its own knowledge, dependent on an institution it cannot govern. The Melbourne researchers were echoing a broader body of work, including a systematic review by the same group published in 2025, that has repeatedly found the governance question, who owns and controls the system, to be more decisive than the technical question of whether the tool works. A separate 2026 review in AI & Society, examining AI projects in Irish Gaelic, Māori, Guaraní and Inuktitut through the lens of data colonialism, reached the same place by a different route: what divides the initiatives that empower communities from those that reproduce extractive structures is not the model but who holds the data and sets the terms.

The counter-example that everyone in this field cites is Te Hiku Media, a charitable media organisation based in Kaitaia, in the far north of New Zealand's North Island, and belonging collectively to the Far North iwi of Ngāti Kuri, Te Aupōuri, NgāiTakoto, Te Rarawa and Ngāti Kahu. Facing the same problem every Indigenous community faces, that the large technology firms had little commercial interest in a language spoken by a few hundred thousand people, Te Hiku built its own. Through a crowdsourcing campaign called Kōrero Māori, it gathered more than three hundred hours of labelled speech in ten days, from more than 2,500 people reading over 200,000 phrases, and in 2021 released an automatic speech-recognition model for te reo Māori that reportedly reached around ninety-two per cent accuracy, outperforming the offerings of far larger companies on the language. The decisive move was not technical but legal. Te Hiku placed the resulting data and models under what it calls the Kaitiakitanga Licence, a bespoke instrument built on the Māori concept of guardianship, which keeps data sovereignty inside the community, prohibits uses that would surveil or discriminate, and ensures the data is used for the benefit of Māori. The organisation refused, publicly and repeatedly, to hand its speech corpus to outside firms, on the grounds that the community had gathered the data as a taonga, a treasure held in trust, not as a commodity to be sold.

Te Hiku is the proof of concept for the Geneva argument, because it demonstrates that Indigenous-governed AI is not a contradiction in terms. The community built the tool, kept the data, wrote the licence, and set the terms of use. What Te Hiku had, that most communities do not, was the technical capacity, the funding and the pre-existing organisational strength to do all of that itself. The Melbourne researchers' warning is aimed at the far more common situation, in which the community lacks the capacity to build its own system and the institution holding the material builds one instead, positioning itself, however benevolently, as the permanent intermediary between a people and its own language.

Every strand of this returns to the conditions under which the material was collected in the first place, and here the historical record is not ambiguous. The great archives of Indigenous cultural material were assembled overwhelmingly during the late nineteenth and twentieth centuries, in a period when the collecting institutions operated on the salvage assumption, the belief that Indigenous peoples were vanishing and that their cultures had to be recorded before they disappeared. The people recorded were rarely in any position to refuse, frequently were not asked, and could not conceivably have consented to uses that had not been invented. Fewkes did not, and could not, obtain Peter Selmore's agreement for a use case that would arrive one hundred and thirty years later.

This is what makes the informed-consent argument so difficult to wave away. When an institution says that its collection is lawfully held and that it is free to license or release it, the claim is legally true and ethically hollow, because the original acquisition would fail every element of the standard that institution would now apply to a living research subject. Contemporary research ethics require that consent be informed, specific, revocable and given by someone with the authority to give it. The archival material fails on all four counts. It was gathered without meaningful information about future use, without specificity, without any mechanism of revocation, and often from individuals who held the knowledge under community protocols that did not give them the personal authority to alienate it. A speaker might share a song with a visiting anthropologist without possessing the right, under his own community's law, to release it to the world.

The AI moment forces this latent problem into the open because it dramatically raises the stakes of the downstream use. For a century the material mostly sat inert, and the injustice of its acquisition, while real, was static. Digitisation made it copyable. Machine learning makes it generative. A model trained on a corpus of restricted ceremonial knowledge does not merely store that knowledge; it can produce new outputs in its style, answer questions about it, and disseminate approximations of it to anyone who asks. The original failure of consent is thereby compounded and multiplied, and the community's ability to enforce its own protocols, already eroded by digitisation, collapses entirely. The advocates in Geneva were not raising a historical grievance for its own sake. They were pointing out that the historical grievance is now the input to an industrial process.

Where the Law Stops Short #

The obvious question is why existing law does not already resolve this, and the answer is that the relevant instruments were built for adjacent problems and stop short of the AI training set.

The most significant recent development is the WIPO Treaty on Intellectual Property, Genetic Resources and Associated Traditional Knowledge, adopted at a diplomatic conference in Geneva in May 2024 after decades of negotiation. It is the first WIPO treaty to deal with the interface between intellectual property and Indigenous knowledge, and its central mechanism, in Article 3, requires patent applicants to disclose the source or origin of genetic resources and associated traditional knowledge on which an invention is based. This matters, but its reach is narrow. It applies to patents, not to training data. It addresses the specific problem of patents being granted over Indigenous knowledge without acknowledgement, not the general problem of that knowledge being ingested by a model. And it enters into force only after fifteen states ratify it. Malawi deposited its instrument on 5 December 2024 and Uganda followed on 9 July 2025. That is the entire tally. Thirteen further ratifications are required, and more than two years after adoption the treaty is still not in force. The treaty is a disclosure requirement for one narrow channel of appropriation, not a consent requirement for the broad one, and even that narrow requirement is not yet law anywhere.

In the United States, the closest analogue is the Native American Graves Protection and Repatriation Act, whose revised regulations took effect in January 2024 and notably strengthened the requirement that institutions obtain consent from lineal descendants or tribes before exhibiting or conducting research on covered items. Those revisions prompted several major museums to close or cover Native American displays overnight while they sought the necessary consent. But NAGPRA governs human remains, funerary objects, sacred objects and objects of cultural patrimony in physical custody. It does not cleanly reach digitised sound, transcribed oral history or a language corpus, and it certainly does not reach a model trained on them. The statute that forced museums to reckon with consent for physical objects has no obvious purchase on the intangible material now flowing into AI systems.

Copyright, the tool a technology company would most naturally invoke, cuts the wrong way entirely. Much archival Indigenous material is old enough to be in the public domain, which under copyright law means it may be freely used, and the ongoing legal battles over whether training AI on copyrighted material constitutes fair use are, from an Indigenous-knowledge perspective, almost beside the point. The community's objection is not that its copyright has been infringed; it is that copyright never captured the interest at stake. A song can be simultaneously in the copyright public domain and subject to strict community protocols about who may perform it. The law that governs the copy has nothing to say about the protocol. This is the misalignment that CARE, OCAP and the Traditional Knowledge Labels were invented to address, and it is why the Geneva advocates were appealing to human-rights instruments rather than intellectual-property ones. The rights they are asserting do not fit inside the categories the technology industry recognises.

Procurement as the Real Decision Point #

If the frameworks exist and the law lags, the practical question becomes where in the process an obligation could actually bite, and here the answer is less about technology than about timing. Malloy's Governing analysis ends not with a call for better algorithms but with a call for meaningful tribal consultation before AI systems are designed and deployed, rather than notification after the fact. This is a deceptively large claim. In the ordinary institutional workflow, an organisation decides to acquire or build an AI system, selects a vendor, signs a contract, and only then, if at all, convenes a consultation about ethics and community impact. By that stage the architecture is fixed, the money is committed, and the consultation can influence little beyond the wording of a usage policy. Authority exercised after procurement is not authority at all; it is public relations. For tribal control to be real, the community's right to say no, or to set conditions, has to be available at the point where saying no would actually change the outcome, which is before the institution commits to building the thing.

This reframes the debate away from the familiar terrain of algorithmic transparency and bias auditing, which are downstream remedies, and toward the upstream decision that determines whether a system gets built at all. It aligns with the free, prior and informed consent standard in a way that most technology governance conspicuously does not, because the word that governance frameworks routinely underweight is prior. Consent obtained after deployment is not prior consent. A consultation convened to smooth the reception of a decision already taken is not consent in any meaningful sense. The whole argument is about procedural sequencing: who is in the room, and at what stage, is where Indigenous authority is either honoured or hollowed out.

What earlier involvement looks like in practice is beginning to be documented. A March 2026 paper by Dora Zhao and colleagues describes a series of co-design workshops with twenty-two public school educators in Hawai'i, convened around the educators' own concerns about cultural misrepresentation and bias in AI systems. Its argument is that auditing an AI system should be understood as a community-oriented process rather than the work of isolated individuals. The tools that emerged from those workshops were shaped by the participants' concerns because the participants were in the room while the tools were still taking shape, which is precisely the condition that late-stage consultation cannot reproduce.

The practical implications are concrete. A university library deciding whether to make its Indigenous collections available to an AI vendor would, under this logic, be obliged to consult the relevant communities before issuing the tender, not after signing it. A national museum contemplating a generative-AI interface to its holdings would need to bring the source communities into the design while fundamental choices, what is included, what is excluded, who governs access, remain open. Most institutions do the opposite, treating community consultation as a late-stage validation exercise, which is exactly the failure mode Malloy identifies.

Sovereignty as a Precondition, Not a Courtesy #

The most ambitious of the recent contributions tries to move the whole conversation from ethics to architecture. A 2026 paper in AI & Society setting out a pluralistic, participatory approach to global AI governance argues that Indigenous knowledge systems and the right of Indigenous peoples to self-determination, underpinned by free, prior and informed consent and the CARE Principles, must be foundational to AI regulation rather than an optional addition to it. Its foundational move is to treat the right of a cultural community to exclude, limit, or set conditions on AI use of its knowledge not as an ethical enhancement but as a structural requirement of the system. In that framing, the ability of a community to say no is not a feature to be added if resources permit. It is a precondition of the system being legitimate at all. A knowledge system that cannot represent and enforce the exclusion is, on this account, defective by design, in the same way that a database without access controls would be considered defective.

This is the intellectual heart of what the Geneva advocates were reaching toward, and it inverts the default. The prevailing institutional posture treats tribal authority as something to be accommodated where feasible: a nice-to-have, honoured when a community is organised enough to demand it and convenient enough to grant. The sovereignty argument insists on the reverse. Authority over the knowledge is the starting condition. The burden is not on the community to prove why its knowledge should be withheld, but on the institution and the developer to establish that they have the right to use it at all. This is the same inversion that free, prior and informed consent performs in every other domain of Indigenous rights: consent is presumed absent until it is actively and legitimately given, rather than presumed present until someone objects.

The tradition this draws on is now substantial and cannot be dismissed as marginal. The Indigenous Protocol and Artificial Intelligence Position Paper, produced in 2020 out of workshops led by the scholar and artist Jason Edward Lewis and involving more than thirty researchers and artists, argued years ahead of the current wave that AI must be designed in partnership with specific communities rather than around assumed universal values, that Indigenous communities must retain full control over their own data, and that ethical scrutiny must extend across the entire development process. That work has since grown into Abundant Intelligences, an Indigenous-led international research programme rethinking what AI could be if it were placed inside Indigenous knowledge systems rather than extracting from them. UNESCO has moved in the same direction, releasing guidance on Indigenous data sovereignty in AI and a report on Indigenous people-centred artificial intelligence developed with communities across Latin America and the Caribbean, insisting that the digitisation of Indigenous data must guarantee self-determination, governance and free, prior and informed consent. The direction of travel in the normative literature is unmistakable. It has simply not yet been translated into binding obligation.

That translation is what the July 2026 session in Geneva was ultimately about. The advocates were not asking whether AI can be used on Indigenous knowledge, a question the archives and the models have already answered. They were asking whether the institutions holding these materials carry any enforceable duty to treat tribal authority as a precondition of deployment rather than a gesture of goodwill. On the current state of the law the answer is mostly no. There are principles without teeth, labels without reach, a treaty confined to patents and still short of the ratifications it needs, a repatriation statute confined to physical objects, and a copyright regime that measures the wrong thing. What there is not, yet, is a rule that says an archive must ask before it lets a model in.

The wax cylinders in Calais have travelled a long way from the horn Fewkes spoke into. They have been catalogued, restored, digitised, labelled and, in the Passamaquoddy case, partly returned. Whether the next generation of that material, the recordings not yet governed by a Traditional Knowledge Label, the collections not yet subject to a differential-access system, the languages not yet defended by a community strong enough to write its own licence, is treated as a training set or as a trust will depend on decisions that institutions are making right now, mostly without asking. The people who sang into the machine could not consent to what would be done with their voices. The least their descendants are owed is the standing to decide.

References #

Tim Green UK-based Systems Theorist & Independent Technology Writer

Tim explores the intersections of artificial intelligence, decentralised cognition, and posthuman ethics. His work, published at smarterarticles.co.uk, challenges dominant narratives of technological progress while proposing interdisciplinary frameworks for collective intelligence and digital stewardship.

His writing has been featured on Ground News and shared by independent researchers across both academic and technological communities.

**ORCID:** [0009-0002-0156-9795](https://orcid.org/0009-0002-0156-9795)
**Email:** [tim@smarterarticles.co.uk](mailto:tim@smarterarticles.co.uk)

Listen to the free weekly [SmarterArticles Podcast](https://www.smarterarticles.fm)
── more in #artificial-intelligence 4 stories · sorted by recency
── more on @kerri j. malloy 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/scraped-without-aski…] indexed:0 read:29min 2026-08-01 ·