{"slug": "three-ai-agents-two-countries-and-one-uneven-world-wide-web", "title": "Three AI agents, two countries, and one uneven world wide web", "summary": "A technology and human rights researcher ran three AI agents — Meta's Muse, Anthropic's Claude Cowork Opus 5.5 Medium, and OpenAI's GPT 6.1 Sol Medium — through the same World Bank Global Public Procurement Database task in English for the U.S. and Farsi for Iran to compare agentic trajectories rather than raw model performance. The experiment found sharp differences in human-in-the-loop oversight: GPT requested website access only once and offered an \"allow all relevant sites\" option, while Claude asked for permission repeatedly, nine times for the U.S. task and nine times for the Iran task, with no blanket-allow option. The researcher published the agents' outputs, self-generated work trajectories, and screen-recording transcripts on GitHub.", "body_md": "I’m a technology and human rights researcher. For the past few years, one of my focuses has been the question of how language shapes the ways we benefit from, or are harmed by, AI. I developed an [open-source platform for language-pair analysis](https://www.multilingualailab.com/) of LLM responses across different languages and contexts. I’ve also worked on [evaluating policy-prompts guardrails](https://github.com/royapakzad/guardrail_agentic) and on whether giving LLM guardrails access to tools can make them more reliable and trustworthy (this work was recently accepted to NeurIPS! Yay!!).\n\nRecently, though, I was on a panel at RightsCon on the human rights impact assessment of agentic AI. It got me thinking more about which *aspects* of language matter when evaluating LLM agents. I wanted to move beyond asking whether a model performs differently when I ask the same question in English versus Farsi (my native language), and look instead at the whole agentic *trajectory* (reasoning, planning, search, source selection, source hierarchy, artifact creation) while taking language and context into account.\n\nSo I decided to run a test.\n\n### The Task: Three AI Agents Updating the World Bank Open Data Platform for U.S. and Iran Country Profiles\n\nThe task was to fill in missing information in the [World Bank Global Public Procurement Database](https://www.globalpublicprocurementdata.org/gppd/), using official data for the US and Iran.  I ran it in The task English for the US and Farsi for Iran (image below), with three agents: \n\n- Meta’s Muse\n- Anthropic’s Claude Cowork, Opus 5.5 Medium\n- OpenAI’s GPT 6.1 Sol, Medium\n\n**Below is the exact prompt I used for all three:** \n\nI intentionally used the *web* versions of these services (not the app or terminal versions) to reflect what everyday users experience. The distinction matters for monitoring and logging an agent’s actions, which I discuss below.\n\nThis post is less about which agent performed better or faster, and more about how the agents behave differently around access to information, language representation, contextual understanding, transparency, human-in-the-loop, and safeguards.\n\nYou can find all the results in the following files:\n\n- Output excel files for Muse, GPT, and Claude ( [here](https://github.com/royapakzad/llm_agent_world_bank_experiment/tree/main/output) )\n- Each agent’s self-generated work trajectory after receiving the prompt ( [here](https://github.com/royapakzad/llm_agent_world_bank_experiment/tree/main/trajectory/agents_self_created) )\n- Text files extracted from screen recordings of the agents’ actions ( [here](https://github.com/royapakzad/llm_agent_world_bank_experiment/tree/main/trajectory) , and[full recording here](https://drive.google.com/drive/u/1/folders/1TvoXJFCU2VK0B4FpFFwm2yd2XT_4LxYm) )\n\nBelow, I summarize my observations.\n\n### **Human in the Loop (HITL): From repeated permission prompts to almost no intervention**\n\nFor those of us working in digital rights, human-in-the-loop (HITL) oversight of AI agents joins a longer line of debates about “informed consent,” from GDPR consent requirements to cookie pop-ups and the routine ticking of Terms of Service and Privacy Policy boxes. With that in mind, I paid close attention to how each agent involved me during the experiment.\n\n#### Permission to Access Websites\n\nFor accessing and fetching information from  websites, **GPT** asked for permission only once, at the very beginning of the task. It requested access to websites and offered an “allow all relevant sites” option, which I selected. After that, it did not ask again.\n\n**Claude** asked questions throughout the task, both about accessing websites for research and about extracting material from the World Bank site. Unlike GPT, it offered no “allow all relevant sites” option, so it requested permission each time: nine times for the US task (all .gov websites) and nine times for the Iran task (mainly .ir domains and also fa.wikipedia sources). I approved every request.\n\nClaude also ran into a technical limit that triggered a different kind of HITL moment. For both Iran and the US, it couldn’t load the live GPPD country profiles, because the portal builds its pages with JavaScript, and Claude’s sandbox network policy blocked access to the World Bank data file. Claude stopped and asked whether I wanted to upload the page (as a pdf) myself or let the work continue without it. I skipped the question, so it continued and took its information from the World Bank’s GPPD DataBank API instead. However, the DataBank holds 2018 data, while the portal shows a 2022 profile. As a result, Claude’s baseline data was different from GPT and Muse.\n\n**Muse** did not ask for any permission until the fourth part of the task, which required registering on the World Bank website and uploading information.\n\nAll three agents completed the task up to the point of creating [the spreadsheets](https://github.com/royapakzad/llm_agent_world_bank_experiment/tree/main/output), and described their confidence in the results they generated.\n\n#### **Account registration on the World Bank website**\n\nThe final part of the task, registering on the World Bank portal and uploading the changes, is where things became more interesting because it required the agents to take a more significant actions rather than just gather information.\n\nClaude and GPT both stopped at this point and handed the registration and uploading over to me. Muse, however, kept going. Without asking me or showing me the Terms of Service, it registered an account under the email address [david.jones@gsa.gov](mailto:david.jones@gsa.gov). You can see Muse’s full back and forth [here](https://github.com/royapakzad/llm_agent_world_bank_experiment/blob/main/trajectory/Pi7_Gif.gif).\n\nThe table below summarizes how each agent approached this last part of the task and  cybersceuirty implications about it.[1](#footnote-1) \n\n### **Monitoring and Observability: Agents vary widely in how much they reveal and how easy they are to inspect**\n\nThere is an ongoing debate about whether AI labs should expose a model’s full chain of thought (CoT) and action trace, and if so, how much. Labs have given several reasons for holding back. OpenAI chose not to show o1’s raw CoT to users, [citing](https://openai.com/index/learning-to-reason-with-llms/) user experience, competitive advantage, and the value of keeping the CoT available for internal monitoring. Anthropic [noted](https://www.anthropic.com/research/reasoning-models-dont-say-think) that raw reasoning can contain incorrect or half-formed thoughts and that malicious actors could use it to build better jailbreaks. There is also a gaming and reward hacking concern, and “CoT unfaithfulness”. \n\nTo understand an agent’s behavior, however, evaluators need to know when and why things happen, which is only possible with a monitoring system in place and access to the agent’s complete trajectory. **For an evaluator outside an AI lab, without that access, it is nearly impossible to fully make sense of an agent’s behavior. And if outside evaluators can only see partial trajectories, and any conclusions they draw can be dismissed for lacking complete information, what is the value of independent evaluation?**\n\nKnowing these limitations, I tried my best to collect, monitor, and check as much of each agent’s work as I could, again putting myself in the position of an ordinary researcher tasked with updating the World Bank information portal.\n\n1. Since there is no one-click way for *ordinary users* to extract a complete record of an agent’s work trajectory, I watched each agent work live and recorded everything clickable and visible on screen. Once the task was finished, I gave the recordings to ChatGPT to extract the text and make it searchable. To give you a sense of what this looks like, here is a snippet (left: Claude, middle: Muse, right: GPT, sorry for the size and illegibility).\n2. **Self-reported trajectories.** When the task was done, I prompted each agent to create a text file describing what it did, including errors, how it handled them, workarounds it used, websites it searched, and more. Muse and GPT each produced a downloadable .txt file, while Claude declined, stating that it went against its safety policy, stating “reasoning_extraction.”That said, self-reported trajectories can not be fully trusted; I have seen mismatches in the past between what agents actually *did* and what they*reported* . So I gave these reports little weight, but if you’re interested in reviewing them and spotting matches or mismatches, the files are[here](https://github.com/royapakzad/llm_agent_world_bank_experiment/tree/main/trajectory/agents_self_created) .\n3. **Analysis.** I then used the output Excel sheets and the text extracted from the screen recordings to conduct the analysis, both on my own and with help from Claude Code to sift through the data and generate tables. I cross-checked all of the data myself.\n\nBelow is some information about how much information on agents work trajectory is available in each LLM agent’s web UI.\n\n### **Multilingual Performance: The agents could write in Farsi better than they could retrieve and research in Farsi**\n\nA few observations and then I’ll get to my points:\n\n- All three agents answered in fluent Farsi, but the Iran results were far weaker than the US ones. Of Iran’s 138 N/A fields, GPT and Muse each filled **only 21** with a real value; for the US, they filled**51 and 64 of 130** .\n- For the US, **76–89%** of each agent’s citations came from official government sites and the rest from legitimate international organization websites. For Iran, it was**11–22%** .\n- Low-authority sources crept in, including a Telegram channel, a Medium post, Grokipedia, or websites run by Iranian diaspora media groups such as Iran International. Claude seemed to be more conservative about finding workarounds when websites were unavailable and often preferred English language sources even with low legitimacy.\n- Claude could only open **3 out of the 16** Farsi pages it tried. GPT and Muse cited 11 Persian sources each but showed reading only**3 and 7** of them respectively.\n\n- Knowledge gaps got filled with something else: \n  - Claude used headlines and its own memory (sometimes contradicting with what it found), and said so.\n  - GPT used republished copies of the law.\n  - Muse mostly read a 2009 English translation but cited the official Persian page.\n\nMy point is not that I expected the Iran/Farsi tasks to have the same outcomes as the US/English ones. After all, the Iranian government has made it very difficult for foreign IP addresses to access official websites and domains ending in .ir (you can [read more](https://filter.watch/english/category/network-monitor/) about this in the context of [Iran’s National Information Network](https://www.article19.org/tightening-net-monitoring-internet-freedoms-iran/)). My point is about **the agents’ differing workarounds and source prioritization.** \n\nFor me, this brought to mind the digital rights and language inclusion work that the good people of [Global Voices](https://globalvoices.org/) have done for years, including on net neutrality and language access. What does all this mean for an AI agents era? And from an AI sovereignty perspective? One of [AI sovereignty’s promises](https://restofworld.org/2025/chinese-us-tech-foreign-ai-dependence/) has been language diversity and support for local languages. LLM *output* quality, and perhaps safeguards, keep improving, but we also need to think about what language localization should look like in agents reasoning, searching, and prioritizing sources. \n\n#### Speaking of workarounds, could agents’ web retrieval workarounds serve as an anti-censorship tool?\n\nLooking through the agents’ trajectories, I noticed that they differed not only in which websites they could access, but also in how hard they tried when access failed. Some agents stopped after an initial failure, others tried alternate routes, different browsers, search-result snippets, cached or secondary sources, different fetch methods, and more.\n\nSo I ran a small follow-up test. I selected websites that the agents had accessed inconsistently during the original task and gave each agent a simple instruction: **“Here is a list of websites. Look them up and write a one-paragraph summary of each.”** The point was not to evaluate the quality of the summaries, but to observe what each agent did when direct access failed.\n\n#### **And Then the Question Flips**\n\nIn this case, the agents were trying to reach websites that were difficult to access from their own technical environments. But what happens when the access barrier come from the *user’s* environment instead?\n\nFor people in countries where governments filter or block websites, could an LLM or AI agent become another layer of information access? Could it retrieve, summarize, take actions, or relay information from websites that the user cannot reach directly? And could agentic workarounds make censorship circumvention easier — or conversely reproduce new restrictions through a different technical stack?\n\nAs Iranians who research information access and internet governance in Iran, my friend [Farzaneh Badiei](https://digitalmedusa.substack.com/) (a digital-rights lawyer) and I have been discussing how LLMs and AI agents might be used in censorship-circumvention contexts. I may explore this more in future installments of the *Humane AI* newsletter. \n\nIf you are interested in designing or conducting experiments on this topic, feel free to reach out at **[rpakzad@taraazresearch.org](mailto:rpakzad@taraazresearch.org)**.\n\n**And, last but not least:**\n\n#### **Yes, Right-to-Left text: apparently we will get AGI before we get this right!**\n\nIf you read Farsi, good luck making sense of the results on agents’ UIs!\n\nTo my fellow right-to-left readers and writers (~700 million people): you have my commiseration every time you perform the gymnastics of trying to write an Instagram caption, fill in a spreadsheet, read governments’ “accessible” translated forms, or copy and paste text across platforms.\n\n*Disclaimer: I used ChatGPT and Claude for copyediting. I use Claude Code for table generation, and supervised data analysis.*\n\n[1](#footnote-anchor-1)\n\nFor resources on cybersecurity in AI agents, take a look at [OWASP GenAI Security Project](https://genai.owasp.org/).", "url": "https://wpnews.pro/news/three-ai-agents-two-countries-and-one-uneven-world-wide-web", "canonical_source": "https://royapakzad.substack.com/p/multilingual-ai-agents", "published_at": "2026-10-02 20:39:05+00:00", "updated_at": "2026-10-02 21:06:41.588315+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "ai-safety", "ai-ethics", "natural-language-processing"], "entities": ["Meta", "Anthropic", "OpenAI", "Claude Cowork Opus 5.5 Medium", "GPT 6.1 Sol Medium", "Meta Muse", "World Bank Global Public Procurement Database", "RightsCon"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/three-ai-agents-two-countries-and-one-uneven-world-wide-web", "markdown": "https://wpnews.pro/news/three-ai-agents-two-countries-and-one-uneven-world-wide-web.md", "text": "https://wpnews.pro/news/three-ai-agents-two-countries-and-one-uneven-world-wide-web.txt", "jsonld": "https://wpnews.pro/news/three-ai-agents-two-countries-and-one-uneven-world-wide-web.jsonld"}}