Researcher tests Meta, Anthropic and OpenAI agents on US-Iran data task
A technology and human rights researcher ran an experiment comparing three AI agents—Meta's Muse, Anthropic's Claude Cowork, and OpenAI's GPT 6.1 Sol—on a shared task: filling gaps in the World Bank's Global Public Procurement Database for US and Iran country profiles. The prompt was given in English for the US portion and Farsi for the Iran portion, using the web versions of each tool to mimic typical user experience. The researcher examined differences in reasoning, search behavior, source selection, transparency and safeguards across the agents rather than ranking speed or accuracy.
GoKawiil's interpretation of the reporting above, not reported fact.
The test suggests that an AI agent's behavior—what sources it trusts, how transparent it is, how it handles human oversight—may shift depending on the language and country context of a task, not just the underlying model's general capability. This could matter for human rights assessments of agentic AI, since inconsistent safeguards or source hierarchies across languages might create uneven risks for users outside English-speaking contexts, according to the researcher who presented related work at RightsCon.
- Three agents—Meta Muse, Anthropic's Claude Cowork, and OpenAI's GPT 6.1 Sol—were tested on the same data-entry task.
- The task used English for US data and Farsi for Iran data to probe language-based differences in agent behavior.
- The comparison focused on transparency, source selection, and safeguards rather than simply which agent was fastest or most accurate.
Source: royapakzad.substack.com — Roya Pakzad, 2026-10-02
Published there as: “Three AI agents, two countries, and one uneven world wide web”
Read the original report → The summary and analysis above are GoKawiil's own, written from reporting by the source above. Facts and quotes belong to the original publisher.