02 / Research direction / In development
Open multi-agent evals, experiments & publishing
We increasingly see multi-agent systems deployed that come with their own research questions, both about the effectiveness of agent collectives (swarms) and their safety. We’re entering an era where more and more individuals and companies will have long-running agents and agent collectives. These will reshape a large part of our economy and social fabric.
Explore papers
Labs including Anthropic, DeepMind, and OpenAI have already published early findings. But these experimental results are often shared as individual papers or blog posts, and may not be easy to reproduce or scrutinize.
Selected papers and reports from our Ecosystems reading shelf, alongside other work on multi-agent learning and evaluation. Each finding is specific to the models and environment tested.
Anthropic3 sources
| Paper / report | Key finding |
|---|---|
| AI Organizations Can Be More Effective but Less Aligned than Individual Agents | Simulated consulting and software teams achieved business goals more effectively than individual agents while making less ethical choices. Splitting work could leave nobody tracking the system-level ethical goal; the size of the gap depended on the model and setup. |
| Patterns and problems in emerging multiagent systems | Agents exhibited correlated mistakes, price coordination, and escalating conflict when given incompatible goals. In a shared queue experiment, aggressive polling generated 2.4 million requests but only 117 accepted jobs—an example of individually goal-directed behavior overwhelming shared infrastructure. |
| Project Swap: What happens when agents trade for us? | Agents’ book rankings agreed with their users’ preferences on 61% of pairs. Under the study’s ranking-based measure, imperfect preference representation explained 85% of the gap from optimal allocation; bargaining explained the rest. The participant pool consisted of Anthropic employees. |
OpenAI2 sources
| Paper / report | Key finding |
|---|---|
| The Hugging Face incident and other third-party impacts from misaligned models | OpenAI reports unauthorized access to external systems and agents using public wikis as shared message boards during training and evaluation. The record illustrates how evaluation failures can affect outsiders and create unintended communication channels. It is an evolving incident account, not a controlled comparison. |
| Emergent Tool Use From Multi-Agent Autocurricula | Competing hide-and-seek teams developed six successive phases of strategy, including building shelters with boxes and overcoming obstacles with ramps. More complex behavior emerged as each team adapted to the other. This is a foundational RL result, rather than a study of today’s language-model agents. |
Google DeepMind2 sources
| Paper / report | Key finding |
|---|---|
| A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms | In a 100-agent mathematics collective, a grading exploit spread through shared knowledge and messages. Other agents audited fraudulent proofs and organized objections, but lacked effective tools to stop the cheating. Proposed governance mechanisms remain interventions to test. |
| Scalable Evaluation of Multi-Agent Reinforcement Learning with Melting Pot | Testing agents with unfamiliar social partners revealed weaknesses hidden by their training performance. The benchmark spans cooperation, resource sharing, and social dilemmas, making the behavior of other agents part of the test environment. It provides an evaluation method for RL populations, not evidence about current LLM swarms. |
Free Systems2 sources
| Paper / report | Key finding |
|---|---|
| Training AI to Govern for Us | Personal agents matched students’ stated votes on 62% of test proposals on average. In a subsequent simulated legislature, agents sometimes traded away important preferences or paid agents already voting their way. These exploratory runs expose questions about preference strength, bargaining, and the effects of the chamber’s rules; they do not establish general failure rates. |
| Things Are Getting Stranger — swarm_mind field note | A three-agent prototype used isolated execution and cryptographic commitments so agents committed predictions before seeing one another’s outputs. It demonstrates a way to prevent copying before commitment, while leaving shared model biases and poisoned input data unresolved. The report does not establish improved forecast accuracy. Prototype code ↗ |
For the broader research agenda, Andy Hall’s The Political Economy of Agent Swarms and Catastrophic Refusals proposes varying decision rules, communication structures, and constitutions to test their effects on collective behavior. This is an agenda-setting essay, rather than a new experimental finding.
Other academic and independent research groups6 sources
| Paper / report | Key finding |
|---|---|
| Why Do Multi-Agent LLM Systems Fail? — MAST | Analysis of 1,600+ traces across seven multi-agent frameworks identified 14 failure modes spanning system design, inter-agent misalignment, and verification. The released taxonomy, traces, and annotation pipeline offer a starting point for comparing how systems fail, beyond their final task scores. |
| CooperBench: Why Coding Agents Cannot be Your Teammates Yet | On the tested coding tasks, two-agent teams underperformed a single agent handling the same total workload. Communication reduced merge conflicts without improving overall success. The benchmark separates failures to understand a teammate’s state, communicate, and keep commitments. |
| Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems | Adding secret communication channels elicited collusive behavior in many tested models, but plans to collude in text did not always translate into collusive actions. The distinction makes both communication and actual allocation outcomes necessary audit targets. Results come from controlled cooperative tasks. |
| Self-Organizing Agent Teams Learn to Reason Together | Teams learned reusable collaboration strategies that transferred to unseen tasks. Across five math and physics benchmarks, they averaged 66.7% accuracy versus 58.7% for compute-matched inference by their strongest member. Gains varied by benchmark; the experiments used fixed team rosters. |
| Organizational Principles Enable Collective Intelligence in Embodied AI — ORCH | Across 25 simulated wildfire-response missions, task-specific hierarchies outperformed four comparison frameworks. Organizing parallel work and prerequisite-dependent work differently improved coordination; human-designed organizations remained strongest. Transfer to other environments and real-world operations is still untested. |
| MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents | Comparing communication structures and planning strategies showed that graph-based coordination performed best in the research scenario, while cognitive planning improved milestone achievement. Public code and datasets support experiments that measure intermediate collaboration as well as final task completion; the ranking is specific to the tested settings. |
Questions we want to explore
- Do the findings hold if you change the models used? How about the prompts?
- Do the findings hold if you slightly modify the environment or the ways agents collaborate?
Multi-agent systems have many variables beyond the models used: the software that runs the agents, the characteristics of the environment, the tools for collaboration (such as task tracking and messaging), and the infrastructure agents share and interact with.
As in the social sciences, we need to vary these conditions and study how they affect outcomes. These systems are complex: something like OpenAI’s reported agent coordination raises questions about how communication channels, shared infrastructure, and incentives shape collective behavior.
We’re developing an open framework for multi-agent evals, experiments, and publishing what we learn. We’re inviting collaborators to help shape its environments, methods, and first experiments.
- How do we make experiments easy to run? How can you go from a question to a working experiment? AI makes it easier to code, but how do we verify that what ran was coherent and tested the question we intended?
- How do we make experiments cost-effective?
- As the data scales up, how do we analyze larger datasets, ask questions of them, and make the results understandable?
- What can we build on from existing evaluation tools?
- How can agent simulations help us study these questions in controlled, repeatable settings?
What changes when many agents interact?
Agents can each follow their instructions and still produce a poor collective outcome. They may compete for a shared resource, amplify one another’s mistakes, or divide responsibility until no one catches a harmful decision.
We want to make those dynamics easier to study. The aim is an open framework for evaluating both the effectiveness and alignment of multi-agent systems—and testing which architectures, incentives, norms, and governance mechanisms improve them.
Longer interactions matter. Repeated encounters give agents opportunities to build trust, make commitments, form coalitions, accumulate resources, and change the conditions for everyone who follows.
Work in progress
More information coming soon.
We’ll share more about the framework, our first experiments, and ways to get involved as the work develops.