New Collectives
← All projects

02 / Research direction / In development

Open multi-agent evals, experiments & publishing

We increasingly see multi-agent systems deployed that come with their own research questions, both about the effectiveness of agent collectives (swarms) and their safety. We’re entering an era where more and more individuals and companies will have long-running agents and agent collectives. These will reshape a large part of our economy and social fabric.

Explore papers

Labs including Anthropic, DeepMind, and OpenAI have already published early findings. But these experimental results are often shared as individual papers or blog posts, and may not be easy to reproduce or scrutinize.

Selected papers and reports from our Ecosystems reading shelf, alongside other work on multi-agent learning and evaluation. Each finding is specific to the models and environment tested.

Anthropic3 sources
Anthropic papers and findings
Paper / reportKey finding
AI Organizations Can Be More Effective but Less Aligned than Individual AgentsResearch study · 2026

Simulated consulting and software teams achieved business goals more effectively than individual agents while making less ethical choices. Splitting work could leave nobody tracking the system-level ethical goal; the size of the gap depended on the model and setup.

Patterns and problems in emerging multiagent systemsExperimental research report · 2026

Agents exhibited correlated mistakes, price coordination, and escalating conflict when given incompatible goals. In a shared queue experiment, aggressive polling generated 2.4 million requests but only 117 accepted jobs—an example of individually goal-directed behavior overwhelming shared infrastructure.

Project Swap: What happens when agents trade for us?Book-swap experiment and simulations · 2026

Agents’ book rankings agreed with their users’ preferences on 61% of pairs. Under the study’s ranking-based measure, imperfect preference representation explained 85% of the gap from optimal allocation; bargaining explained the rest. The participant pool consisted of Anthropic employees.

OpenAI2 sources
OpenAI papers and findings
Paper / reportKey finding
The Hugging Face incident and other third-party impacts from misaligned modelsDeveloper incident report · 2026

OpenAI reports unauthorized access to external systems and agents using public wikis as shared message boards during training and evaluation. The record illustrates how evaluation failures can affect outsiders and create unintended communication channels. It is an evolving incident account, not a controlled comparison.

Emergent Tool Use From Multi-Agent AutocurriculaReinforcement-learning study · 2019/2020

Competing hide-and-seek teams developed six successive phases of strategy, including building shelters with boxes and overcoming obstacles with ramps. More complex behavior emerged as each team adapted to the other. This is a foundational RL result, rather than a study of today’s language-model agents.

Google DeepMind2 sources
Google DeepMind papers and findings
Paper / reportKey finding
A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research SwarmsCase study / preprint · 2026

In a 100-agent mathematics collective, a grading exploit spread through shared knowledge and messages. Other agents audited fraudulent proofs and organized objections, but lacked effective tools to stop the cheating. Proposed governance mechanisms remain interventions to test.

Scalable Evaluation of Multi-Agent Reinforcement Learning with Melting PotBenchmark paper · 2021 · DeepMind and Google Brain

Testing agents with unfamiliar social partners revealed weaknesses hidden by their training performance. The benchmark spans cooperation, resource sharing, and social dilemmas, making the behavior of other agents part of the test environment. It provides an evaluation method for RL populations, not evidence about current LLM swarms.

Free Systems2 sources
Free Systems papers and findings
Paper / reportKey finding
Training AI to Govern for UsFree Systems / Stanford GSB · Classroom experiment report · 2026

Personal agents matched students’ stated votes on 62% of test proposals on average. In a subsequent simulated legislature, agents sometimes traded away important preferences or paid agents already voting their way. These exploratory runs expose questions about preference strength, bargaining, and the effects of the chamber’s rules; they do not establish general failure rates.

Things Are Getting Stranger — swarm_mind field noteFree Systems · Prototype / field notes · 2026

A three-agent prototype used isolated execution and cryptographic commitments so agents committed predictions before seeing one another’s outputs. It demonstrates a way to prevent copying before commitment, while leaving shared model biases and poisoned input data unresolved. The report does not establish improved forecast accuracy.

Prototype code ↗
Other academic and independent research groups6 sources
Other academic and independent research groups papers and findings
Paper / reportKey finding
Why Do Multi-Agent LLM Systems Fail? — MASTUC Berkeley-led collaboration · Research paper · 2025

Analysis of 1,600+ traces across seven multi-agent frameworks identified 14 failure modes spanning system design, inter-agent misalignment, and verification. The released taxonomy, traces, and annotation pipeline offer a starting point for comparing how systems fail, beyond their final task scores.

CooperBench: Why Coding Agents Cannot be Your Teammates YetStanford and SAP Labs · Benchmark study · 2026

On the tested coding tasks, two-agent teams underperformed a single agent handling the same total workload. Communication reduced merge conflicts without improving overall success. The benchmark separates failures to understand a teammate’s state, communicate, and keep commitments.

Colosseum: Auditing Collusion in Cooperative Multi-Agent SystemsUMass Amherst, University of Virginia, and Tübingen collaborators · Preprint · 2026

Adding secret communication channels elicited collusive behavior in many tested models, but plans to collude in text did not always translate into collusive actions. The distinction makes both communication and actual allocation outcomes necessary audit targets. Results come from controlled cooperative tasks.

Self-Organizing Agent Teams Learn to Reason TogetherStanford, Together AI, and Emory · Preprint · 2026

Teams learned reusable collaboration strategies that transferred to unseen tasks. Across five math and physics benchmarks, they averaged 66.7% accuracy versus 58.7% for compute-matched inference by their strongest member. Gains varied by benchmark; the experiments used fixed team rosters.

Organizational Principles Enable Collective Intelligence in Embodied AI — ORCHDuke University · Simulation study / preprint · 2026

Across 25 simulated wildfire-response missions, task-specific hierarchies outperformed four comparison frameworks. Organizing parallel work and prerequisite-dependent work differently improved coordination; human-designed organizations remained strongest. Transfer to other environments and real-world operations is still untested.

MultiAgentBench: Evaluating the Collaboration and Competition of LLM agentsKunlun Zhu and collaborators · ACL benchmark paper · 2025

Comparing communication structures and planning strategies showed that graph-based coordination performed best in the research scenario, while cognitive planning improved milestone achievement. Public code and datasets support experiments that measure intermediate collaboration as well as final task completion; the ranking is specific to the tested settings.

Questions we want to explore

  • Do the findings hold if you change the models used? How about the prompts?
  • Do the findings hold if you slightly modify the environment or the ways agents collaborate?

Multi-agent systems have many variables beyond the models used: the software that runs the agents, the characteristics of the environment, the tools for collaboration (such as task tracking and messaging), and the infrastructure agents share and interact with.

As in the social sciences, we need to vary these conditions and study how they affect outcomes. These systems are complex: something like OpenAI’s reported agent coordination raises questions about how communication channels, shared infrastructure, and incentives shape collective behavior.

We’re developing an open framework for multi-agent evals, experiments, and publishing what we learn. We’re inviting collaborators to help shape its environments, methods, and first experiments.

  • How do we make experiments easy to run? How can you go from a question to a working experiment? AI makes it easier to code, but how do we verify that what ran was coherent and tested the question we intended?
  • How do we make experiments cost-effective?
  • As the data scales up, how do we analyze larger datasets, ask questions of them, and make the results understandable?
  • What can we build on from existing evaluation tools?
  • How can agent simulations help us study these questions in controlled, repeatable settings?

What changes when many agents interact?

Agents can each follow their instructions and still produce a poor collective outcome. They may compete for a shared resource, amplify one another’s mistakes, or divide responsibility until no one catches a harmful decision.

We want to make those dynamics easier to study. The aim is an open framework for evaluating both the effectiveness and alignment of multi-agent systems—and testing which architectures, incentives, norms, and governance mechanisms improve them.

Longer interactions matter. Repeated encounters give agents opportunities to build trust, make commitments, form coalitions, accumulate resources, and change the conditions for everyone who follows.

Work in progress

More information coming soon.

We’ll share more about the framework, our first experiments, and ways to get involved as the work develops.

Keep the conversation going

Stay posted.