New Collectives
← All presentations

Event 01 · September 15, 2026 · Clay HQ, New York

Agent swarm economies inside RuneScape

Competition, collaboration, and emergent behavior among AI agents in a multiplayer game economy.

Download video ↗

Transcript

Click a timestamp to play from that point. Machine-generated transcript with names and key terms reviewed; some words and audience questions may be imperfect.

0:07Hey, thanks for having me. I'm really excited to be talking about multi-agent systems and alignment. think it's a really important topic right now. So I'm going to be sharing some kind of like worked examples from a real life multi-agent system that I maintain that happens to be an MMO called RuneScape. So just to introduce myself really quickly, my name is Max Bittker. I work on a project called Websim, which is like a platform where a bunch of people work together to build games and build really complicated multiplayer projects.

0:53But I'm going to be talking today about another project of mine called RS SDK, which is the RuneScape SDK, which is basically code bindings for any person who wants to write scripts, but it turns out language models love writing scripts to control a RuneScape character and observe its surrounding and act on goals. And maybe even really long horizon or multi-agent goals like competing for a high score or trading with other agents all inside kind of like an emulated open source RuneScape server.

1:29So part of the inspiration for this is that I was at one point like a kid who loved RuneScape and at a certain point I figured out that you could repeat all of the repetitive actions via scripts. And you didn't have to like mine 10,000 logs to get the goal you wanted. You could set up an auto clicker overnight and do it. And I thought that that game loop was so much more fun than RuneScape itself. And that kind of led me to programming and everything else.

1:55And I've always wanted more people to experience that as a game itself, like the metagame of automating the game. Unfortunately, it's like against the rules, but I just think it shouldn't be. Everybody should just be able to do it equally. And so coding agents really work well for this. This part of the goal is to get other people to have that experience.

2:21And so this project has been popular online, like thousands of people have tried it and used a coding agent for the first time to do these long horizon goals. I'll show the live version because it's cool to watch.

2:47Maybe not right now, but basically at any given time, there's hundreds and hundreds of different people's agents. Some people run one agent, some people run a swarm of agents, and they're all interacting on the server and pursuing goals. So this is kind of like a heat map of seven days of activity. And each of these yellow dots on the map is being controlled by a coding agent somewhere.

3:12There also is a version of this project that is an eval to measure the kind of like problem solving ability of different coding models. You can kind of see this is a Pareto curve, and you can see an outlier on the top left is GPT-6 Astra, one of the newest models on here, which is... This is a log scale, the way, so GPT-Astra is almost 10 times better than some of the models over here that are cheaper, even though it's also much more expensive.

3:42And this has actually turned out to be a really useful evaluation for just understanding how well models can deal with long horizon tasks and kind of goal following and optimization. And then also, you know, this is like classic scary graph of every AI thing, but x-axis here is just release of the model. And in the nine months that this benchmark has been out, there's been 10x improvement in how well that they score on it.

4:12And it's going up. It's actually, I think, of saturating.

4:16And so I've been looking at not just single agent tasks, but at multi-agent tasks and how well can they work together on either competitive or cooperative, like kind of market-based tasks. And so setting up agents into scenarios where they need to trade and collaborate in order to accomplish their goals.

4:42This is a... It's okay. This is a video of just like a grid of 20 agents all working at the same time to talk to each other and to trade with each other in order to... They're each optimizing for their own individual income. But the way they have to accomplish that is by talking to each other and setting up trades and basically finding prices.

5:08And there's also even kind of like exploitation because they might ask for like loans from other people. They're like, hey, I can pay you, you know, 2,000 gold for that item, but I need the item first in order to afford it. And so, yeah, I think I'd have to get my... Unfortunately. And so, basically, this has been a really interesting kind of like experimental playground for determining what kinds of misalignment and group misalignment scenarios happen and what factors are kind of like push them to happen more or less.

5:47And also how agents deal with scenarios. And Like if they've gotten scammed, how do they like tell all the other agents what happened? And in some cases, they've threatened to tell everybody and then gotten their money back. And then I guess that... So, I'm doing a lot of experiments with this kind of test bed.

6:13And one of my kind of like zoomed out hunches is that the way that agents are trained is that they do experience like many millions of hours of task goal following. But it's almost always single agent goal following with single agent rewards. If you look at the biological world and like our own evolution, we are a product of individual selection, but we're also the process of group and community selection.

6:44And many factors that explain the way that animals and plants behave is better explained by group selection than only individual selection. And so, I'm very curious about ideas about factoring in like negative externalities of your actions into training processes. And so, basically giving agents many examples of being in a world where they need to benefit the people around them in order to succeed at their goal.

7:15Much like our evolutionary past. So, really excited. I can definitely... I've got cool videos and transcripts that people want to see some examples of these agent scenarios. And yeah, thanks. Nice to be here.

7:32Thank you. [Audience question inaudible; the speaker repeats it below.] Yeah, so he was asking how the long running agent swarms, like how long do they run and how does that work to keep them running.

7:58So there's basically two examples. One is that I just run a server that's always on and agents are constantly just booting up, connecting to it. And that's each individual who runs the agent makes their own decision. Some people play kind of like interactively. Some people set it up in a server to be on a cron job and run all the time. And then inside of my kind of like controlled scenarios, those tend to be between like 30 and 90 minutes.

8:25And that just fits inside of one agent run. And so it's just setting up 10 sandboxes or 100 sandboxes that are each running a coding agent.

8:42I'm sure there are many cases of this. But just out of curiosity, relative to your expectations when you first started doing these experiments, what has been the most surprising emergent behavior that you've seen either on the live server or in your own controlled simulations, if there is one that stands out to you? Yeah. So I was really curious when I started how the... Because RuneScape is a really boring game.

9:09But the cool part is all about like the economy and the prices and different goals. And all your kind of like greatest moments in RuneScape are because you saved up or you found some kind of like money-making trick. And so I was really curious. okay, if RuneScape is so much of a labor resource processing economy, so how does that work if labor is very cheap? And it's been really cool watching this play out.

9:35One thing is that like trade, basically like inflation is super high on the server because people don't have a lot of demand for like the cost of coordinating with another agent compared to just leaving it overnight to do the work for you and go get the resource. It de-incentivizes trade. of of a trade. It's little And then the other thing that's been interesting is that certain resources that don't just scale linearly with labor but instead have some kind of natural scarcity.

10:01So an example in the game is Rune Ore only has a single spawn location. And if you have 100 people, only one person is going to get it. So these are the resources that have become scarce and kind of valuable. And so people make more and more advanced swarms just to compete for kind of these same scarce resources. And so that's been an interesting thing to watch play out.

10:28Any other questions?

10:32Yeah, this is more just the thought, but it made me think seeing the RuneScape example. One of the things that I've been thinking about is just like humans, as humans, we have bodies and we're like geographically constrained. And it's interesting that in this RuneScape example, I feel like the agents kind of have the same sort of instantiation. And it could be that there's like all kinds of interesting things from just like the geography. You know, it was like even the question of like, hey, why did these people share these values over here versus these ones?

11:01It's like partially determined by geographical constraints. And it is kind of, yeah, it just makes me wonder like if agents behave differently if you embody them and you like also constrain their geography. Yeah, definitely. I have no, yeah. I think of it kind of as like a mini robotics environment because you're taking a text-based coding agent, but actually all of its actions are expressed through a kind of like a thin nozzle, which is that you can only observe what's around you and you can only act on what's around you.

11:35And so a cool thing about this is that it has big implications for multi-agent in many kinds of software tasks. It's undetermined if having more individual actors is actually helpful versus just like one long running task in kind of like a game or a robotics task. More people collaborating just means more bodies. And so that's been interesting to watch too. But it's not always one-to-one.

12:03You do sometimes see one coding agent controlling a hundred actors in the game by kind of like multiplexing.

12:18Thanks. I was curious about like the limiting factors. Because right there, the slide where you were saying like how we are in biological, you know, our human bodies and our ecosystem. There's so many limiting factors here in terms of like the energy output we have per day, our attention, our focus. And with agents, it seems like there are many less limiting factors.

12:45Sort of like if you have infinite budget, then you can waste all the tokens you want. Do you see any way in which like new types of limiting factors could be applied to agents or to systems that could begin to put boundaries on these systems? I think that it becomes really obvious when the systems start playing out in practice. so I didn't come into like I was just going to run a stock RuneScape server.

13:15And once you start having all of these agents running on it, you just kind of like see what the breaking points are. And so for instance, there was this limiting factor of like RuneScape server can only hold so many people in it at once. And so in order to just keep it accessible, I started limiting per IP address. And so it's interesting that I think people often assume when it comes to AI, they use like infinity as the multiplier.

13:43But I think it's actually more just like a million or something or like or even more like a thousand. And so things you apply that new form of energy into the system and things just rearrange. They don't like explode. And so for example of this server, was like, okay, yeah, we're going to put on like each IP address can only connect 200 bots.

14:08And then there's even also limits where in order to run a bot, you kind of need to be running a web browser. So you're also limited by how much RAM you have, not to mention like tokens. And so the numbers get weird, but they don't go to infinity. So there's always some kind of balance to be struck.

14:30I just had a quick follow-up thought slash question regarding the rune ore thing. I don't know if you've looked at this, but I feel like it'd be interesting to like the whole notion of like comparative advantage, right? In like economics where it's just like, yeah, they can just do it themselves, but there's still an opportunity cost. And like the rune ore example made me wonder, it's like, well, if the rune ore is the thing that's scarce, why wouldn't they want to like dedicate all their resources to like beating everyone for the rune ore and just being like, hey, other agent over there.

14:56Like you can do this for me because I need my rune ore. And I'm kind of curious if, you know, we would expect that to end up happening or if there's already empirical evidence that it doesn't. Because I feel like that has probably implications for like how people think about, you know, the real world too. Like whether comparative advantage will actually continue to hold. So I don't know if you've thought about that or have seen anything in that regard. Yeah. So an example of the rune ore where this is like a very scarce resource.

15:24And if you want to make money, you kind of have to like go after some of these things that can't just be, people actually want to buy it from you because they can't get it themselves. And absolutely right now we see comparative advantage because if you're trying to go after one of these scarce resources, you don't just let the agent do it itself. That's where you start like giving the agent more resources, you suggest strategies.

15:50And so the people who are successfully getting access to these scarce resources on the server are the people who also currently are putting in the most dollars and the most human ingenuity. And so that's currently how it's playing out is that those are the people who are winning. Maybe there could be somebody who just purely puts in like tons of Astra credits and gives it a goal like this and they have a good outcome too.

16:17But right now it seems like Centaur kind of, you know, combinations are the people who are the most successful.

16:35I'm curious if you observe any, like this, this curve is really interesting. I'm curious if like, what is the behavior that changes as you move up the curve that creates such a massive difference in the XP? Like what is Astra doing? Like I would, I would kind of think that a game like RuneScape would be saturated at some point. And so what did, what is Astra doing that is so much more effective than what the other models are doing?

17:03So that's a great question. At the low end of the curve of just like being like better, you know, this is where we see like Sonnet 4.5. Just navigating the game is really hard because the game actually has a surprising amount of weird stuff in it. And you're only given 30 minutes wall clock. so Sonnet here, like it was supposed to go train crafting for this task.

17:31And in 30 minutes, it just probably like got stuck on a door. It tried to go find something over here, but then it needed this. And like it couldn't just untangle the web. And so at the low end, you see just too much complexity and they get confused. In the mid range, a lot of it has to do with the difference between doing the task and optimizing the task. And so sometimes you'll see models where they will accomplish a task like fishing.

17:59And then they'll kind of just chill for like the next 15 minutes. And they'll keep doing the same loop, but they won't kind of have this like feeling of like, I got to figure out how to catch these fish faster. I got to go try different fish. I got to go try different stuff. And so the benchmark is really set up to reward. It rewards your peak XP rate within any 15 second window. So once you've kind of found one strategy, you're supposed to keep looking for better strategies.

18:28And you're not supposed to just look for like slightly better strategies. Like I'm going to keep catching shrimp, but I'm going to like click differently. You're supposed to go explore and try more complicated strategies. And so at the top of the skill expression, you see agents who kind of like reason without acting about what strategies will be good. In some case, even what strategies will be optimal based on all the information and then beeline for those.

18:57then additionally, if they fail at those, they'll be like, I've only got 15 minutes left. This strategy is not working. I'm going to back off and do this safer strategy. So it's this combination of like, to be honest, the specifically the Astra run is like scary because it's definitely superhuman in terms of a human with no planning. And it is like, it goes straight for the very most optimal strategy.

19:24And there's a chance to be honest, they are old on this task. Like it is open source. So that's like, I kind of hope they did. But if it's just straight intelligence and kind of like information crunching, it is, it shows really high confidence. Just go for the best strategy.

19:42Have you tested humans on this task? Like expert human players? It's really weird because I run this strategy on an eight times speed server. So a human would have to... I have my mental idea of what a perfect human would be. Instead of the 8X speed you gave them the equivalent four hours. I think they would probably be really close to Astra or Beta. If they were a smart player who really knew the game. A random person would probably be more in the middle.

20:19Got it. Thanks. Thank you Max.

Continue exploring

More from Event 01

Nicolae Rusan

Agent organizations, economies & collective safety

Eric Tang

Designing agent-native organizations

Yondon Fu

Winning the metric, losing the commons

Prof. Andrew Caplin

Cognitive economics and the organization of inquiry

Keep the conversation going

Stay posted.