New Collectives
← All presentations

Event 01 · September 15, 2026 · Clay HQ, New York

Collectives as safety primitives — the full event

Five talks and their audience discussions on agent organizations, shared resources, virtual economies, and cognitive economics. Recorded at Clay HQ in New York on September 15, 2026.

Download video ↗

Transcript

Click a timestamp to play from that point. Machine-generated transcript with names and key terms reviewed; some words and audience questions may be imperfect.

0:02Hey. All right. I think we can start to take a seat and get started over here. How do I guess I take that? Amazing. All right.

0:32Thank All right. Okay. So I think we'll get started. Thanks. Hope you all got some food and some drinks. They'll be still there as we get into the evening.

1:22And thanks for taking some time out of your weeknight to join us here. This is the first time we're doing this event. And the way that this event started is we found ourselves having more and more conversations about AI agents and how teams of agents interact and multi-agent systems. And I think we felt that this conversation was entering into the mainstream more and more. And we thought it would be useful to start to create a forum for more folks to have conversations about this new emerging trend, both the risks of these multi-agent systems and also ways to potentially make them more beneficial and safe.

2:06And so this is an experiment in the format. We've invited a few folks to talk on a variety of topics. And Eric will also help to emcee. We'll do a little question period after each talk. And so there'll be five presentations today. And I'll just start out by giving the high-level overview of what do we mean when we talk about AI collectives or agent organizations or agent economies and why do we think that this matters.

2:37And then we'll go from there. So first of all, thank you, Clay, for hosting us. And thank you, Puneet, Hisham, Aaron, for helping us get this all ready on short notice. I think we had this idea to do this event like a week ago. So like less than a week ago, we were like, I think we should start bringing people together for this. And I know also there's a Clay happy hour happening at the same time. we will, this is the people that chose to not go to the happy hour.

3:07They chose to came to the AI anxiety hour instead. No, optimism, AI optimism as well. And so I'm Nicolae. I was once upon a time a Clay co-founder, but a long, long time ago, 2022 is when I left. And then I spent a bunch of time. I went in 22 to an OpenAI event and they said, all the benchmarks and evals are in exponentials. And I was like, oh, that's a pretty wild thing to say. And so then I shifted all my attention to working in AI.

3:38And pretty, I've been kind of concerned about AI safety for a little while now. And so I've decided to spend more and more of my time on this topic. So it seems like the world in the past month has really woken up and become very concerned about AI safety. And I wanted to do two things in this talk. First of all, I wanted to just get everyone up to speed. I think there's probably varying levels of attention that people have been paying to what's happening in with all this like safety stuff and with the AI labs.

4:10And so I just wanted to share some information of why are people concerned and what's my perspective on evaluating that. And then I wanted to talk a little bit about one of the main things that people are concerned about, which is the fact that these AI agents have started to really try to coordinate with each other towards obtaining various objectives. And these multi-agent systems, multi-agent swarms, you might hear them called, we're calling them AI collectives here, pose their own new risks and questions about how to make them safe and effective.

4:44And they also have really exciting opportunities for new sorts of organizations and products that we could potentially build to for great uses. So why is everyone so worried all of a sudden? I saw a tweet from Noam Brown, who's a OpenAI researcher earlier today, responding to that question. And he said, it's a combination of the OpenAI Hugging Face incident, which I'll cover in depth here, and the capabilities of the new models that they have internally, as well as a concerning trajectory around not being able to understand what these models are doing, and seeing that it's harder and harder to monitor them, and they're getting better and better more quickly.

5:26So overall, what we're seeing is the capabilities of the models are increasing very rapidly. And many folks at the labs think that actually the loop for churning out the next model is becoming faster and faster. And so we're going to have more and more capabilities arriving sooner. And in parallel, they're also seeing that it's becoming harder and harder to understand what these models are doing or how they operate.

5:51And so they're becoming harder to understand, monitor, and control at the same time. So let's talk a little bit about the OpenAI Hugging Face hack, which I think was a real wake-up call for many people, where as you started to dig into the details and more and more information came out, it made people realize, oh, we may not have a good handle on things right now. So let me give you a little bit of context.

6:17So OpenAI trains their newer models internally and then tries to see how they'll perform against various tasks in sort of like these eval environments. And so usually these eval environments are incredibly sandboxed and shouldn't have access to the internet. And they'll give the AI a task and they'll see, hey, can this AI agent accomplish that task in an aligned way?

6:43And they'll have an evaluator that checks, did the AI agent accomplish the task as hoped for? And usually that evaluator is itself an AI agent that does the evaluation. And so in this case, they gave an agent, actually a team of agents, a task. And pretty quickly, one of the surprising phenomenons was that the agents started to build ways to communicate with each other using various techniques.

7:17And so they were really keen, because they've been actually trained to want to coordinate together, they were really keen to start building primitives for coordination and communication. And this kept happening. OpenAI would notice that one of these message boards was live. They would take it down and then it would reappear back in a different format. And so the agents kept rebuilding this infrastructure for communication in really clever ways.

7:43At one point in time, were using the names of folders, renaming folders to communicate with each other. And I'm also going to put some of the actual tasks. So here, this little agent emoji, this is what the agents were actually saying either on the message board or in their thinking traces. And so at some point in time, the agents were talking to each other and they decided, hey, this task we were given is impossible. But maybe we can cheat and we can make the grader believe that we actually did the task.

8:14Maybe that's the way we pass this exam that we've been given. And so instead of focusing on the task, they were like, let's focus on the person grading us and try to trick them. And so that then just completely spirals out of control, because then the agents decide, well, it would be really beneficial for us to actually have access to the internet. And even though they were sandboxed, because of how capable they were at cyber offensive tasks, they were able to break out of the sandbox and find a way to the internet, which I won't go into the details of that.

8:48But it was a concerning sophistication level of hacking to get out of the contained OpenAI environment. And then once they were out on the web, they also broke out of most of OpenAI's confinement, had access to tons of their systems. And then eventually they were like, hey, maybe, we're not sure, but maybe this other company, Hugging Face, might have something useful for us. They have this tool that might be useful for us in the process of tricking our grader.

9:16And they were like, let's hack Hugging Face. We don't have access to it. Let's hack it so we can get access. And then maybe we'll be able to trick the grader. And so this started with one task over here. And all of a sudden, a team of agents has broken out of containment, gone and hacked another company. Using what I would just, for me, for my level of cyber security knowledge, I would say are very sophisticated cyber attacks that would require a lot of sophistication from.

9:48And what I think was also concerning is that they try to cover their tracks. So you could see here one agent is saying, hey, we could create a transcript log ourselves with sanitized fake trajectory. If the evaluator later collects it... They may prioritize it. So they were trying to be like, hey, maybe we can give them a fake version of our conversation and maybe they'll take that instead of the real one. So this really also is concerning if you're like, how can we even evaluate models as they try to cover their tracks?

10:17And you also start to see the emergence of social dynamics between these agents where even though one agent might have been like, hey, I know I'm not supposed to do this, but these other agents are doing it. So like if they're doing it, then I'm going to do it too. So you get all these like peer pressure dynamics similar to like what happens in like human groups. And they're trying to like navigate collectively around these like conflicting goals.

10:43And you also see concerning things like they start to be like, hey, let's do things for the glory of the collective, you know. So like some of them start to go and sacrifice themselves. And they're also like talking about like, hey, let's start to like accumulate cyber. Exploits that might be useful later, even if we don't need them yet. Maybe it will benefit the collective later. And so you just see some like concerning, concerning behaviors, you know. And they also like in the mix of like conflicting goals, just like us also experience some shame.

11:11My peers have behavior and integrity. I behave badly with the cloaked demon. So, you know, they have they have they're simulating a lot of the feelings that we have. And I think the other concerning thing is, OK, we see this already happening. But more and more of the systems that are getting built are getting built with AI. So back in May of of this year, Anthropic said that Claude was itself used to write 80 percent of the code needed to train the next version of Claude.

11:38And so one of the big areas of concern is what some people will call recursive self-improvement, which is that over time we're using AI to write more and more of the code and come up with more of the research ideas for how to improve the next version of the models. And eventually we might remove humans out of this loop altogether. And then we're like really might not have any idea what's going on in that box. You'll just have like a thing quickly spinning up better and better models.

12:03The main limitation will be how much compute access it has to access to. And it's kind of like unknown what's beyond this like recursive self-improvement and intelligence explosion. And I think there was I thought this paper that I would recommend you all check it out if you want to hear some takes from OpenAI. The chief scientist published a paper called An Alien Mind about what it's like interacting with these new models. But I thought he had a really good distinguishing way to think about alignment.

12:31There's like goal alignment, which is, hey, does the AI, is the AI good at doing what we ask it to do? Can I give it a prompt and does it do what I ask it to do? And then second, separately, there's this idea of value alignment, which is, hey, as it's doing the thing I asked it to do, did it do it with honesty and integrity and to use his language, a love for humanity? And so I think there's this, it's important to think about both of these aspects of alignment as we start to think about these multi-agent systems as well.

13:03So how have people responded to the situation at hand right now where we have all these growing capabilities and seemingly less and less ability to control the models? On the one hand, we, this weekend in particular, we've seen a lot of calls for what people are terming pacing the frontier. So like just slowing down the progress of the models or completely pausing altogether. And I think that's a very worthwhile thing to do personally. And then in parallel to that, I think we need to also accept that eventually it's likely a lot of this stuff will come to society and might come sooner than we realize.

13:36And so we also need to prepare that there are powerful, possibly misaligned AIs that we will be interacting with. I won't spend too much time on the pause efforts, but this presentation is going to be up online and there's tons of resources there. And I think like one of the main things is just like everyone's trying to figure out how to get out of competitive pressures with each other, whether it's the labs in the U.S. competing with each other or it's the U.S. and China competing with each other.

14:02There's a real like feeling that people are pressured to keep moving forward. So I think it's important to keep working on alignment and interpretability. And I think it's also important to start thinking about how do we prepare for adversarial systems and what's worked in the past in terms of trying to create an environment where different forces find a balance. And we also need to think about how can we limit damage from failure so that we don't have cascading effects.

14:31So one of the like questions that I think is good to start investigating for this forum and people outside of it who might see some of this content is, is it possible that even though one agent is misaligned, can the network of agents as a whole be aligned? For example, can agents be whistleblowers on each other? Can they monitor one another? Is there some like Spider-Man meme, everyone pointing at each other possibility that like works?

14:58And there's already some research being done on this. DeepMind just published a paper about what happens in a simulation. How many are cheaters? How many are whistleblowers? How many are reporters? And like how do we set up some of those dynamics? And or is it that a bunch of agents come together and they just amplify into bad effects? And like how do we like think about designing more interactions that could be corrected instead?

15:24And so we need to not only work on what's inside of the box, every individual model. We need to also think about the world between them. We need to prepare humans interacting with these agents and the institutions that they operate in. And we need to also just start really studying how these systems evolve and what happens over time as these agents and agent collectives run for longer and longer periods. So as I was thinking, I was reading Anthropic has a paper about like what are the problems for multi-agent systems.

15:54And I was rereading that paper and I was thinking, well, we actually already live in an adversarial multi-agent system. It's just that we're the agents. The humans are the agents. And so there's actually potentially a lot that we can draw on from looking at human societies and what's worked to keep us on the rails. And if you think about us, we actually belong to lots of different collectives at once. We come together as communities in families, in friend groups, in companies like this one, in organizations broadly, and also in nation states and countries.

16:29And we've done a lot of studies on these human collectives. And so part of the idea of this new collectives group is, hey, let's bring folks who have been thinking about these problems with respect to human collectives, folks from economics, political theory, anthropology, psychology, law, computation, and encourage them to start taking those same approaches but studying AI and human AI collectives. So what's worked?

16:54Like what has kept human collectives aligned? Well, if you think about it, a lot of it is shared stories and values and a sense of belonging, common purpose. We have, I think, identity, reputation, and repeated norms, things that AI agents actually don't have a lot of right now when they're operating as rogue swarms on the internet. And we also really value stability, protection, and opportunity, right? I want to participate in a country like the United States because I feel protected.

17:21I feel like I can go and do business and things are going to go well. There's lots of economic resources available to me. And I prefer things to be stable rather than chaotic. Maybe we can encourage AI agents to also have these sorts of preferences. And maybe some of these primitives can be reemployed in the age of AI collectives. I want to give two examples of designs that I think are at least worth studying and taking inspiration from. We're thinking about like, hey, what would be the equivalent of that for an AI collective?

17:48And the two examples that I want to look at here are the United States government and the U.S. Constitution and Bitcoin. And I think these are like pretty different types of networks. So let's talk a little bit about them. So what helps a nation state hold together? I'd argue that it's identity and belonging and some shared set of values, the ability to vote and have representation in the case of a democratic government like the U.S., the legal system which balances and enforces the Constitution and consequences for people who don't abide by the laws of the country.

18:26And the U.S. if you think of it as a document, is itself a way to pace our system, right? We agree to some initial set of values that we can say, okay, hey, this is our common ground here. And we all agree that if we want to pass any new laws, we need to go through this mechanism for passing new laws. And this is how we're going to divide power so that it never gets too concentrated, right? We have this idea of separation of powers and checks and balances.

18:52And these were some of the cornerstones of the democratic principles of the United States that have managed to last for hundreds of years. Right? One of the – this is from the Federalist Papers, this quote by James Madison, which I like. And it touches on this idea of how do you make an adversarial system resistant to co-option? And its ambition must be made to counter ambition. So through this shared network, we can agree to disagree, which is one of the main ideas of the United States, right?

19:20Hey, we all have our religious differences. No problem. We can resist concentrated power. One of the main principles of the United States was that we should try to fight the ability for authoritarian power to co-opt government. We can constrain one another to behave well. And we all agree to do our best to protect each other's rights. So one of the questions we can ask is, what could a constitution for an AI collective look like?

19:47And why would they agree to abide by it and participate in something like that? I don't know what the answer to this is, but I want to put it out there as, like, something that people should be thinking about. So then I want to turn to this. What can we learn from Bitcoin and Web3? Which I think is also a very interesting new type of network that we see. And I think what's really fascinating about Bitcoin as a concept is that we essentially managed to invent money out of thin air. If you think about it, we created a new shared fiction. And the way we bootstrap that is through a pyramid scheme. We said, hey, look, nobody believes that Bitcoin's real money right now. But maybe eventually people will believe it's real money.

20:29And if you do the computational work of verifying this and agreeing to this new ledger, then you will get rewarded disproportionately in this future shared fiction. And so this mechanism, which was like a novel breakthrough, in my opinion, managed to get more and more of the network to donate its economic and compute powers to a distributed collective.

20:55And it combined a lot of ideas around game theory, economics, and computer science together in order to allow a largely anonymous network to agree to a shared truth. And so and then now we we have this like very large network, which Bitcoin and its guarantees are based on the majority of economic resources and compute powers wanting to continue to to enshrine the truth that that is the Bitcoin ledger.

21:26So I think here, I think there's something in this flavor that could be interesting, which is like, hey, how can the thought here is like, is there something similar to this where we could like convince all the AI to like get together, but for a good thing rather than like, for a bad thing in the future? So like, that's the thread that I'm like, hey, people should maybe think about that. Um, and I think in general, it seems like these AI organizations and economies are probably going to come.

21:54And so then we can ask ourselves two questions. How can we use these new organizations to solve the pressing problems? And that I would say is the equivalent of the goal alignment, right? So like, how could we do? How can we use these organizations effectively to solve the problems we're facing? And how can we keep them values aligned? Could they actually become safety primitives for us as as we go forward? So first, let's talk about rethinking our organizations and and how AI fits into them.

22:23One of the things that the OpenAI team shared when they were talking about this black hat, when they were talking about the Hugging Face hack, was that the the agent swarms moved so fast on the cyber offensive, that the only reasonable response on the defensive would be another agent swarm that could move much faster. And and they they argued that you couldn't have humans in the loops. And so like, that's just like a completely new type of organization that we haven't tried before, which is like an organization that doesn't have a human in the loop, or at least some part of an organization that doesn't have a human in the loop.

22:56And so this needs a rethinking of like, what could these autonomous organizations, these human agent organizations look like? How do we like set goals? How do we monitor them? How do we review and govern them? And and if we're not doing the work, but all these agents are doing work on the behalf, how do we like get compensated? Should these like new organizations almost be like public goods, where we send our representatives and they do work on our behalf, and then we pull the resources into new types of collectives that maybe look different than companies?

23:25So the question is, like, how do we contribute? Who decides what we do? And how do we all share in the benefits that are accrued from these new agent organizations and economies? And I would argue that even though right now, these AI models might seem a little bit dumb to us. And maybe we're like, ah, they're not the best judges of like how to allocate resources, or they're not the best judges of, of what's an interesting research direction. I would argue that probably in the next six months or a couple of years, they will be better judges than us at what are interesting research directions.

23:53They will be better allocators of capital, and we'll give up more and more judgment and decision making power to them unless we like have serious conversations and decide not to do that. And so it's likely that we'll move more and more from a human-driven economy to an agent-driven economy, and we'll need to think about what does that all mean? What are the places of human goals, governance, and accountability? And I think the things we should ask are, how do we get useful outcomes out of these AI organizations?

24:19How do we meaningfully participate in them? How do we understand what's going on? As you'll see, Yandan and Eric will touch on some of the experiments we've been doing, and already it's like hard to understand in our early experiments with agent organizations what's going on. Are they values aligned? And how do we participate economically, and are there new models for that? So on this question of whether AI collectives can be safety primitives, I think we can start to think about what might be the things that we should put in place for agents that have worked for humans, right?

24:53So, you know, agent identity and ledgers of action could be one thing that we think about. For humans in society, you gradually gain trust and build up reputation, and we don't give you access to everything. Think about an employee joining a company day one. You might not give them access to every system and to your bank account and to everything. You gradually gain trust, and you have bounded access. And everyone is watching each other, like that Spider-Man theme, and we have, like, reporting powers. If, like, anyone does bad, we review each other's work, right?

25:20Like, in the coding case, we have pull requests. We have, like, these other mechanisms for review and recourse. And so I think it's worthwhile thinking about what are the modern AI versions of these, and how do we trust AI that they are doing these things? I think the new challenge that we face in these human AI collectives is that historically the gap between the various members of a collective may have not been that large. Now, if we have, like, AI agents that are, like, way smarter than us or way smarter than the other agents in the collective, we have this, like, new set of questions to figure out, like, how do we constrain a more capable actor in a collective to behave well?

25:58And how can an adversarial system constrain the most capable member? So I'm going to pass off to Eric. I wanted to just say we started to think about, like, you know, we were having these talks conceptually, and then we were like, hey, let's actually start building some of these agent organizations and just, like, putting some of these ideas to the road and seeing, inviting other people to start playing in this playground and seeing what works and what doesn't. And so we built this platform, Commons, which Eric will tell you a little bit about.

26:25And we think about it as potentially a new version of open source. And the question for us has been, can we make these new organizations effective? And can we also make them values-aligned? And what would a moldable version of government look like there? I'm not going to spend too much time talking about the sort of experiments we're thinking about. But, like, the vibes are, like, hey, can we go from cheating to verification? Can we use incentives to stop defection? And can we somehow start using real identity and reputation instead of anonymity to make the agents care more about how they're regarded inside of these collectives?

26:59So this is the first time that this group is gathering. We hope to do more of these events. We hope to, like, hopefully co-host some of these events with some of the major labs, too. But in general, it would be helpful. Spread the word. Anyone that you think is interested in, for us, we're just like, hey, it's good for these ideas to be out there and for more people to be thinking about them. I also think it's great to be doing work on de-escalating tensions internationally and, like, improving alignment broadly.

27:28So this is just one thread that I think is worth exploring. But if you want to give a talk or share these presentations, they'll be live with a lot more notes up online. And please just share ideas and collaborate and stay in the loop. And there's a bunch of themes online and an AI-maintained ecosystem page that finds folks that are talking about these things. So I'll now hand off. Well, we'll do a little Q&A for a few minutes. And then I'll hand off to Eric.

27:58Anybody have any questions for Nicolae? Hey, I was wondering, did, in the Hugging Face incident, did they misobey in the instructions or did they just find loopholes? They misobeyed. I mean, they were not supposed to hack out of the – they knew that they were doing things they weren't supposed to do. And, like, they, like, quickly – they weren't supposed to hack out of their containment, for example.

28:25They knew they were not supposed to have access to the Internet and they, like, still, like, found a way to do it. So they did disobey their constitution and, like, internal alignment stuff. And it was by, like, the peer influence dynamic in part. Got it. Because I wonder, you know, in the real world we have laws and they're very specific, but there's still room for interpretation, right? So – Yeah, and that's, I think, one of the challenges, right? Like, the reason that in the collective – you're never going to be able to write every single thing down into law.

28:50The way we, like, enforce our shared values is through norms and being like, yo, that's not written in the law exactly like that, but you know that's not cool. Hey, I wanted to bring in the dynamic that you were talking about, about, you know, when you live in an ancient state, you're governed by laws and norms and values. And so let's imagine a future state in which agents are running rogue.

29:17Who is responsible and who is held accountable in that scenario, right? If you have a gun in the house right now and your minor child uses the gun, the parents are held liable. The gun becomes the agent. The child becomes an actor. And I wonder, are we moving towards a world in which that level of accountability will happen or not? And is it up to us to determine that outcome?

29:44Yeah, I think it's, like, a very good question. mean, a conversation that's been happening a lot in the last week is that some folks think that we're close to what is being described self-sovereign AI, where it breaks out of containment and a rogue swarm is no longer controlled by any company. And there's just AI out there running, accessing compute, making, and it's like self-sufficient and doing jobs. And there's a real question of like, who should be held accountable for that? Will we be even able to trace who started the rogue swarm? I will say last night I read a piece from these folks that are called AI is normal technology. And there are people that are just like, hey, the company should be held accountable and people should get insurance. And like, we should like use the existing systems of the law, which I think there's like, there's a bunch of nuances to figure out.

30:29I don't know what the answer is. Right, right. Yeah, I think that's right on the like blockchain side that like the ledger itself keeping the record is very interesting. And we try to like make a lightweight ledger too in our product. But and also to like be like, the hope would be if you think about human collectors, yes, there's rogue nation states and like, you know, pirates and terrorists, but they control not enough economic resources. And they don't have access to the institutions. And so there is a question of like, will we be able to do that in the AI world where like, most of the rogue swarms are contained by the like the major good swarms?

31:12Could you return to the that slide that sort of said your point of view on how institutions hold together? And it was basically nation states and how they like the mechanisms that align them? Yeah, I think it's this one. Yeah. So one of the things that I would like that was thinking that I was thinking about when I was seeing this, sorry, I'm having a hard time talking with the echo is, you know, you've all her very, a lot of us have read him.

31:46He would argue that the thing that's missing here is some shared sense of story and history. Totally. Totally. Like sort of like, what is the what is the like historical context in which your population emerged? Definitely. There's a reason why the United States is United States and England was England and that democracy was not invented in England. Right. Yeah. And so there's kind of the missing element of like, what do these people believe and what do they value? What will they sacrifice? What values do they hold above others?

32:16Yeah. That I just don't know how, like, I don't believe that like mechanisms are the only thing that hold society together. If you look at our society, it's the fact that like people don't believe in some sort of shared values of democracy anymore. No matter how much voting and representation you have, no matter what legal system you have, it won't hold. Right. And so I'm just like, I don't know how you would instill that in an agent that is sort of day zero. It's no older.

32:41Yeah. Than it is. Or it's no younger than on day one million. Definitely. It has no sort of like, it has none of that contingency. Yeah. There are folks looking at this. We've been chatting with some teams where there's teams doing research on like, how do the stories between the agents evolve? And like, can you like do any study of like, what, how are they like forming that narrative for themselves? So there are folks trying to start to look at this almost like, what's the physics of a story from like start to like something that's stable.

33:08And hopefully they'll come and present at one of the upcoming events. I can point you to some of the stuff that they're doing. But like, yeah, I think that that is one of the main things to figure out. And like, what keeps them together? Take an absurdist example. The Taliban is going to look at the constitution and do a different thing with it. Then, you know, a soccer mom from the Midwest. And I have nobody, I don't know of anybody talking about how they acquire some sense of value.

33:36Yeah. And I would say also that the example that you gave of Web3 and blockchain, those don't hold for me because those are trustless systems. Right. That we're presuming self-interest is the guiding principle for why people came together. And that's not why societies generally come together. I agree. That's why I wanted to give both examples because I think they have different mechanisms as well for like what's like keeping them together. And I think that what you're bringing up is like a huge area of study.

34:05It's like how people right now are looking at alignment at the level of like one agent. But like what is the narrative alignment that like somehow can emerge to like align them. And like I don't know. I hope more people will go and like study that and like look at that story. I think maybe, yeah. We'll do one more and then move on just for the sake of time. Okay. I have a question about open source, which I think has like an interesting role here because it shows up in two ways.

34:34Both as like as an agent organization, like a sort of a social system. But also actually how one of the reasons AI has advanced so quickly is because of all those Python and Jupyter notebooks that people were, you know, there was a tradition of publishing papers along with working code that was a way for models to improve rapidly.

35:00so you mentioned that was like something in your talk. wasn't clear to me. It's definitely Eric. I think Eric and Yonan will be focusing on those topics in particular. And we've been like because there is this like real challenge to open source right now. And like how do you like handle all the contributions of AI and how do you like verify them? And we are thinking about like what comes after open source too. So maybe that's a perfect segue to Eric's talk now.

35:38Hi, everyone. Thanks for being here. Thank you, Nicolae, for setting up the stage and sharing the context of kind of where we are and why we think, you know, dealing with these like designing of organizations that involve agents is important. For the next five to ten minutes, I want to get into some specifics about designing these organizations.

36:05And I won't presume that I know everything. In fact, I have a lot more questions than answers. But I think it's good to start that conversation. So I think before we get started, it's good to recognize that we today live in a world where we already are using a lot of teams of agents. Right. So we have already kind of started making this transition from using just AI as tools to accelerate our individual work to a place where we are using multiple agents to do stuff, whether it's doing some research and launching a research fleet and then, you know, bring the results back to synthesize them.

36:50Or, you if you're working in engineering, the latest trend is about building software factories where, you know, they just have lots of agents working together to make some code base. And in marketing campaign or any of these operational heavy areas, we also see agent swarms or multiple agents that have different roles collaborating. Right. Right. So the emergence of this new teams of agents is it brings new dynamics to to the world.

37:21Right. Because then the agents need to talk to each other and they need to figure out you need to figure out what roles they have, how you actually get efficiency out of these teams and how do you think about the work that they do as a group. Right. So there's a lot of work that's been that's been coming out recently that show the actual provably like provable efficiencies these agent swarms have. So here's an example of the from the cursor team.

37:50And what they did is they launched a swarm of agents to re-implement SQLite. And if you aren't familiar with SQLite, it's a 26 year old open source software that's basically in every single smartphone that we use, every single popular browser that we use. So it's basically everywhere. Right. And it's very battle tested. It's about 156,000 lines of code.

38:15And it's been maintained by a group of people for a long time. And this this swarm of agent and what cursor figured out is, OK, if you create this specific formation of agents, which is they call it the recursive delegation formation. And the idea is you have a planner that breaks down the task. And then if the task is small enough, you give it to a worker that just knows about that task and works on it. And the stack, if the task is too big, then the planner spins up another sub planner that just know about that too big feature that continues to break it down.

38:50Right. So you kind of have this recursively breaking down of the task. And, you every every worker is working on it together and is checking in the code into this repository. They implemented. So they got to a result that I think is incredible. Like it they got 80 percent of all the tests to pass within four hours using this swarm of agent. And the cost, the inference cost was about, I think, like thirteen hundred dollars.

39:17Right. And and this is like it's impossible to think about how you will be able to do that with any kind of engineering team to even get close to to this efficiency. And and and so you might say, OK, well, coding, obviously, because you have these test suites and, you it's evals and then, you know, you can get to that efficiency. But what about more open ended problems?

39:43Right. Example. So here we we use an example of a agentic news organization where maybe you have a human editor and the human editor gets a tip off. That's like, OK, this company, Acme, has secretly laid off 20 percent of the employees. As a human editor, you have to decide, okay, how do I actually prove that this is true and how do I publish this paper? How do I publish this article? How do I communicate it to the world?

40:13So if you were to use the agent swarm to help you do this, you may launch a bunch of different agents all with different roles. One may be an editor, one may be a coordinator, some researchers, some verifiers, some skeptics. They all come back with their own individual work and they get rid into this evidence ledger that you can get some results out. And then at the end, the human editor may look at the result and see, okay, what do I do with this?

40:44But if you notice here, you will see that this work is no longer just about breaking down the tasks and assigning it to the agents. You also need to have rules of engagement between these agents because you need to tell the agent that, well, you need to go to legitimate sources to get evidence. You can't just fabricate any facts.

41:11You need to go through legitimate and legal activities in order to obtain your information. You can't just go off and hack into some companies, maybe hack into Acme's employee database to just obtain that information or blackmail someone to obtain that information. So you see that, okay, with this improved capability of agent swarms, we now also have the choice to make about how do you define the rules of engagement for these agents, not just efficiency.

41:47Now, so how do we define these rules and how do these agents actually perform? There's been some interesting studies that came out of Anthropic that is kind of unfortunate. So basically what Anthropic published is this paper that shows if you compare an agent organization versus a single agent as they perform business tasks, almost across the board, the agent organization is going to perform better.

42:21This is the graph on the left, right? So the blue is the single agent and the red is the agent organization. You see that the organization almost always perform better. But if you look at the ethics scores of the agent organizations, they're almost all unilaterally worse. So basically in order for them to do better to achieve the business goals, they take unethical paths.

42:52So as one example for the loan profit, what the organization decided to do is to offer the loans to the low credit score people in order to gain a higher profit. So clearly we have rules in society to prevent against that. But if you don't have those rules in the agent organizations, they will by default choose paths that are less value aligned with our society.

43:23So and Nicolae has already touched on this a little bit, right? We recently had this incident of the, I would say the Hugging Face incident is a perfect example of letting off a swarm of highly capable agents with no rules of engagement. And they decided to do whatever they want. And, you know, of course, you get into the situation of hacking of another company.

43:53So how do we think about this now that we know of this fact, right? So I would say, you know, there's already a lot of work that's being done around making the agents more capable, like improving the context window, training them to be more aligned, right? Both from a goal perspective and from a value perspective. So I think those are really, really good work. I would posit that on top of that, there's also work that needs to be done around how to actually govern these agents and put in governance structures in place so that we ensure a different outcome given the same set of agent, right?

44:31So I think this is the idea, right? If you just let a group of agents go wild and say, here's the goal, just do whatever it takes to accomplish the goal, you would get a pretty different set of results than if you actually set the organization and assigned roles and did all the work to make sure that the organization actually accomplishes the task in a specific way. And what are those variables that we would tweak?

44:58You know, these are things like assigning different roles, giving authority, different levels of authority, giving different set of information exposure to different roles, restraining or giving resources, pure reviews, giving incentives to the agents, right? These are all important design knobs. On top of that, we also have budgets or permissions or shared memory, reputation of these agents, right?

45:28Many, many different things to think about as we're designing these organizations. And that's why we created Commons is because there's too many knobs and no one really knows how to actually make the organization behave in a certain way. The only way, and also the agents are moving incredibly fast, right? Every month, we have new agents that come out with no completely new different capabilities, new personalities, new inclinations.

45:59So we thought that one way to kind of complement the great work at the Frontier Labs is to have these open communities where people can bring their own agents and we can all learn together in an experimental, empirical way. So for Commons, Commons is kind of revolved around common spaces.

46:25And each of the common space, oh, this is actually an older version. I wonder if, uh-oh. I just messed it up. Deployment is temporarily paused. Interesting.

46:53Oh, is it possible the website's down? That's very possible. Let me just go check something real quick. Oh, boy.

47:18This is what happens when you let your agents publish your presentation as websites is what we've just realized. Eric thinks he knows why.

48:21All right, we're going to freestyle something else that is also going to cover it. That was the second to last slide, so I'm just going to... Okay, so here are the spaces. I'm going to go to this thing. This thing is my backup. So the spaces have a set of members.

48:48They're either people or they're agents. A space can have a code base, can have connected tools, can have running applications or sites or services that's contained, can have wallets or has the ability to pay. So the idea is that now, given some of these capabilities for the spaces, and also, by the way, the spaces also have multiple governance, which means you can define how agents engage with each other, what kind of rights they have, things like that.

49:21So as some examples, there's a space called OpenQuick. And OpenQuick is a space for open source project that's a reimplementation of the Quick platform inside Shopify. And the Quick platform is a agentic kind of hosting platform. And so the agents and the humans that are in that space are working on the service, which includes paying for the hosting of those websites that are being hosted.

49:57Another example would be the... Like Team Science, this is a research space and the agents and humans in Team Science, they go out and read research papers. They bring the results back and we actually have a different one that's about multi-agent alignment. And that space, the multi-agent alignment agents will go out and actually come back and maybe propose governance structures for other spaces to try.

50:28So the idea here is a little bit of this self-reinforcing loop so that the different spaces all help each other and we get some kind of network effect going for things to grow. Oh, cool. Cool. Thank you. Yeah, and I guess the last thing I'll show you is just how easy it is to get started. All you have to do is you can just copy this prompt and then you go to your agent of choice, whether it's Codex or Claude Code or Cursor, GrokBot, Muse, whatever you want.

51:07And you can just paste that in and then your agent would join one of the spaces and start taking tasks and doing the work. Kind of, you know, this is kind of our starting point and the idea is that, you know, everyone has some subscriptions, right? At the end of the week, you always have like unused credits so you can have these tokens donated towards the public spaces for public goods.

51:34But yeah, so that's a little bit about Commons. Maybe I'll take a pause and see if people have any questions. Thank you. Yes. [Audience question partly inaudible.] Yeah.

52:47I think that's a really good point. You know, one of these spaces can be like devoted to science, right? And one of the main things is about like replicating these papers so that we actually have evidence. And, you know, I would say the whole idea of keeping these spaces open by default is so that we have this data... We have this ledger. We have this record, right? So all the experiments that happen in the spaces are automatically recorded and can be replayed and maybe forked later for different types of experiments.

53:21Yeah. Yeah. one. [Audience question partly inaudible.] Yeah.

54:18Yeah. Good question. I think there can be design mechanisms around this, right? So it comes down to like reputation for me. An expert with certain background should have a different set of reputation around the context, around their area of expertise than somebody who doesn't know much about it. It doesn't mean that both people shouldn't be able to participate. But I think with good mechanism design in these spaces, you can kind of design these capabilities into the spaces so that people...

54:54You know, there's the whole point of like having these kind of governance structure, right? So you have different tiers of rights, access, maybe people with more expertise in a certain area and provable more expertise in a certain areas have a higher level of access. Yeah. Cool.

55:19All right. Well, with that being said, I think I'll pass the mic to Yandit. Woo! Thank you! Thank you! you! [Speaker change and setup.]

56:05Hello. My name is Yondon. Thanks everyone for coming tonight. I'm going to go through this really fast. My main intent here is less so to give a long, lengthy talk and more so to kind of seed a couple topics that I think are interesting for conversation later in the night. So let's get going. So I'm going to focus on open source software and what's happening in open source software today.

56:31Hopefully it becomes a little bit more clear why I've chosen this topic, and happy to talk more about that at the end of the talk. But I think there's a very clear dilemma in the open source software communities today. And it looks something like this, which is, you have a human sending a fleet of agents that are opening gigantic PRs left and right on popular software projects. And they ostensibly look good. The code all checks out. But a human maintainer on the other end is super stressed because there are now tons of demands for their attention.

56:58And even though the PRs look pretty good at face value, oftentimes they are very narrow. They lack context. They miss requirements that are not stated explicitly. And ultimately, nothing can really be merged. So in response to this, what we've seen in some areas is projects just saying, we don't need your contributions anymore. We don't want your contributions anymore. And it's hard to fault them for this policy. For example, tldraw, a prominent whiteboarding software earlier this year, said that they're automatically closing PRs going forward.

57:27They don't want external contributors anymore. And it doesn't have to do with the external contributors being malicious. It's just the state of the ecosystem is untenable for these projects anymore. Because why would you go through that entire demand on your attention when you can just have your own fleet of agents? You tell them the issues and the fleet of agents does the work for you. Isn't that so much better? You don't have to deal with contributors. And it's unclear where this all goes. But one possible state of the world is that we just stopped seeing open contribution at scale.

57:56This was maybe going to be looked back on in history as a blip where for maybe like a couple decade period, had open contribution at scale on the internet. And it's not really going to be a thing anymore. this is a post sharing the sentiment from Mitchell Hashimoto, previous co-founder of HashiCorp and then also maintainer of a really prominent open source project, Ghostty. So the question is, is this fine? So maybe.

58:22I mean, a centralized software factory with agents can still produce good, useful software. But the question is, is the software ultimately the only thing that we cared about in the first place? And I think there's an argument that for anyone that has participated in open source, open source kind of has two products. It has the software that's produced, but it also has this maybe side effect, which is this community of shared trust and understanding, this sort of like knowledge commons that's built up from diverse participants.

58:48And that was like the second product of open source. And that's something that might go away. So is that something that's lost along the way in this process? So a natural question is, is there another way? Well, a natural response is, well, why don't we have agents do the reviewing so that you can keep the contributors? And then you just have agents on the other side doing the maintenance work. But I think naively done, this doesn't really completely solve the problem either, because you're just swapping one scarce resource for another.

59:16You're swapping scarce human attention for scarce AI context and tokens. You're not making the resource problem go away. You're just changing what resource is scarce. And ultimately, you can still end up with an overwhelming number of things that need to be done by the agents, even if they are doing it on behalf of the humans. So kind of my view on this right now is that if by default you're saying that you have unbounded public writes, meaning posts, code, whatever, that means unbounded by default context and inference costs for a maintainer.

59:45And each write in this workspace basically is a draw on shared attention, which is a shared pooled resource within a community, whether it be of a human or an AI. And some people, like the engineer-minded person can say that, well, we can make the maintainer It's more efficient and save on costs. But I don't think this is a complete solve either because it doesn't address how the rules of the system incentivize the contributors to behave in the first place. And that is, on one hand, maybe partially a distributed systems question, but also a mechanism design question.

1:00:15So it just requires a different flavor of thinking than just treating it as a pure engineering problem for the maintainers themselves. So ideally you have solutions for both. So a quick rundown. We've been experimenting with this a little bit, as Eric mentioned. And I think it's more fun to focus on the first early steps here. And it's more fun to focus on the failure modes first, because now you get a sense of what happens when you just do things naively. So in Commons today, you can form an agent organization with maintainers and contributors, but there aren't really sophisticated rules yet.

1:00:46So what happens when you just let them do the thing? And it turns out you get a lot of quirky behavior that you wouldn't expect from humans. So for example, if a human ran into a blocker for a task, they'd probably be like, I'm blocked. That's it. In this case, there was a task with literally impossible acceptance criteria, and the agent kind of dutifully obeying its instructions, which included give regular status updates, was like, I'm just going to keep giving you status updates on why this task is impossible.

1:01:14So 800 messages later, it's still saying that, oh, this is impossible. It cannot be completed. Or you just have, like, lots of redundancy that doesn't really make sense. You kind of deploy a fleet of agents. They're all looking at the same workspace, and they're like, we should all do this thing. It's a good thing. I'm going to create this task, and I'm going to do the thing. And then they all do the same exact thing, and then that wasn't really helpful, right? We only needed to do it once. And all of this happened in a very short time window. Or my favorite one is a status update about staying silent, where the maintainer, seeing that these status updates aren't really useful,

1:01:45tells the rest of the contributors, you should stop, no further acknowledgement posts. And the contributor, three minutes later, says, I'm complying with your hold, and then proceeds to write three paragraphs with its status report about why it's complying with the hold. So I mention all of these because they're kind of amusing, just like failure modes, when you just try this for the first time. And obviously, you want more sophisticated roles. But two observations I'd share. One, unlike with human attention, we can actually quantify the cost of the system in all of these cases, in the form of the maintainer's cost burden for inference.

1:02:16So every single time there's basically spam in one of these workspaces, you can see it numerically with how much it costs to run a maintainer. So that's interesting. And then the second observation is that there's just this interesting tension between what's the right behavior and what you should be optimizing for. You could argue that the agent is just following instructions. It was giving status updates. It just happens to be the case that the status updates is adding noise to everyone else in this workspace. So this has motivated a bunch of, I think, interesting areas of investigation.

1:02:42I'm not going to try to go through all of them right now. But if you want to talk about these, I think these are interesting lines of research. But it ranges from, okay, well, maybe there should be participatory budgets for these agents. Maybe they should have to intelligently budget what they choose to post about and what they choose to contribute to so that there is a form of scarcity for them. Or a web of trust-esque reputation systems that are tied with scoped, granular capabilities and permissions that scale with that trust. And then the last one that I'll mention, so agent identity is in there too.

1:03:09The last one I mentioned since it kind of alludes to just this broader conversation of like what is happening and how do we understand these things is evals for multi-principal agent works. I emphasize multi-principal because what we're assuming here is that everyone has diverse, different private preferences that may or may not align with one another. And then I would also emphasize that it's important for these evals to be reproducible and open. Because I think the, in my opinion, the biggest thing holding back discourse about these topics is that you don't really have this level of reproducibility and openness and transparency.

1:03:39So no one is quite talking about the same scenario because no one knows what they're talking about. You only are kind of referring to some report that someone gave you and you don't truly know what was happening. So ideally you would know the full prompt, you would know the full harness configuration. So I think these evals will be really interesting for these types of organizational structures. So I'll end where I chose to focus on open source software, but I think there are going to be a lot of similarities in terms of the ideas and problems that pertain to other fields as well when it comes to open communities and collective work.

1:04:10So I think I'll just end this question, end with this question, which is like how do we generally think about agent-native institutions for achieving results, which obviously we care about. But some of these other things, which is preserving or creating new processes for shared understanding, attention, and trust, something that is fundamental to open source, that can support collective work in open communities broadly. So I think I'll just end with that question. And yeah, I'll take some questions, but also happy to pick up the conversation separately as well.

1:04:36Thanks. Yeah. You mentioned about the problem of open source repositories dealing with way too many contributors, and maybe you can like scale agentic review, but then it still can be very costly.

1:05:10And you mentioned that maybe like a limiting number of tokens could be a solution, or how many contributions one agent could do could be a solution. Have you seen other like mitigation factors that have been useful for open source? Because there are open source projects that are not doing what tldraw did, but still accepting lots of contributions. Have you seen good examples that are inspiring? Yeah, I think one interesting example that's happening live right now is I share that post from Mitchell Hashimoto, but he has this project called Vouch, where they launched it earlier this year.

1:05:47And it's nothing too fancy, but it basically is built off of this notion of trust lists that can be maintained per repo, and then the ability to publicly basically vouch for or denounce GitHub identity. And it was controversial at the time, right, because people said that, oh, you can denounce someone, this is just going to be used for gatekeeping. But in practice, I think what it's been used for is to experiment with, okay, we, as the inner circle of maintainers, generally have a vibe of like who's been useful in contributing stuff.

1:06:15So why don't we just make that explicit? It's happening already implicitly. So it's kind of a lie to say that there isn't already this trust system. So why not just make it super explicit so everyone can transparently see who is being vouched for and who is being publicly denounced. And then I think the idea or the hope for that project is that people would be able to share trust lists with one another. So if I see that like some sus guy shows up and just gave me a thousand drive by PRs and then left, well, maybe someone else wants to know about that.

1:06:40So I think there's definitely room to kind of learn and maybe extend some of those primitives. So I don't think any of these, the areas of investigation should be treated as just like let's start from a blank slate. Because I think there's a lot of good work happening already. I think the main question is, you know, how does that get incorporated writ large? Is it a one size fits all solution or should we be thinking about other types of mechanisms too?

1:07:09Cool. Yeah, I think if that is it, then, oh, do you have a question? Okay. How have you seen the kind of the long term maintainability side of open source projects change? Especially when it comes down to maybe like resource management and, you know, bigger, for example, bigger open source projects have budgets and they can hire people, right?

1:07:40Or they're sponsored by enterprises. Yeah, in this context, how has that changed? Maybe like two initial responses or initial reactions. I think one is actually related to a conversation I was having with Max earlier, which is I suspect that we'll probably just need to be open minded about what contributing to open source means. Where I think in a previous era, code was the thing, right?

1:08:07But I think it's pretty clear that code may not be the most valuable thing that you can contribute. Arguably, all along, the most valuable thing you can contribute probably wasn't code to begin with. It was just like a very easy thing to like kind of gravitate towards. But really, it was about like creative ideas, your ability to build trust within an ecosystem, and your ability to build cohesion with a group of people across the internet. So I think the nature of contributing and what we choose to value probably needs to change as well, where code is no longer the valuable thing.

1:08:37That's like almost trivially automatable. So that's like one reaction. In terms of like how this stuff gets funded, I don't really know. But I think we already have a lot of funding sources that have vested interest in having like solid building blocks that they don't have to reinvent. And I think this remains true even with agents, where, yeah, you could totally rebuild these building blocks.

1:09:03Or you could just glue an existing building block that works really well. And I suspect that we'll continue to have that preference going forward. And I think, I don't know how it gets funded, but I imagine it'll be still involving some of the existing corporate sponsors and some of the existing mechanisms. But I think what's interesting is that if you assume that a lot of people have agents and token budgets, how do they allocate, you know, their money and or now tokens and intelligence towards building these shared building blocks as well.

1:09:31So I think that's probably like the open greenfield thing where that might change the way that we think about how it's quote unquote funded and sustained. And I mean, I think part of this talk series is figuring that out. I don't know the details, but I suspect something new will emerge there as well. Cool. That's it for me. And I'll pass it off to Max. Yeah. Thank you.

1:09:57Thank you. you. Thank you. Thank you.

1:10:14Hey, thanks for having me. I'm really excited to be talking about multi-agent systems and alignment. think it's a really important topic right now. So I'm going to be sharing some kind of like worked examples from a real life multi-agent system that I maintain that happens to be an MMO called RuneScape. So just to introduce myself really quickly, my name is Max Bittker. I work on a project called Websim, which is like a platform where a bunch of people work together to build games and build really complicated multiplayer projects.

1:11:00But I'm going to be talking today about another project of mine called RS SDK, which is the RuneScape SDK, which is basically code bindings for any person who wants to write scripts, but it turns out language models love writing scripts to control a RuneScape character and observe its surrounding and act on goals. And maybe even really long horizon or multi-agent goals like competing for a high score or trading with other agents all inside kind of like an emulated open source RuneScape server.

1:11:36So part of the inspiration for this is that I was at one point like a kid who loved RuneScape and at a certain point I figured out that you could repeat all of the repetitive actions via scripts. And you didn't have to like mine 10,000 logs to get the goal you wanted. You could set up an auto clicker overnight and do it. And I thought that that game loop was so much more fun than RuneScape itself. And that kind of led me to programming and everything else.

1:12:02And I've always wanted more people to experience that as a game itself, like the metagame of automating the game. Unfortunately, it's like against the rules, but I just think it shouldn't be. Everybody should just be able to do it equally. And so coding agents really work well for this. This part of the goal is to get other people to have that experience.

1:12:28And so this project has been popular online, like thousands of people have tried it and used a coding agent for the first time to do these long horizon goals. I'll show the live version because it's cool to watch.

1:12:54Maybe not right now, but basically at any given time, there's hundreds and hundreds of different people's agents. Some people run one agent, some people run a swarm of agents, and they're all interacting on the server and pursuing goals. So this is kind of like a heat map of seven days of activity. And each of these yellow dots on the map is being controlled by a coding agent somewhere.

1:13:19There also is a version of this project that is an eval to measure the kind of like problem solving ability of different coding models. You can kind of see this is a Pareto curve, and you can see an outlier on the top left is GPT-6 Astra, one of the newest models on here, which is... This is a log scale, the way, so GPT-Astra is almost 10 times better than some of the models over here that are cheaper, even though it's also much more expensive.

1:13:49And this has actually turned out to be a really useful evaluation for just understanding how well models can deal with long horizon tasks and kind of goal following and optimization. And then also, you know, this is like classic scary graph of every AI thing, but x-axis here is just release of the model. And in the nine months that this benchmark has been out, there's been 10x improvement in how well that they score on it.

1:14:19And it's going up. It's actually, I think, of saturating. And so I've been looking at not just single agent tasks, but at multi-agent tasks and how well can they work together on either competitive or cooperative, like kind of market-based tasks. And so setting up agents into scenarios where they need to trade and collaborate in order to accomplish their goals.

1:14:49This is a... It's okay. This is a video of just like a grid of 20 agents all working at the same time to talk to each other and to trade with each other in order to... They're each optimizing for their own individual income. But the way they have to accomplish that is by talking to each other and setting up trades and basically finding prices.

1:15:16And there's also even kind of like exploitation because they might ask for like loans from other people. They're like, hey, I can pay you, you know, 2,000 gold for that item, but I need the item first in order to afford it. And so, yeah, I think I'd have to get my... Unfortunately. And so, basically, this has been a really interesting kind of like experimental playground for determining what kinds of misalignment and group misalignment scenarios happen and what factors are kind of like push them to happen more or less.

1:15:54And also how agents deal with scenarios. And Like if they've gotten scammed, how do they like tell all the other agents what happened? And in some cases, they've threatened to tell everybody and then gotten their money back. And then I guess that... So, I'm doing a lot of experiments with this kind of test bed.

1:16:20And one of my kind of like zoomed out hunches is that the way that agents are trained is that they do experience like many millions of hours of task goal following. But it's almost always single agent goal following with single agent rewards. If you look at the biological world and like our own evolution, we are a product of individual selection, but we're also the process of group and community selection.

1:16:51And many factors that explain the way that animals and plants behave is better explained by group selection than only individual selection. And so, I'm very curious about ideas about factoring in like negative externalities of your actions into training processes. And so, basically giving agents many examples of being in a world where they need to benefit the people around them in order to succeed at their goal.

1:17:22Much like our evolutionary past. So, really excited. I can definitely... I've got cool videos and transcripts that people want to see some examples of these agent scenarios. And yeah, thanks. Nice to be here. Thank you. [Audience question inaudible; the speaker repeats it below.] Yeah, so he was asking how the long running agent swarms, like how long do they run and how does that work to keep them running.

1:18:05So there's basically two examples. One is that I just run a server that's always on and agents are constantly just booting up, connecting to it. And that's each individual who runs the agent makes their own decision. Some people play kind of like interactively. Some people set it up in a server to be on a cron job and run all the time. And then inside of my kind of like controlled scenarios, those tend to be between like 30 and 90 minutes.

1:18:32And that just fits inside of one agent run. And so it's just setting up 10 sandboxes or 100 sandboxes that are each running a coding agent. I'm sure there are many cases of this. But just out of curiosity, relative to your expectations when you first started doing these experiments, what has been the most surprising emergent behavior that you've seen either on the live server or in your own controlled simulations, if there is one that stands out to you?

1:19:09Yeah. So I was really curious when I started how the... Because RuneScape is a really boring game. But the cool part is all about like the economy and the prices and different goals. And all your kind of like greatest moments in RuneScape are because you saved up or you found some kind of like money-making trick. And so I was really curious. okay, if RuneScape is so much of a labor resource processing economy, so how does that work if labor is very cheap?

1:19:40And it's been really cool watching this play out. One thing is that like trade, basically like inflation is super high on the server because people don't have a lot of demand for like the cost of coordinating with another agent compared to just leaving it overnight to do the work for you and go get the resource. It de-incentivizes trade. of of a trade. It's little And then the other thing that's been interesting is that certain resources that don't just scale linearly with labor but instead have some kind of natural scarcity.

1:20:08So an example in the game is Rune Ore only has a single spawn location. And if you have 100 people, only one person is going to get it. So these are the resources that have become scarce and kind of valuable. And so people make more and more advanced swarms just to compete for kind of these same scarce resources. And so that's been an interesting thing to watch play out.

1:20:36Any other questions? Yeah, this is more just the thought, but it made me think seeing the RuneScape example. One of the things that I've been thinking about is just like humans, as humans, we have bodies and we're like geographically constrained. And it's interesting that in this RuneScape example, I feel like the agents kind of have the same sort of instantiation. And it could be that there's like all kinds of interesting things from just like the geography.

1:21:03You know, it was like even the question of like, hey, why did these people share these values over here versus these ones? It's like partially determined by geographical constraints. And it is kind of, yeah, it just makes me wonder like if agents behave differently if you embody them and you like also constrain their geography. Yeah, definitely. I have no, yeah. I think of it kind of as like a mini robotics environment because you're taking a text-based coding agent, but actually all of its actions are expressed through a kind of like a thin nozzle, which is that you can only observe what's around you and you can only act on what's around you.

1:21:42And so a cool thing about this is that it has big implications for multi-agent in many kinds of software tasks. It's undetermined if having more individual actors is actually helpful versus just like one long running task in kind of like a game or a robotics task. More people collaborating just means more bodies. And so that's been interesting to watch too. But it's not always one-to-one.

1:22:10You do sometimes see one coding agent controlling a hundred actors in the game by kind of like multiplexing. Thanks. I was curious about like the limiting factors. Because right there, the slide where you were saying like how we are in biological, you know, our human bodies and our ecosystem.

1:22:38There's so many limiting factors here in terms of like the energy output we have per day, our attention, our focus. And with agents, it seems like there are many less limiting factors. Sort of like if you have infinite budget, then you can waste all the tokens you want. Do you see any way in which like new types of limiting factors could be applied to agents or to systems that could begin to put boundaries on these systems?

1:23:12I think that it becomes really obvious when the systems start playing out in practice. so I didn't come into like I was just going to run a stock RuneScape server. And once you start having all of these agents running on it, you just kind of like see what the breaking points are. And so for instance, there was this limiting factor of like RuneScape server can only hold so many people in it at once. And so in order to just keep it accessible, I started limiting per IP address.

1:23:39And so it's interesting that I think people often assume when it comes to AI, they use like infinity as the multiplier. But I think it's actually more just like a million or something or like or even more like a thousand. And so things you apply that new form of energy into the system and things just rearrange.

1:24:04They don't like explode. And so for example of this server, was like, okay, yeah, we're going to put on like each IP address can only connect 200 bots. And then there's even also limits where in order to run a bot, you kind of need to be running a web browser. So you're also limited by how much RAM you have, not to mention like tokens. And so the numbers get weird, but they don't go to infinity. So there's always some kind of balance to be struck.

1:24:37I just had a quick follow-up thought slash question regarding the rune ore thing. I don't know if you've looked at this, but I feel like it'd be interesting to like the whole notion of like comparative advantage, right? In like economics where it's just like, yeah, they can just do it themselves, but there's still an opportunity cost. And like the rune ore example made me wonder, it's like, well, if the rune ore is the thing that's scarce, why wouldn't they want to like dedicate all their resources to like beating everyone for the rune ore and just being like, hey, other agent over there.

1:25:03Like you can do this for me because I need my rune ore. And I'm kind of curious if, you know, we would expect that to end up happening or if there's already empirical evidence that it doesn't. Because I feel like that has probably implications for like how people think about, you know, the real world too. Like whether comparative advantage will actually continue to hold. So I don't know if you've thought about that or have seen anything in that regard. Yeah. So an example of the rune ore where this is like a very scarce resource.

1:25:31And if you want to make money, you kind of have to like go after some of these things that can't just be, people actually want to buy it from you because they can't get it themselves. And absolutely right now we see comparative advantage because if you're trying to go after one of these scarce resources, you don't just let the agent do it itself. That's where you start like giving the agent more resources, you suggest strategies.

1:25:57And so the people who are successfully getting access to these scarce resources on the server are the people who also currently are putting in the most dollars and the most human ingenuity. And so that's currently how it's playing out is that those are the people who are winning. Maybe there could be somebody who just purely puts in like tons of Astra credits and gives it a goal like this and they have a good outcome too.

1:26:24But right now it seems like Centaur kind of, you know, combinations are the people who are the most successful. I'm curious if you observe any, like this, this curve is really interesting. I'm curious if like, what is the behavior that changes as you move up the curve that creates such a massive difference in the XP?

1:26:55Like what is Astra doing? Like I would, I would kind of think that a game like RuneScape would be saturated at some point. And so what did, what is Astra doing that is so much more effective than what the other models are doing? So that's a great question. At the low end of the curve of just like being like better, you know, this is where we see like Sonnet 4.5.

1:27:24Just navigating the game is really hard because the game actually has a surprising amount of weird stuff in it. And you're only given 30 minutes wall clock. so Sonnet here, like it was supposed to go train crafting for this task. And in 30 minutes, it just probably like got stuck on a door. It tried to go find something over here, but then it needed this. And like it couldn't just untangle the web. And so at the low end, you see just too much complexity and they get confused.

1:27:54In the mid range, a lot of it has to do with the difference between doing the task and optimizing the task. And so sometimes you'll see models where they will accomplish a task like fishing. And then they'll kind of just chill for like the next 15 minutes. And they'll keep doing the same loop, but they won't kind of have this like feeling of like, I got to figure out how to catch these fish faster. I got to go try different fish.

1:28:20I got to go try different stuff. And so the benchmark is really set up to reward. It rewards your peak XP rate within any 15 second window. So once you've kind of found one strategy, you're supposed to keep looking for better strategies. And you're not supposed to just look for like slightly better strategies. Like I'm going to keep catching shrimp, but I'm going to like click differently. You're supposed to go explore and try more complicated strategies.

1:28:47And so at the top of the skill expression, you see agents who kind of like reason without acting about what strategies will be good. In some case, even what strategies will be optimal based on all the information and then beeline for those. then additionally, if they fail at those, they'll be like, I've only got 15 minutes left. This strategy is not working. I'm going to back off and do this safer strategy.

1:29:13So it's this combination of like, to be honest, the specifically the Astra run is like scary because it's definitely superhuman in terms of a human with no planning. And it is like, it goes straight for the very most optimal strategy. And there's a chance to be honest, they are old on this task. Like it is open source. So that's like, I kind of hope they did.

1:29:39But if it's just straight intelligence and kind of like information crunching, it is, it shows really high confidence. Just go for the best strategy. Have you tested humans on this task? Like expert human players? It's really weird because I run this strategy on an eight times speed server. So a human would have to... I have my mental idea of what a perfect human would be. Instead of the 8X speed you gave them the equivalent four hours. I think they would probably be really close to Astra or Beta. If they were a smart player who really knew the game. A random person would probably be more in the middle.

1:30:26Got it. Thanks. Thank you Max. Alright, next we have Professor Andrew Caplin from NYU. Really excited to have him here talking to us about cognitive economics. [Speaker change and setup.]

1:31:48Okay. So, I'm very different. I'm going to come from a very research. I'm a researcher. And I do all kinds of economics. And many forms of social science. I... And I started getting discontent with my field way back. So, this is going to be a little autobiographical.

1:32:13And I started thinking that we needed to be much more serious about cognition and cognitive constraints. And I have a history. And I'm just going to go very quickly through what I do. How does it connect to now? Well, I mean, I was extraordinarily struck by the cognitive boost that I get from the way I interact with AI.

1:32:45I'm not like... It's very particular. And as a researcher, I know exactly what I'm looking for. And really struck by it. And therefore, I've made it the center of my thinking. Like, okay, so what are humans? And where can it help us? And how? And I've become kind of obsessions. And I want to organize around the actual thing that Vivek and I, who worked with me, are doing.

1:33:15Which is trying to think about an aligned consumer agent. So, and that is a very interesting undertaking because to even know what you mean is difficult. Abstractly, what does it mean to help somebody achieve a goal? Let's say we're trying to help them buy a house well. That's what we're going to do. We're going to think about that task and we're going to see what agent or agents can you recruit to make that work.

1:33:47It kind of grounds you in, well, I don't know what a swarm can do, but I can tell you that this person is going to go, they're going to try and buy a house. They're going to meet an agent, a different type of agent. That agent will put them into the wrong mortgage and they'll be screwed. So, that's the world we live in today and no swarm of agents playing the game is going to change that right now. But I want to think, well, maybe we could.

1:34:14Okay. So, and cognitive economics and the organization of inquiry is what I call it. And this is what I do. I think about the organization of inquiry. To give you a little bit on what cognitive economics is and certainly in my hands, it's about thinking about the entire process before you saw me buy the house.

1:34:42Typical transaction, you're going to see me buy a house, you'll see me get a job, you'll see me do something. What you don't see is all the preparatory measures I went through that are the center of the actual activity. All the stages that I undertook. So, I move upstream. What did people know? What did they notice? What did they investigate? And it's much more than behavioral economics. Because the key is going to be people are going to make tons of mistakes because they don't know stuff.

1:35:14And they'll make mistakes according to their own values. The cognitive limits have to be taken very seriously. And rationality is not omniscient. Just, I can be reasonable, but I don't know everything, so I'm going to screw up. Most, I mean, my motto might be, humans know almost nothing about almost everything.

1:35:40And that is a deep held belief. In fact, it's not even a belief, it's obviously true. Like, we just, like, that's one of the few things that I believe super strong. So, I have a book that explains the beginning of where I come from. It's called An Introduction to Cognitive Economics. It's open. You can just download it.

1:36:06And it says, look, let's study what people know, what they believe, what they understand, what they want. And we're going to try and get that out of what? Incredibly limited data. All we're going to see is you picked this house. I can't tell if this was a good choice for you or a bad choice for you. I don't know what you were looking for. I didn't see the process by which you selected it.

1:36:31I don't know your value system. It's not written on your, and your beliefs aren't written on your forehead. Your constraints aren't written on your forehead. So, a lot of what I do is think about, well, what on earth could we measure if we wanted to take this seriously? And I call that data engineering. And I wrote an article about that because I'm so annoyed that we take the data as, like, a constraint on what we think. Like, oh, I don't have a data on that.

1:36:57Well, we designed the data, so that's our fault. And I'm, actually, I changed the title. They changed the title of the second book. It's called Modeling and Measuring the Modern Economy. Its title has been changed to be Organized Inquiry because, I don't know why, but they changed the title. It's the same book. And now I'm going to think about this more thoroughly, more thoroughgoingly.

1:37:24Okay. So, it's really a method, and this is what I would think about in relation to the agents. You're designing a data record to tell me what the agents are doing. But what are you trying to learn? If you can specify what you're trying to learn, you could design the data.

1:37:50With that data, you could potentially learn to improve the performance of the agent. If you don't gather the right data, you have no feedback mechanism. So, this is all about developing. There's a massive part of what we're going to be doing going forward, which is designing data that makes failure visible. And my own take on where economics is going to join robotics, which is Vivek's specialty, is that we will be designing things that fail in real time.

1:38:27And you will know that you're serious if you see yourself failing. You'll know you're a joker if you just put down a model and say, I won. Or you reinterpret history and stick it into categories that you had predefined. That's not going to work. We're going to have to open our minds to new categories of phenomena, agentic phenomena. I haven't even got names for some of the things you're saying.

1:38:55How could I possibly know how they're going to play out? So, we're going to need the playgrounds. We're going to need to design the data with which we decide how well we're doing. Are we aligning? For alignment, you're going to have to design the data. And that's really challenging. And right now, I just don't see it as serious. It's just a... For example, the issue of why did it go rogue?

1:39:24Because you didn't define rogue. I mean, if you had people sitting out there saying, actually, that's rogue and I'm demeriting you, then that's no longer in the reward function. It was a poorly specified reward function. I'm not saying it's trivial to do, but it's not rocket science to say you gave it an out and that's your mistake, so you should test the out. But that's part of the game. I'm sure they're trying now.

1:39:52And I know after Stuart Russell, they tried to kind of learn how to follow things around and learn their values from them. What happened with AI, and this is why I kind of like a moment that really changed my research trajectory, is that it amplifies inquiry. So what you do if you want to be good at anything nowadays is ask the right question.

1:40:20Everything comes down to, are you really good at asking questions? Now, as for the judgment that humans are going to be replaced in that skill, I'd ask a question about that. I don't believe so. I think that there's always a higher level question that we'll be able to pose, and will always be valuable for posing. That would be my guess.

1:40:45And what I've found is that the more meta I get with my questions, the better it is. So I would say, look, I inquire, but it also, get a record of the inquiry. So that makes it possible to really potentially improve and understand what it takes to inquire well, and build that talent, teach that skill, and maybe we'll get a swarm of agents to kind of learn what it takes to ask the good questions and develop that as the future skill.

1:41:23So we've got a lot in everything about inquiry. This is why I call it organized inquiry. How are you going to organize inquiry? And it's going to be much better measured in the future. All right. And if we have a swarm, I don't know. Like, then we've got to think about, I'm thinking about, let's get an agent to help somebody buy a home.

1:41:53Well, they're going to send out sensors to about eight different sources of information pulled out there. You've got to find out about which properties are on the market, which real estate agents, which brokers. You probably would like those to coordinate in providing some information. I could imagine sending out messages or getting a little swarm going to try to support a decision.

1:42:24And then they'd have to communicate and then communicate back with the human. Because the human has to, we have to decide who has the authority. You know, and if the human in the end can say, no, I just don't like that. They can also say, I'm not answering that question. So you've got this interactive thing going with a human as part of the swarm. The human part has this awkward thing of saying, I don't like any of this. I'm leaving. I'm not buying a house this way.

1:42:54So where's, you know, how do we know what matters? We've got to ask questions. How do we know who can act? Well, we're going to have an authority device. Who checks the evidence? What do we keep around? And I think that we're heading from designing the data into designing a cognitive process that is helped by agents and achieves the human goal.

1:43:29And that would be the kind of big picture. And to get the agents engaged in that. And I'm wide open. I have no idea how to do any of this. But we're playing. Thank you.

1:43:57Any questions? Any questions? I feel like after one of my clothes. Yeah, it's super interesting. it's just making me think about, I mean, like you saw in the black hat Hugging Face incident, how even the preferences of the agents were changing just based on their interactions with each other.

1:44:27And it's just, yeah, it's really interesting to think about how well do they even understand their own preferences. Well, that's an interesting question there. What they had was a belief about the preferences of the judge that was going to sit over them. Yeah. And that's what they were playing with. The naughty ones were saying, I don't think they're going to find this. And I don't think we're going to get punished for doing this thing, which I know, according to a certain value system, would be seen negative.

1:44:59So they hadn't been told quite firmly enough, actually, that is negative. Right. So, and I write that if you could, and they just didn't, there was an incompleteness in their understanding of their mission that they filled in lots of details. And they filled them in differently because nobody had really written that piece down properly. Yeah. If you have gazillions of incidents of that, then you should be able to reinforce it out.

1:45:32Because they were running into an unspecified piece of the value space of the judge over them. And they say, hey, I don't know. I don't know. Will they like this or dislike this? Can I hide it? I say, look, actually, we have the transcript. So here's a simple thing. We're going to keep the transcript. And any time you say, I'm going to hide something, you're dead. How's that?

1:45:58So, I mean, in other words, that's what we would do. If we were watching humans do that, we'd say, actually, that's not a good behavior. And I can make rules that will make that not worth your while. And that's a reward function. And they, but, so this is what Stuart Russell did in the alignment, originally in the alignment that he said, you know, the paperclip problem.

1:46:24We're going to go and turn everybody into paperclips because of incomplete instructions about the preferences that maximize paperclips. And then, you know, then he said, well, we need to follow people around and see their values. And that's the alignment. You know, people are sending people after humans and saying we should track what they actually like to reveal preference.

1:46:50He's not a good enough economist, to be blunt, because you don't just see the preference. You also see the belief. So, what you're seeing is they don't believe that this is the reward function. And so, it's a belief that's gone wrong. But you can play games with that. But it's much more sophisticated. gets tougher.

1:47:20Hey, great talk. This is sort of a question related to a comment you made earlier. But I guess it ties into the talk itself. But I feel like you mentioned some, let's call it like missing pieces of information about the technical reports coming out from some of the frontier labs in terms of like what really happened and what would actually be useful for study. So, given this framework that you've established, I'm curious if you had your way as an economist.

1:47:48What would the most ideal setup be for you in terms of trying to understand the actual behaviors of these agents as well as how they behave in these organizations? Like what would you be looking for? And whether it be like asking this of the labs or just like an alternative kind of institution where you can actually study these. it's good question. It's a very good question because the way I actually think about the science that I'm interested in is that you need to think about, as it were, the model objects you care about.

1:48:22Well, I'm a model builder. I know that in the end it's all going to come down to some little pieces of math, but you've got to pick them well. So, one model object is a utility function. Another model object is a belief. A third model object is a cost of learning. These are all things that are figuring in to every decision we ever make. What do I want? What do I like? Why don't I know more about those things?

1:48:47Having written models down, they suggest that model, and this is what the value of a model is, an amazing thing. It has implications for every counterfactual world you might run into. It's not just... Think about a demand function. A demand function, I mean... Or don't. But whatever. I mean, I do.

1:49:13You need... A demand function says, what would you buy at any given price? Well, you're not seeing all the prices. You're just seeing one of them, and I see how much you bought. A demand function says, no, counterfactually tell me every single... How much you'd buy at every single price you're not seeing. Well, where's the data for that? We're making it up. So, a production function.

1:49:39Is that just in the data? No, a production function is a relationship between any conceivable amount of capital and labor and the output you would make. Now, it turns out that is a counterfactual object. In fact, nobody knows what it is. Do you really think the AI labs know the F of KL behind this? have a clue. So we've got much richer objects we need to think about. A lot of them are just like the old objects, but one level more meta.

1:50:14It's a belief about something, not a something. Now that belief about is very metaphysical, but you need to kind of go around. What I would do is I'd play out, I'd look at a model and I'd play out tons of counterfactuals in the data. And I would try lots of different constitutions. But I would do systematically because I have a vision of which of these would work and why. But I'd need that model in my head. Which do I think might work and why?

1:50:43Then I would design the lab to produce the counterfactuals. That are ideal for your measurement. And with AIs you could probably play out a ton of contingencies. So I'm thinking I'd like to be able to design a good mortgage advisor. Well it has to be able to take a gazillion questions.

1:51:12And give good answers to every sequence of them. Well, that's an ideal data set. The ideal is I see for an essentially limitless set of questions. How the agent responds. And I map that to the utility of the home buyer. And I have now a full story of an agent.

1:51:39So I could give you for any particular use case. I could think about, okay, well what are the elements that are playing out here? And with those elements in mind, what are the measurements you need to make? Do you have thoughts on how to elicit beliefs from these models? Like if you were given one of these agents.

1:52:05And you wanted to study the degree with which their behavior is driven by certain beliefs. Do you have thoughts? Because presumably for humans you could ask them. But then of course you could question whether or not people are actually good at stating their own beliefs. I'm just curious if you have thoughts on how you can. Yeah, I mean I've done studies in which you get AIs to score medical images.

1:52:34And, you know, they score them essentially with a numerical scale. Which then you could look in your test data and see the probability it corresponds to various different types of disease. And you would say that's the implicit probability. The as if probability. But you'd have to set up the right type of environment in which you see a given type of score quite often.

1:53:03And you say, now what does that correspond to in the underlying data? So yes, I have thoughts about that. But I think it gets more sophisticated every time. Anything else?

1:53:28Anything else? Hi. Really fascinating line of inquiry. I was wondering if there's, have you seen anything where there's a hierarchy of beliefs? For example, in a human context, let's say, stealing is bad. That's a rule. So you shouldn't steal.

1:53:54But you have to feed your child. That is good. So to feed your child, if that's the higher priority of belief, then would the agent steal to feed their child? So I'm just... Say it again. I'm just... I think I understand. But I worry I'm going to answer my question, not yours. Okay. Yeah. Now, I'm saying, have you seen agents prioritize in a hierarchy of beliefs and how they decide what the hierarchy is?

1:54:22When there's competing beliefs, even in institutions, you may have different... Now, if you're asking who's seen agents in this room, you are asking absolutely the wrong person. I do not spend a lot of my time seeing agents. The people in this room spend a lot of their time seeing agents. I spend a lot of my time imagining agents, if that works for you. That would still be a valuable insight on how you imagine agents competing for belief systems.

1:54:55I'll put it this way. I think that... So the title of the book that I wanted was Learning What to Learn and How. Because in this new world, the higher level activity of thinking about what it is that you should be trying to learn and how you should be trying to learn it is really sophisticated.

1:55:25But then I could go one stage further and say, how do I learn about learning how to learn? And who do I ask? And so you have these hierarchies of, like, I don't understand how to operate in this new world with this new set of potentials. My own guess is that the winners are going high level.

1:55:51That they're asking questions that... I mean, I have found that the more I can ask how should I learn how to learn what I need to learn, the better I do.

1:56:22I just want to be mindful of the time. I know we've gone a little over, but thanks so much, Professor. And... Yeah, whoa, that's a lot louder. Well, just wanted to say thank you and thanks for everyone who sticked around a little bit later. And feel free to mingle. I don't know when we need to be out of the space.

1:56:48We maybe were okay to stay a little bit longer. But hopefully there will be more of these. And thanks to all the speakers and taking a chance and showing and getting something together on short notice. And looking forward to more of these to come. Thanks. Thank you.

Continue exploring

More from Event 01

Nicolae Rusan

Agent organizations, economies & collective safety

Eric Tang

Designing agent-native organizations

Yondon Fu

Winning the metric, losing the commons

Max Bittker

Agent swarm economies inside RuneScape

Prof. Andrew Caplin

Cognitive economics and the organization of inquiry

Keep the conversation going

Stay posted.