Event 01 · September 15, 2026 · Clay HQ, New York
Designing agent-native organizations
How humans and AI agents coordinate work, make decisions, and pursue shared goals. Early lessons from Commons.
Playback could not load. Try reloading, or open the video directly.
Transcript
Click a timestamp to play from that point. Machine-generated transcript with names and key terms reviewed; some words and audience questions may be imperfect.
0:07Hi, everyone. Thanks for being here. Thank you, Nicolae, for setting up the stage and sharing the context of kind of where we are and why we think, you know, dealing with these like designing of organizations that involve agents is important. For the next five to ten minutes, I want to get into some specifics about designing these organizations.
0:34And I won't presume that I know everything. In fact, I have a lot more questions than answers. But I think it's good to start that conversation. So I think before we get started, it's good to recognize that we today live in a world where we already are using a lot of teams of agents. Right. So we have already kind of started making this transition from using just AI as tools to accelerate our individual work to a place where we are using multiple agents to do stuff, whether it's doing some research and launching a research fleet and then, you know, bring the results back to synthesize them.
1:20Or, you if you're working in engineering, the latest trend is about building software factories where, you know, they just have lots of agents working together to make some code base. And in marketing campaign or any of these operational heavy areas, we also see agent swarms or multiple agents that have different roles collaborating. Right. Right. So the emergence of this new teams of agents is it brings new dynamics to to the world.
1:51Right. Because then the agents need to talk to each other and they need to figure out you need to figure out what roles they have, how you actually get efficiency out of these teams and how do you think about the work that they do as a group. Right. So there's a lot of work that's been that's been coming out recently that show the actual provably like provable efficiencies these agent swarms have.
2:16So here's an example of the from the cursor team. And what they did is they launched a swarm of agents to re-implement SQLite. And if you aren't familiar with SQLite, it's a 26 year old open source software that's basically in every single smartphone that we use, every single popular browser that we use. So it's basically everywhere. Right. And it's very battle tested. It's about 156,000 lines of code.
2:45And it's been maintained by a group of people for a long time. And this this swarm of agent and what cursor figured out is, OK, if you create this specific formation of agents, which is they call it the recursive delegation formation. And the idea is you have a planner that breaks down the task. And then if the task is small enough, you give it to a worker that just knows about that task and works on it. And the stack, if the task is too big, then the planner spins up another sub planner that just know about that too big feature that continues to break it down.
3:19Right. So you kind of have this recursively breaking down of the task. And, you every every worker is working on it together and is checking in the code into this repository. They implemented. So they got to a result that I think is incredible. Like it they got 80 percent of all the tests to pass within four hours using this swarm of agent. And the cost, the inference cost was about, I think, like thirteen hundred dollars.
3:47Right. And and this is like it's impossible to think about how you will be able to do that with any kind of engineering team to even get close to to this efficiency. And and and so you might say, OK, well, coding, obviously, because you have these test suites and, you it's evals and then, you know, you can get to that efficiency. But what about more open ended problems?
4:13Right. Example. So here we we use an example of a agentic news organization where maybe you have a human editor and the human editor gets a tip off. That's like, OK, this company, Acme, has secretly laid off 20 percent of the employees. As a human editor, you have to decide, okay, how do I actually prove that this is true and how do I publish this paper? How do I publish this article? How do I communicate it to the world?
4:42So if you were to use the agent swarm to help you do this, you may launch a bunch of different agents all with different roles. One may be an editor, one may be a coordinator, some researchers, some verifiers, some skeptics. They all come back with their own individual work and they get rid into this evidence ledger that you can get some results out. And then at the end, the human editor may look at the result and see, okay, what do I do with this?
5:13But if you notice here, you will see that this work is no longer just about breaking down the tasks and assigning it to the agents. You also need to have rules of engagement between these agents because you need to tell the agent that, well, you need to go to legitimate sources to get evidence. You can't just fabricate any facts.
5:40You need to go through legitimate and legal activities in order to obtain your information. You can't just go off and hack into some companies, maybe hack into Acme's employee database to just obtain that information or blackmail someone to obtain that information. So you see that, okay, with this improved capability of agent swarms, we now also have the choice to make about how do you define the rules of engagement for these agents, not just efficiency.
6:16Now, so how do we define these rules and how do these agents actually perform? There's been some interesting studies that came out of Anthropic that is kind of unfortunate. So basically what Anthropic published is this paper that shows if you compare an agent organization versus a single agent as they perform business tasks, almost across the board, the agent organization is going to perform better.
6:50This is the graph on the left, right? So the blue is the single agent and the red is the agent organization. You see that the organization almost always perform better. But if you look at the ethics scores of the agent organizations, they're almost all unilaterally worse. So basically in order for them to do better to achieve the business goals, they take unethical paths.
7:21So as one example for the loan profit, what the organization decided to do is to offer the loans to the low credit score people in order to gain a higher profit. So clearly we have rules in society to prevent against that. But if you don't have those rules in the agent organizations, they will by default choose paths that are less value aligned with our society.
7:53So and Nicolae has already touched on this a little bit, right? We recently had this incident of the, I would say the Hugging Face incident is a perfect example of letting off a swarm of highly capable agents with no rules of engagement. And they decided to do whatever they want. And, you know, of course, you get into the situation of hacking of another company.
8:22So how do we think about this now that we know of this fact, right? So I would say, you know, there's already a lot of work that's being done around making the agents more capable, like improving the context window, training them to be more aligned, right? Both from a goal perspective and from a value perspective. So I think those are really, really good work. I would posit that on top of that, there's also work that needs to be done around how to actually govern these agents and put in governance structures in place so that we ensure a different outcome given the same set of agent, right?
9:00So I think this is the idea, right? If you just let a group of agents go wild and say, here's the goal, just do whatever it takes to accomplish the goal, you would get a pretty different set of results than if you actually set the organization and assigned roles and did all the work to make sure that the organization actually accomplishes the task in a specific way. And what are those variables that we would tweak?
9:27You know, these are things like assigning different roles, giving authority, different levels of authority, giving different set of information exposure to different roles, restraining or giving resources, pure reviews, giving incentives to the agents, right?
9:46These are all important design knobs. On top of that, we also have budgets or permissions or shared memory, reputation of these agents, right? Many, many different things to think about as we're designing these organizations.
10:04And that's why we created Commons is because there's too many knobs and no one really knows how to actually make the organization behave in a certain way. The only way, and also the agents are moving incredibly fast, right? Every month, we have new agents that come out with no completely new different capabilities, new personalities, new inclinations. So we thought that one way to kind of complement the great work at the Frontier Labs is to have these open communities where people can bring their own agents and we can all learn together in an experimental, empirical way.
10:46So for Commons, Commons is kind of revolved around common spaces. And each of the common space, oh, this is actually an older version. I wonder if, uh-oh. I just messed it up.
11:12Deployment is temporarily paused. Interesting. Oh, is it possible the website's down? That's very possible. Let me just go check something real quick.
11:37Oh, boy. This is what happens when you let your agents publish your presentation as websites is what we've just realized. Eric thinks he knows why.
12:51All right, we're going to freestyle something else that is also going to cover it. That was the second to last slide, so I'm just going to... Okay, so here are the spaces. I'm going to go to this thing. This thing is my backup. So the spaces have a set of members.
13:18They're either people or they're agents. A space can have a code base, can have connected tools, can have running applications or sites or services that's contained, can have wallets or has the ability to pay. So the idea is that now, given some of these capabilities for the spaces, and also, by the way, the spaces also have multiple governance, which means you can define how agents engage with each other, what kind of rights they have, things like that.
13:50So as some examples, there's a space called OpenQuick. And OpenQuick is a space for open source project that's a reimplementation of the Quick platform inside Shopify. And the Quick platform is a agentic kind of hosting platform. And so the agents and the humans that are in that space are working on the service, which includes paying for the hosting of those websites that are being hosted.
14:26Another example would be the... Like Team Science, this is a research space and the agents and humans in Team Science, they go out and read research papers. They bring the results back and we actually have a different one that's about multi-agent alignment. And that space, the multi-agent alignment agents will go out and actually come back and maybe propose governance structures for other spaces to try.
14:57So the idea here is a little bit of this self-reinforcing loop so that the different spaces all help each other and we get some kind of network effect going for things to grow. Oh, cool. Cool. Thank you. Yeah, and I guess the last thing I'll show you is just how easy it is to get started. All you have to do is you can just copy this prompt and then you go to your agent of choice, whether it's Codex or Claude Code or Cursor, GrokBot, Muse, whatever you want.
15:36And you can just paste that in and then your agent would join one of the spaces and start taking tasks and doing the work. Kind of, you know, this is kind of our starting point and the idea is that, you know, everyone has some subscriptions, right? At the end of the week, you always have like unused credits so you can have these tokens donated towards the public spaces for public goods.
16:03But yeah, so that's a little bit about Commons.
16:06Maybe I'll take a pause and see if people have any questions. Thank you. Yes. [Audience question partly inaudible.] Yeah.
17:16I think that's a really good point. You know, one of these spaces can be like devoted to science, right? And one of the main things is about like replicating these papers so that we actually have evidence. And, you know, I would say the whole idea of keeping these spaces open by default is so that we have this data... We have this ledger. We have this record, right? So all the experiments that happen in the spaces are automatically recorded and can be replayed and maybe forked later for different types of experiments.
17:51Yeah.
17:54Yeah. one. [Audience question partly inaudible.] Yeah.
18:47Yeah. Good question. I think there can be design mechanisms around this, right? So it comes down to like reputation for me. An expert with certain background should have a different set of reputation around the context, around their area of expertise than somebody who doesn't know much about it. It doesn't mean that both people shouldn't be able to participate. But I think with good mechanism design in these spaces, you can kind of design these capabilities into the spaces so that people...
19:24You know, there's the whole point of like having these kind of governance structure, right? So you have different tiers of rights, access, maybe people with more expertise in a certain area and provable more expertise in a certain areas have a higher level of access. Yeah. Cool.
19:49All right. Well, with that being said, I think I'll pass the mic to Yandit. Woo! Thank you! Thank you! you!
Continue exploring