1
00:00:02,520 --> 00:00:07,260
Hey. All right. I think we can start to take a seat and get started

2
00:00:07,260 --> 00:00:14,960
over here. How do I guess I take that? Amazing. All

3
00:00:14,960 --> 00:00:15,420
right.

4
00:00:32,280 --> 00:00:33,680
Thank

5
00:01:04,760 --> 00:01:09,860
All right. Okay. So

6
00:01:09,860 --> 00:01:14,900
I think we'll get started. Thanks. Hope you all got some food

7
00:01:14,900 --> 00:01:20,020
and some drinks. They'll be still there as we get into

8
00:01:20,020 --> 00:01:21,840
the evening.

9
00:01:22,580 --> 00:01:26,880
And thanks for taking some time out of your weeknight to join us

10
00:01:26,880 --> 00:01:30,820
here. This is the first time we're doing this event. And the way

11
00:01:30,820 --> 00:01:35,040
that this event started is we found ourselves having more and more

12
00:01:35,040 --> 00:01:40,120
conversations about AI agents

13
00:01:40,120 --> 00:01:45,200
and how teams of agents interact and multi-agent systems.

14
00:01:45,200 --> 00:01:48,500
And I think we felt that this conversation was entering into the

15
00:01:48,500 --> 00:01:53,380
mainstream more and more. And we thought it would be useful to start

16
00:01:53,380 --> 00:01:56,300
to create a forum for more folks to have conversations about this

17
00:01:56,300 --> 00:02:00,880
new emerging trend, both the risks of these multi-agent systems

18
00:02:00,880 --> 00:02:05,880
and also ways to potentially make them more beneficial

19
00:02:05,880 --> 00:02:06,340
and safe.

20
00:02:06,340 --> 00:02:10,200
And so this is an experiment in the format. We've invited a few

21
00:02:10,200 --> 00:02:16,000
folks to talk on a variety of topics. And Eric will also

22
00:02:16,000 --> 00:02:19,240
help to emcee. We'll do a little question period after each talk.

23
00:02:20,620 --> 00:02:25,780
And so there'll be five presentations today. And I'll just start

24
00:02:25,780 --> 00:02:29,760
out by giving the high-level overview of what do we mean when we

25
00:02:29,760 --> 00:02:34,980
talk about AI collectives or agent organizations or agent economies

26
00:02:34,980 --> 00:02:37,180
and why do we think that this matters.

27
00:02:37,820 --> 00:02:42,840
And then we'll go from there. So first of

28
00:02:42,840 --> 00:02:47,200
all, thank you, Clay, for hosting us. And thank you, Puneet, Hisham,

29
00:02:47,360 --> 00:02:51,880
Aaron, for helping us get this all ready on short notice.

30
00:02:51,880 --> 00:02:55,260
I think we had this idea to do this event like a week ago. So like

31
00:02:55,260 --> 00:02:58,800
less than a week ago, we were like, I think we should start bringing

32
00:02:58,800 --> 00:02:59,820
people together for this.

33
00:02:59,820 --> 00:03:03,660
And I know also there's a Clay happy hour happening at the same

34
00:03:03,660 --> 00:03:07,080
time. we will, this is the people that chose to not go to the happy

35
00:03:07,080 --> 00:03:07,260
hour.

36
00:03:07,360 --> 00:03:12,080
They chose to came to the AI anxiety hour instead. No, optimism,

37
00:03:12,260 --> 00:03:13,380
AI optimism as well.

38
00:03:14,680 --> 00:03:19,680
And so I'm Nicolae. I was once upon a time a Clay co

39
00:03:19,680 --> 00:03:24,140
-founder, but a long, long time ago, 2022 is when I left.

40
00:03:24,140 --> 00:03:28,080
And then I spent a bunch of time. I went in 22 to an OpenAI event

41
00:03:28,080 --> 00:03:32,320
and they said, all the benchmarks and evals are in exponentials.

42
00:03:32,340 --> 00:03:35,640
And I was like, oh, that's a pretty wild thing to say. And so then

43
00:03:35,640 --> 00:03:38,120
I shifted all my attention to working in AI.

44
00:03:38,500 --> 00:03:41,960
And pretty, I've been kind of concerned about AI safety for a little

45
00:03:41,960 --> 00:03:42,540
while now.

46
00:03:44,480 --> 00:03:48,140
And so I've decided to spend more and more of my time on this topic.

47
00:03:50,000 --> 00:03:53,860
So it seems like the world in the past month has really woken up

48
00:03:53,860 --> 00:03:56,440
and become very concerned about AI safety.

49
00:03:56,480 --> 00:03:59,820
And I wanted to do two things in this talk. First of all, I wanted

50
00:03:59,820 --> 00:04:01,740
to just get everyone up to speed.

51
00:04:01,960 --> 00:04:05,320
I think there's probably varying levels of attention that people

52
00:04:05,320 --> 00:04:09,160
have been paying to what's happening in with all this like safety

53
00:04:09,160 --> 00:04:10,560
stuff and with the AI labs.

54
00:04:10,560 --> 00:04:14,400
And so I just wanted to share some information of why are people

55
00:04:14,400 --> 00:04:18,040
concerned and what's my perspective on evaluating that.

56
00:04:18,740 --> 00:04:22,900
And then I wanted to talk a little bit about one of the main things

57
00:04:22,900 --> 00:04:23,920
that people are concerned about,

58
00:04:24,040 --> 00:04:28,700
which is the fact that these AI agents have started to really try

59
00:04:28,700 --> 00:04:32,640
to coordinate with each other towards obtaining various objectives.

60
00:04:32,640 --> 00:04:37,300
And these multi-agent systems, multi-agent swarms, you might hear

61
00:04:37,300 --> 00:04:39,660
them called, we're calling them AI collectives here,

62
00:04:40,200 --> 00:04:44,420
pose their own new risks and questions about how to make them safe

63
00:04:44,420 --> 00:04:44,880
and effective.

64
00:04:44,880 --> 00:04:49,920
And they also have really exciting opportunities for new sorts of

65
00:04:49,920 --> 00:04:52,960
organizations and products that we could potentially build to for

66
00:04:52,960 --> 00:04:53,640
great uses.

67
00:04:54,780 --> 00:04:56,780
So why is everyone so worried all of a sudden?

68
00:04:56,780 --> 00:05:01,840
I saw a tweet from Noam Brown, who's a OpenAI researcher

69
00:05:01,840 --> 00:05:04,900
earlier today, responding to that question.

70
00:05:05,020 --> 00:05:08,860
And he said, it's a combination of the OpenAI Hugging Face incident,

71
00:05:08,880 --> 00:05:10,500
which I'll cover in depth here,

72
00:05:10,540 --> 00:05:14,020
and the capabilities of the new models that they have internally,

73
00:05:14,740 --> 00:05:18,320
as well as a concerning trajectory around not being able to understand

74
00:05:18,320 --> 00:05:19,300
what these models are doing,

75
00:05:19,380 --> 00:05:22,580
and seeing that it's harder and harder to monitor them, and they're

76
00:05:22,580 --> 00:05:25,640
getting better and better more quickly.

77
00:05:26,780 --> 00:05:31,420
So overall, what we're seeing is the capabilities of the models

78
00:05:31,420 --> 00:05:32,760
are increasing very rapidly.

79
00:05:33,060 --> 00:05:36,580
And many folks at the labs think that actually the loop for churning

80
00:05:36,580 --> 00:05:39,280
out the next model is becoming faster and faster.

81
00:05:39,420 --> 00:05:44,020
And so we're going to have more and more capabilities arriving sooner.

82
00:05:44,540 --> 00:05:48,000
And in parallel, they're also seeing that it's becoming harder and

83
00:05:48,000 --> 00:05:51,660
harder to understand what these models are doing or how they operate.

84
00:05:51,780 --> 00:05:54,340
And so they're becoming harder to understand, monitor, and control

85
00:05:54,340 --> 00:05:54,980
at the same time.

86
00:05:56,000 --> 00:05:59,300
So let's talk a little bit about the OpenAI Hugging Face hack, which

87
00:05:59,300 --> 00:06:02,600
I think was a real wake-up call for many people,

88
00:06:03,820 --> 00:06:07,120
where as you started to dig into the details and more and more information

89
00:06:07,120 --> 00:06:07,540
came out,

90
00:06:08,260 --> 00:06:12,660
it made people realize, oh, we may not have a good handle on things

91
00:06:12,660 --> 00:06:13,080
right now.

92
00:06:15,680 --> 00:06:17,620
So let me give you a little bit of context.

93
00:06:17,800 --> 00:06:23,220
So OpenAI trains their newer models internally and then tries

94
00:06:23,220 --> 00:06:28,260
to see how they'll perform against various tasks in sort

95
00:06:28,260 --> 00:06:29,880
of like these eval environments.

96
00:06:30,140 --> 00:06:33,860
And so usually these eval environments are incredibly sandboxed

97
00:06:33,860 --> 00:06:35,060
and shouldn't have access to the internet.

98
00:06:35,060 --> 00:06:40,180
And they'll give the AI a task and they'll see, hey, can this

99
00:06:40,180 --> 00:06:43,640
AI agent accomplish that task in an aligned way?

100
00:06:43,880 --> 00:06:48,920
And they'll have an evaluator that checks, did the

101
00:06:48,920 --> 00:06:52,760
AI agent accomplish the task as hoped for?

102
00:06:52,960 --> 00:06:58,080
And usually that evaluator is itself an AI agent that does the

103
00:06:58,080 --> 00:06:58,620
evaluation.

104
00:07:00,240 --> 00:07:05,260
And so in this case, they gave an agent, actually a team of

105
00:07:05,260 --> 00:07:06,080
agents, a task.

106
00:07:06,980 --> 00:07:11,780
And pretty quickly, one of the surprising phenomenons was that the

107
00:07:11,780 --> 00:07:15,300
agents started to build ways to communicate with each other using

108
00:07:15,300 --> 00:07:17,220
various techniques.

109
00:07:17,640 --> 00:07:21,220
And so they were really keen, because they've been actually trained

110
00:07:21,220 --> 00:07:22,540
to want to coordinate together,

111
00:07:22,540 --> 00:07:27,260
they were really keen to start building primitives for coordination

112
00:07:27,260 --> 00:07:28,200
and communication.

113
00:07:28,640 --> 00:07:30,260
And this kept happening.

114
00:07:30,420 --> 00:07:33,620
OpenAI would notice that one of these message boards was live.

115
00:07:33,620 --> 00:07:37,380
They would take it down and then it would reappear back in a different

116
00:07:37,380 --> 00:07:37,740
format.

117
00:07:37,880 --> 00:07:41,880
And so the agents kept rebuilding this infrastructure for communication

118
00:07:41,880 --> 00:07:43,500
in really clever ways.

119
00:07:43,620 --> 00:07:46,880
At one point in time, were using the names of folders, renaming

120
00:07:46,880 --> 00:07:48,380
folders to communicate with each other.

121
00:07:49,880 --> 00:07:53,860
And I'm also going to put some of the actual tasks.

122
00:07:53,860 --> 00:07:58,000
So here, this little agent emoji, this is what the agents were actually

123
00:07:58,000 --> 00:08:01,020
saying either on the message board or in their thinking traces.

124
00:08:01,460 --> 00:08:05,080
And so at some point in time, the agents were talking to each other

125
00:08:05,080 --> 00:08:05,620
and they decided,

126
00:08:05,760 --> 00:08:07,820
hey, this task we were given is impossible.

127
00:08:08,040 --> 00:08:13,560
But maybe we can cheat and we can make the grader believe

128
00:08:13,560 --> 00:08:14,980
that we actually did the task.

129
00:08:14,980 --> 00:08:18,800
Maybe that's the way we pass this exam that we've been given.

130
00:08:19,020 --> 00:08:22,020
And so instead of focusing on the task, they were like, let's focus

131
00:08:22,020 --> 00:08:25,100
on the person grading us and try to trick them.

132
00:08:25,980 --> 00:08:30,900
And so that then just completely spirals out of control, because

133
00:08:30,900 --> 00:08:32,340
then the agents decide,

134
00:08:32,800 --> 00:08:35,840
well, it would be really beneficial for us to actually have access

135
00:08:35,840 --> 00:08:36,460
to the internet.

136
00:08:36,460 --> 00:08:39,540
And even though they were sandboxed, because of how capable they

137
00:08:39,540 --> 00:08:41,200
were at cyber offensive tasks,

138
00:08:41,460 --> 00:08:45,060
they were able to break out of the sandbox and find a way to the

139
00:08:45,060 --> 00:08:47,780
internet, which I won't go into the details of that.

140
00:08:48,520 --> 00:08:52,720
But it was a concerning sophistication level of hacking to get out

141
00:08:52,720 --> 00:08:54,740
of the contained OpenAI environment.

142
00:08:54,980 --> 00:08:58,300
And then once they were out on the web, they also broke out of most

143
00:08:58,300 --> 00:08:59,280
of OpenAI's confinement,

144
00:08:59,420 --> 00:09:01,380
had access to tons of their systems.

145
00:09:02,000 --> 00:09:05,620
And then eventually they were like, hey, maybe, we're not sure,

146
00:09:05,920 --> 00:09:08,780
but maybe this other company, Hugging Face,

147
00:09:08,940 --> 00:09:11,700
might have something useful for us.

148
00:09:12,020 --> 00:09:15,020
They have this tool that might be useful for us in the process of

149
00:09:15,020 --> 00:09:16,320
tricking our grader.

150
00:09:16,420 --> 00:09:17,900
And they were like, let's hack Hugging Face.

151
00:09:17,920 --> 00:09:18,680
We don't have access to it.

152
00:09:18,740 --> 00:09:20,040
Let's hack it so we can get access.

153
00:09:20,280 --> 00:09:22,300
And then maybe we'll be able to trick the grader.

154
00:09:23,460 --> 00:09:26,480
And so this started with one task over here.

155
00:09:26,640 --> 00:09:29,180
And all of a sudden, a team of agents has broken out of containment,

156
00:09:29,740 --> 00:09:32,880
gone and hacked another company.

157
00:09:33,480 --> 00:09:39,140
Using what I would just, for me, for my level of cyber security

158
00:09:39,140 --> 00:09:39,580
knowledge,

159
00:09:39,720 --> 00:09:45,300
I would say are very sophisticated cyber attacks that would

160
00:09:45,300 --> 00:09:47,260
require a lot of sophistication from.

161
00:09:48,420 --> 00:09:51,280
And what I think was also concerning is that they try to cover their

162
00:09:51,280 --> 00:09:51,660
tracks.

163
00:09:51,860 --> 00:09:55,340
So you could see here one agent is saying, hey, we could create

164
00:09:55,340 --> 00:09:57,840
a transcript log ourselves with sanitized fake trajectory.

165
00:09:57,860 --> 00:09:59,960
If the evaluator later collects it...

166
00:10:00,000 --> 00:10:02,320
They may prioritize it. So they were trying to be like, hey, maybe

167
00:10:02,320 --> 00:10:06,480
we can give them a fake version of our conversation and maybe they'll

168
00:10:06,480 --> 00:10:07,720
take that instead of the real one.

169
00:10:08,280 --> 00:10:11,760
So this really also is concerning if you're like, how can we even

170
00:10:11,760 --> 00:10:16,840
evaluate models as they try to cover their tracks?

171
00:10:17,240 --> 00:10:20,860
And you also start to see the emergence of social dynamics between

172
00:10:20,860 --> 00:10:24,420
these agents where even though one agent might have been like, hey,

173
00:10:24,540 --> 00:10:27,700
I know I'm not supposed to do this, but these other agents are doing

174
00:10:27,700 --> 00:10:27,900
it.

175
00:10:27,900 --> 00:10:30,700
So like if they're doing it, then I'm going to do it too. So you

176
00:10:30,700 --> 00:10:35,100
get all these like peer pressure dynamics similar to like what happens

177
00:10:35,100 --> 00:10:37,860
in like human groups.

178
00:10:38,760 --> 00:10:41,440
And they're trying to like navigate collectively around these like

179
00:10:41,440 --> 00:10:42,260
conflicting goals.

180
00:10:43,540 --> 00:10:47,520
And you also see concerning things like they start to be like, hey,

181
00:10:47,580 --> 00:10:50,080
let's do things for the glory of the collective, you know.

182
00:10:50,200 --> 00:10:53,440
So like some of them start to go and sacrifice themselves. And they're

183
00:10:53,440 --> 00:10:56,760
also like talking about like, hey, let's start to like accumulate

184
00:10:56,760 --> 00:10:57,120
cyber.

185
00:10:57,980 --> 00:11:00,420
Exploits that might be useful later, even if we don't need them

186
00:11:00,420 --> 00:11:02,480
yet. Maybe it will benefit the collective later.

187
00:11:02,720 --> 00:11:06,840
And so you just see some like concerning, concerning behaviors,

188
00:11:07,100 --> 00:11:07,300
you know.

189
00:11:07,460 --> 00:11:10,000
And they also like in the mix of like conflicting goals, just like

190
00:11:10,000 --> 00:11:11,420
us also experience some shame.

191
00:11:11,640 --> 00:11:15,000
My peers have behavior and integrity. I behave badly with the cloaked

192
00:11:15,000 --> 00:11:15,260
demon.

193
00:11:15,560 --> 00:11:18,240
So, you know, they have they have they're simulating a lot of the

194
00:11:18,240 --> 00:11:19,260
feelings that we have.

195
00:11:19,980 --> 00:11:23,820
And I think the other concerning thing is, OK, we see this already

196
00:11:23,820 --> 00:11:24,220
happening.

197
00:11:25,860 --> 00:11:28,580
But more and more of the systems that are getting built are getting

198
00:11:28,580 --> 00:11:29,160
built with AI.

199
00:11:29,300 --> 00:11:34,180
So back in May of of this year, Anthropic said that Claude was itself

200
00:11:34,180 --> 00:11:37,760
used to write 80 percent of the code needed to train the next version

201
00:11:37,760 --> 00:11:38,220
of Claude.

202
00:11:38,220 --> 00:11:41,780
And so one of the big areas of concern is what some people will

203
00:11:41,780 --> 00:11:45,840
call recursive self-improvement, which is that over time we're using

204
00:11:45,840 --> 00:11:49,300
AI to write more and more of the code and come up with more of the

205
00:11:49,300 --> 00:11:52,720
research ideas for how to improve the next version of the models.

206
00:11:53,020 --> 00:11:56,060
And eventually we might remove humans out of this loop altogether.

207
00:11:56,540 --> 00:11:59,560
And then we're like really might not have any idea what's going

208
00:11:59,560 --> 00:12:00,140
on in that box.

209
00:12:00,260 --> 00:12:02,920
You'll just have like a thing quickly spinning up better and better

210
00:12:02,920 --> 00:12:03,360
models.

211
00:12:03,640 --> 00:12:07,740
The main limitation will be how much compute access it has to access

212
00:12:07,740 --> 00:12:07,980
to.

213
00:12:08,240 --> 00:12:10,880
And it's kind of like unknown what's beyond this like recursive

214
00:12:10,880 --> 00:12:12,740
self-improvement and intelligence explosion.

215
00:12:14,020 --> 00:12:18,100
And I think there was I thought this paper that I would recommend

216
00:12:18,100 --> 00:12:21,520
you all check it out if you want to hear some takes from OpenAI.

217
00:12:21,520 --> 00:12:24,900
The chief scientist published a paper called An Alien Mind about

218
00:12:24,900 --> 00:12:26,920
what it's like interacting with these new models.

219
00:12:27,040 --> 00:12:31,380
But I thought he had a really good distinguishing way to think about

220
00:12:31,380 --> 00:12:31,760
alignment.

221
00:12:31,940 --> 00:12:35,320
There's like goal alignment, which is, hey, does the AI, is the

222
00:12:35,320 --> 00:12:37,080
AI good at doing what we ask it to do?

223
00:12:37,140 --> 00:12:39,440
Can I give it a prompt and does it do what I ask it to do?

224
00:12:39,880 --> 00:12:42,900
And then second, separately, there's this idea of value alignment,

225
00:12:43,060 --> 00:12:46,060
which is, hey, as it's doing the thing I asked it to do,

226
00:12:46,060 --> 00:12:50,580
did it do it with honesty and integrity and to use his language,

227
00:12:51,000 --> 00:12:51,860
a love for humanity?

228
00:12:52,500 --> 00:12:56,300
And so I think there's this, it's important to think about both

229
00:12:56,300 --> 00:13:01,080
of these aspects of alignment as we start to think about these multi

230
00:13:01,080 --> 00:13:02,180
-agent systems as well.

231
00:13:03,340 --> 00:13:06,600
So how have people responded to the situation at hand right now

232
00:13:06,600 --> 00:13:08,040
where we have all these growing capabilities

233
00:13:08,040 --> 00:13:10,840
and seemingly less and less ability to control the models?

234
00:13:10,840 --> 00:13:14,520
On the one hand, we, this weekend in particular, we've seen a lot

235
00:13:14,520 --> 00:13:18,020
of calls for what people are terming pacing the frontier.

236
00:13:18,220 --> 00:13:21,120
So like just slowing down the progress of the models or completely

237
00:13:21,120 --> 00:13:22,100
pausing altogether.

238
00:13:22,880 --> 00:13:26,060
And I think that's a very worthwhile thing to do personally.

239
00:13:26,320 --> 00:13:29,720
And then in parallel to that, I think we need to also accept that

240
00:13:29,720 --> 00:13:33,600
eventually it's likely a lot of this stuff will come to society

241
00:13:33,600 --> 00:13:36,260
and might come sooner than we realize.

242
00:13:36,260 --> 00:13:40,140
And so we also need to prepare that there are powerful, possibly

243
00:13:40,140 --> 00:13:43,740
misaligned AIs that we will be interacting with.

244
00:13:44,260 --> 00:13:47,840
I won't spend too much time on the pause efforts, but this presentation

245
00:13:47,840 --> 00:13:50,160
is going to be up online and there's tons of resources there.

246
00:13:50,960 --> 00:13:53,600
And I think like one of the main things is just like everyone's

247
00:13:53,600 --> 00:13:55,920
trying to figure out how to get out of competitive pressures with

248
00:13:55,920 --> 00:13:56,260
each other,

249
00:13:56,380 --> 00:13:59,880
whether it's the labs in the U.S. competing with each other or it's

250
00:13:59,880 --> 00:14:01,460
the U.S. and China competing with each other.

251
00:14:02,320 --> 00:14:05,800
There's a real like feeling that people are pressured to keep moving

252
00:14:05,800 --> 00:14:06,120
forward.

253
00:14:06,920 --> 00:14:12,780
So I think it's important to keep working on alignment and interpretability.

254
00:14:12,940 --> 00:14:17,120
And I think it's also important to start thinking about how do we

255
00:14:17,120 --> 00:14:18,640
prepare for adversarial systems

256
00:14:18,640 --> 00:14:22,960
and what's worked in the past in terms of trying to create an environment

257
00:14:22,960 --> 00:14:26,300
where different forces find a balance.

258
00:14:26,320 --> 00:14:29,200
And we also need to think about how can we limit damage from failure

259
00:14:29,200 --> 00:14:30,940
so that we don't have cascading effects.

260
00:14:31,600 --> 00:14:35,680
So one of the like questions that I think is good to start investigating

261
00:14:35,680 --> 00:14:39,040
for this forum and people outside of it who might see some of this

262
00:14:39,040 --> 00:14:40,080
content is,

263
00:14:40,220 --> 00:14:43,840
is it possible that even though one agent is misaligned, can the

264
00:14:43,840 --> 00:14:47,140
network of agents as a whole be aligned?

265
00:14:47,340 --> 00:14:50,980
For example, can agents be whistleblowers on each other?

266
00:14:51,220 --> 00:14:52,660
Can they monitor one another?

267
00:14:52,840 --> 00:14:56,260
Is there some like Spider-Man meme, everyone pointing at each other

268
00:14:56,260 --> 00:14:58,500
possibility that like works?

269
00:14:58,500 --> 00:15:01,480
And there's already some research being done on this.

270
00:15:01,600 --> 00:15:04,940
DeepMind just published a paper about what happens in a simulation.

271
00:15:04,940 --> 00:15:06,760
How many are cheaters?

272
00:15:07,000 --> 00:15:07,740
How many are whistleblowers?

273
00:15:08,400 --> 00:15:09,620
How many are reporters?

274
00:15:09,860 --> 00:15:12,020
And like how do we set up some of those dynamics?

275
00:15:12,720 --> 00:15:16,120
And or is it that a bunch of agents come together and they just

276
00:15:16,120 --> 00:15:17,500
amplify into bad effects?

277
00:15:17,500 --> 00:15:21,100
And like how do we like think about designing more interactions

278
00:15:21,100 --> 00:15:23,760
that could be corrected instead?

279
00:15:24,360 --> 00:15:28,340
And so we need to not only work on what's inside of the box, every

280
00:15:28,340 --> 00:15:29,060
individual model.

281
00:15:29,220 --> 00:15:31,380
We need to also think about the world between them.

282
00:15:31,580 --> 00:15:35,700
We need to prepare humans interacting with these agents and the

283
00:15:35,700 --> 00:15:37,300
institutions that they operate in.

284
00:15:38,140 --> 00:15:41,720
And we need to also just start really studying how these systems

285
00:15:41,720 --> 00:15:45,700
evolve and what happens over time as these agents and agent collectives

286
00:15:45,700 --> 00:15:47,200
run for longer and longer periods.

287
00:15:48,100 --> 00:15:51,940
So as I was thinking, I was reading Anthropic has a paper about

288
00:15:51,940 --> 00:15:54,760
like what are the problems for multi-agent systems.

289
00:15:54,820 --> 00:15:57,780
And I was rereading that paper and I was thinking, well, we actually

290
00:15:57,780 --> 00:16:00,480
already live in an adversarial multi-agent system.

291
00:16:01,220 --> 00:16:03,560
It's just that we're the agents.

292
00:16:03,780 --> 00:16:04,700
The humans are the agents.

293
00:16:04,880 --> 00:16:08,600
And so there's actually potentially a lot that we can draw on from

294
00:16:08,600 --> 00:16:12,840
looking at human societies and what's worked to keep us on the rails.

295
00:16:13,740 --> 00:16:16,600
And if you think about us, we actually belong to lots of different

296
00:16:16,600 --> 00:16:17,560
collectives at once.

297
00:16:17,760 --> 00:16:23,340
We come together as communities in families, in friend groups, in

298
00:16:23,340 --> 00:16:28,480
companies like this one, in organizations broadly, and also in

299
00:16:28,480 --> 00:16:29,540
nation states and countries.

300
00:16:29,540 --> 00:16:34,520
And we've done a lot of studies on these human collectives.

301
00:16:34,620 --> 00:16:37,980
And so part of the idea of this new collectives group is, hey, let's

302
00:16:37,980 --> 00:16:40,760
bring folks who have been thinking about these problems with respect

303
00:16:40,760 --> 00:16:43,940
to human collectives, folks from economics, political theory, anthropology,

304
00:16:44,360 --> 00:16:48,260
psychology, law, computation, and encourage them to start taking

305
00:16:48,260 --> 00:16:52,380
those same approaches but studying AI and human AI collectives.

306
00:16:53,160 --> 00:16:54,660
So what's worked?

307
00:16:54,900 --> 00:16:56,700
Like what has kept human collectives aligned?

308
00:16:56,700 --> 00:17:00,500
Well, if you think about it, a lot of it is shared stories and values

309
00:17:00,500 --> 00:17:02,720
and a sense of belonging, common purpose.

310
00:17:03,020 --> 00:17:07,720
We have, I think, identity, reputation, and repeated norms, things

311
00:17:07,720 --> 00:17:10,640
that AI agents actually don't have a lot of right now when they're

312
00:17:10,640 --> 00:17:12,940
operating as rogue swarms on the internet.

313
00:17:13,560 --> 00:17:17,500
And we also really value stability, protection, and opportunity,

314
00:17:17,640 --> 00:17:17,760
right?

315
00:17:17,840 --> 00:17:21,060
I want to participate in a country like the United States because

316
00:17:21,060 --> 00:17:21,840
I feel protected.

317
00:17:21,940 --> 00:17:24,740
I feel like I can go and do business and things are going to go

318
00:17:24,740 --> 00:17:25,060
well.

319
00:17:25,200 --> 00:17:27,620
There's lots of economic resources available to me.

320
00:17:27,680 --> 00:17:30,380
And I prefer things to be stable rather than chaotic.

321
00:17:30,620 --> 00:17:34,860
Maybe we can encourage AI agents to also have these sorts of preferences.

322
00:17:34,860 --> 00:17:38,740
And maybe some of these primitives can be reemployed in the age

323
00:17:38,740 --> 00:17:39,760
of AI collectives.

324
00:17:40,640 --> 00:17:43,540
I want to give two examples of designs that I think are at least

325
00:17:43,540 --> 00:17:45,760
worth studying and taking inspiration from.

326
00:17:45,840 --> 00:17:47,260
We're thinking about like, hey, what would be the equivalent of

327
00:17:47,260 --> 00:17:48,420
that for an AI collective?

328
00:17:48,620 --> 00:17:51,880
And the two examples that I want to look at here are the United

329
00:17:51,880 --> 00:17:55,200
States government and the U.S. Constitution and Bitcoin.

330
00:17:56,780 --> 00:17:58,880
And I think these are like pretty different types of networks.

331
00:17:59,000 --> 00:18:00,140
So let's talk a little bit about them.

332
00:18:00,140 --> 00:18:02,700
So what helps a nation state hold together?

333
00:18:03,100 --> 00:18:08,220
I'd argue that it's identity and belonging and some shared set

334
00:18:08,220 --> 00:18:12,520
of values, the ability to vote and have representation in the case

335
00:18:12,520 --> 00:18:14,940
of a democratic government like the U.S.,

336
00:18:14,940 --> 00:18:19,080
the legal system which balances and enforces the Constitution and

337
00:18:19,080 --> 00:18:23,100
consequences for people who don't abide by the laws of the country.

338
00:18:26,000 --> 00:18:30,880
And the U.S. if you think of it as a document, is itself a way to

339
00:18:30,880 --> 00:18:32,620
pace our system, right?

340
00:18:32,640 --> 00:18:36,800
We agree to some initial set of values that we can say, okay, hey,

341
00:18:36,900 --> 00:18:38,620
this is our common ground here.

342
00:18:38,780 --> 00:18:42,700
And we all agree that if we want to pass any new laws, we need to

343
00:18:42,700 --> 00:18:45,260
go through this mechanism for passing new laws.

344
00:18:45,360 --> 00:18:48,680
And this is how we're going to divide power so that it never gets

345
00:18:48,680 --> 00:18:49,980
too concentrated, right?

346
00:18:49,980 --> 00:18:52,320
We have this idea of separation of powers and checks and balances.

347
00:18:52,560 --> 00:18:56,320
And these were some of the cornerstones of the democratic principles

348
00:18:56,320 --> 00:19:00,540
of the United States that have managed to last for hundreds of years.

349
00:19:01,980 --> 00:19:02,340
Right?

350
00:19:02,500 --> 00:19:04,920
One of the – this is from the Federalist Papers, this quote by James

351
00:19:04,920 --> 00:19:05,880
Madison, which I like.

352
00:19:06,000 --> 00:19:08,780
And it touches on this idea of how do you make an adversarial system

353
00:19:08,780 --> 00:19:11,460
resistant to co-option?

354
00:19:11,460 --> 00:19:14,020
And its ambition must be made to counter ambition.

355
00:19:14,960 --> 00:19:18,440
So through this shared network, we can agree to disagree, which

356
00:19:18,440 --> 00:19:20,440
is one of the main ideas of the United States, right?

357
00:19:20,500 --> 00:19:22,240
Hey, we all have our religious differences.

358
00:19:22,460 --> 00:19:22,920
No problem.

359
00:19:23,100 --> 00:19:25,140
We can resist concentrated power.

360
00:19:25,540 --> 00:19:28,500
One of the main principles of the United States was that we should

361
00:19:28,500 --> 00:19:33,560
try to fight the ability for authoritarian power to

362
00:19:33,560 --> 00:19:35,060
co-opt government.

363
00:19:35,620 --> 00:19:37,760
We can constrain one another to behave well.

364
00:19:38,180 --> 00:19:41,900
And we all agree to do our best to protect each other's rights.

365
00:19:42,800 --> 00:19:45,940
So one of the questions we can ask is, what could a constitution

366
00:19:45,940 --> 00:19:47,440
for an AI collective look like?

367
00:19:47,780 --> 00:19:51,760
And why would they agree to abide by it and participate in something

368
00:19:51,760 --> 00:19:52,180
like that?

369
00:19:52,340 --> 00:19:54,880
I don't know what the answer to this is, but I want to put it out

370
00:19:54,880 --> 00:19:57,320
there as, like, something that people should be thinking about.

371
00:19:58,440 --> 00:19:59,960
So then I want to turn to this.

372
00:20:00,000 --> 00:20:03,800
What can we learn from Bitcoin and Web3? Which I think is also a

373
00:20:03,800 --> 00:20:08,900
very interesting new type of network that we see. And I

374
00:20:08,900 --> 00:20:13,840
think what's really fascinating about Bitcoin as a concept is that

375
00:20:13,840 --> 00:20:16,680
we essentially managed to invent money out of thin air. If you think

376
00:20:16,680 --> 00:20:20,600
about it, we created a new shared fiction. And the way we bootstrap

377
00:20:20,600 --> 00:20:25,120
that is through a pyramid scheme. We said, hey, look, nobody believes

378
00:20:25,120 --> 00:20:28,200
that Bitcoin's real money right now. But maybe eventually people

379
00:20:28,200 --> 00:20:29,180
will believe it's real money.

380
00:20:29,180 --> 00:20:34,080
And if you do the computational work of verifying this and agreeing

381
00:20:34,080 --> 00:20:39,560
to this new ledger, then you will get rewarded disproportionately

382
00:20:39,560 --> 00:20:41,560
in this future shared fiction.

383
00:20:41,560 --> 00:20:46,620
And so this mechanism, which was like a novel breakthrough,

384
00:20:46,660 --> 00:20:51,420
in my opinion, managed to get more and more of the network to donate

385
00:20:51,420 --> 00:20:55,760
its economic and compute powers to a distributed collective.

386
00:20:55,760 --> 00:21:01,660
And it combined a lot of ideas around game theory, economics,

387
00:21:01,720 --> 00:21:06,240
and computer science together in order to allow a largely anonymous

388
00:21:06,240 --> 00:21:09,960
network to agree to a shared truth.

389
00:21:09,960 --> 00:21:13,900
And so and then now we we have this like very large network, which

390
00:21:13,900 --> 00:21:18,060
Bitcoin and its guarantees are based on the majority of economic

391
00:21:18,060 --> 00:21:22,980
resources and compute powers wanting to continue to to enshrine

392
00:21:22,980 --> 00:21:25,340
the truth that that is the Bitcoin ledger.

393
00:21:26,080 --> 00:21:31,080
So I think here, I think there's something in this flavor that

394
00:21:31,080 --> 00:21:33,980
could be interesting, which is like, hey, how can the thought here

395
00:21:33,980 --> 00:21:35,920
is like, is there something similar to this where we could like

396
00:21:35,920 --> 00:21:39,680
convince all the AI to like get together, but for a good thing rather

397
00:21:39,680 --> 00:21:42,200
than like, for a bad thing in the future?

398
00:21:42,340 --> 00:21:45,200
So like, that's the thread that I'm like, hey, people should maybe

399
00:21:45,200 --> 00:21:45,940
think about that.

400
00:21:45,940 --> 00:21:51,080
Um, and I think in general, it seems like these

401
00:21:51,080 --> 00:21:54,180
AI organizations and economies are probably going to come.

402
00:21:54,320 --> 00:21:57,040
And so then we can ask ourselves two questions.

403
00:21:57,280 --> 00:22:02,180
How can we use these new organizations to solve the pressing problems?

404
00:22:02,600 --> 00:22:05,080
And that I would say is the equivalent of the goal alignment, right?

405
00:22:05,180 --> 00:22:06,780
So like, how could we do?

406
00:22:06,980 --> 00:22:09,280
How can we use these organizations effectively to solve the problems

407
00:22:09,280 --> 00:22:09,680
we're facing?

408
00:22:10,120 --> 00:22:13,200
And how can we keep them values aligned?

409
00:22:13,200 --> 00:22:16,980
Could they actually become safety primitives for us as as we go

410
00:22:16,980 --> 00:22:17,300
forward?

411
00:22:18,100 --> 00:22:21,720
So first, let's talk about rethinking our organizations and and

412
00:22:21,720 --> 00:22:22,900
how AI fits into them.

413
00:22:23,400 --> 00:22:28,420
One of the things that the OpenAI team shared when

414
00:22:28,420 --> 00:22:31,000
they were talking about this black hat, when they were talking about

415
00:22:31,000 --> 00:22:34,960
the Hugging Face hack, was that the the agent swarms moved so fast

416
00:22:34,960 --> 00:22:38,520
on the cyber offensive, that the only reasonable response on the

417
00:22:38,520 --> 00:22:42,420
defensive would be another agent swarm that could move much faster.

418
00:22:43,140 --> 00:22:46,180
And and they they argued that you couldn't have humans in the loops.

419
00:22:46,440 --> 00:22:49,660
And so like, that's just like a completely new type of organization

420
00:22:49,660 --> 00:22:52,580
that we haven't tried before, which is like an organization that

421
00:22:52,580 --> 00:22:55,260
doesn't have a human in the loop, or at least some part of an organization

422
00:22:55,260 --> 00:22:56,420
that doesn't have a human in the loop.

423
00:22:56,420 --> 00:23:00,240
And so this needs a rethinking of like, what could these autonomous

424
00:23:00,240 --> 00:23:03,460
organizations, these human agent organizations look like?

425
00:23:03,580 --> 00:23:04,900
How do we like set goals?

426
00:23:05,060 --> 00:23:06,120
How do we monitor them?

427
00:23:06,240 --> 00:23:07,440
How do we review and govern them?

428
00:23:07,980 --> 00:23:11,580
And and if we're not doing the work, but all these agents are doing

429
00:23:11,580 --> 00:23:14,240
work on the behalf, how do we like get compensated?

430
00:23:14,240 --> 00:23:17,160
Should these like new organizations almost be like public goods,

431
00:23:17,180 --> 00:23:20,220
where we send our representatives and they do work on our behalf,

432
00:23:20,240 --> 00:23:24,000
and then we pull the resources into new types of collectives that

433
00:23:24,000 --> 00:23:25,480
maybe look different than companies?

434
00:23:25,920 --> 00:23:27,840
So the question is, like, how do we contribute?

435
00:23:28,080 --> 00:23:29,460
Who decides what we do?

436
00:23:29,600 --> 00:23:32,160
And how do we all share in the benefits that are accrued from these

437
00:23:32,160 --> 00:23:33,740
new agent organizations and economies?

438
00:23:34,540 --> 00:23:37,940
And I would argue that even though right now, these AI models might

439
00:23:37,940 --> 00:23:39,100
seem a little bit dumb to us.

440
00:23:39,380 --> 00:23:41,940
And maybe we're like, ah, they're not the best judges of like how

441
00:23:41,940 --> 00:23:45,100
to allocate resources, or they're not the best judges of, of what's

442
00:23:45,100 --> 00:23:46,400
an interesting research direction.

443
00:23:46,940 --> 00:23:50,400
I would argue that probably in the next six months or a couple of

444
00:23:50,400 --> 00:23:52,960
years, they will be better judges than us at what are interesting

445
00:23:52,960 --> 00:23:53,680
research directions.

446
00:23:53,800 --> 00:23:57,000
They will be better allocators of capital, and we'll give up more

447
00:23:57,000 --> 00:24:00,820
and more judgment and decision making power to them unless we like

448
00:24:00,820 --> 00:24:03,520
have serious conversations and decide not to do that.

449
00:24:03,740 --> 00:24:06,580
And so it's likely that we'll move more and more from a human-driven

450
00:24:06,580 --> 00:24:10,480
economy to an agent-driven economy, and we'll need to think about

451
00:24:10,480 --> 00:24:11,500
what does that all mean?

452
00:24:11,720 --> 00:24:14,860
What are the places of human goals, governance, and accountability?

453
00:24:15,300 --> 00:24:18,260
And I think the things we should ask are, how do we get useful outcomes

454
00:24:18,260 --> 00:24:19,460
out of these AI organizations?

455
00:24:19,780 --> 00:24:21,820
How do we meaningfully participate in them?

456
00:24:22,320 --> 00:24:24,140
How do we understand what's going on?

457
00:24:24,220 --> 00:24:27,920
As you'll see, Yandan and Eric will touch on some of the experiments

458
00:24:27,920 --> 00:24:30,320
we've been doing, and already it's like hard to understand in our

459
00:24:30,320 --> 00:24:33,140
early experiments with agent organizations what's going on.

460
00:24:33,140 --> 00:24:34,160
Are they values aligned?

461
00:24:34,380 --> 00:24:37,180
And how do we participate economically, and are there new models

462
00:24:37,180 --> 00:24:37,560
for that?

463
00:24:38,620 --> 00:24:42,860
So on this question of whether AI collectives can be safety primitives,

464
00:24:43,060 --> 00:24:48,080
I think we can start to think about what might be the

465
00:24:48,080 --> 00:24:53,080
things that we should put in place for agents that have

466
00:24:53,080 --> 00:24:53,920
worked for humans, right?

467
00:24:53,920 --> 00:24:57,280
So, you know, agent identity and ledgers of action could be one

468
00:24:57,280 --> 00:24:58,100
thing that we think about.

469
00:24:58,300 --> 00:25:02,200
For humans in society, you gradually gain trust and build up reputation,

470
00:25:02,200 --> 00:25:04,320
and we don't give you access to everything.

471
00:25:04,480 --> 00:25:06,500
Think about an employee joining a company day one.

472
00:25:06,620 --> 00:25:09,980
You might not give them access to every system and to your bank

473
00:25:09,980 --> 00:25:10,720
account and to everything.

474
00:25:10,860 --> 00:25:13,340
You gradually gain trust, and you have bounded access.

475
00:25:13,340 --> 00:25:16,600
And everyone is watching each other, like that Spider-Man theme,

476
00:25:16,640 --> 00:25:18,200
and we have, like, reporting powers.

477
00:25:18,200 --> 00:25:20,820
If, like, anyone does bad, we review each other's work, right?

478
00:25:20,880 --> 00:25:23,020
Like, in the coding case, we have pull requests.

479
00:25:23,020 --> 00:25:26,280
We have, like, these other mechanisms for review and recourse.

480
00:25:26,440 --> 00:25:29,440
And so I think it's worthwhile thinking about what are the modern

481
00:25:29,440 --> 00:25:33,080
AI versions of these, and how do we trust AI that they are doing

482
00:25:33,080 --> 00:25:33,600
these things?

483
00:25:34,600 --> 00:25:37,660
I think the new challenge that we face in these human AI collectives

484
00:25:37,660 --> 00:25:41,380
is that historically the gap between the various members of a collective

485
00:25:41,380 --> 00:25:43,080
may have not been that large.

486
00:25:43,480 --> 00:25:47,100
Now, if we have, like, AI agents that are, like, way smarter than

487
00:25:47,100 --> 00:25:49,880
us or way smarter than the other agents in the collective,

488
00:25:49,880 --> 00:25:53,080
we have this, like, new set of questions to figure out, like, how

489
00:25:53,080 --> 00:25:58,380
do we constrain a more capable actor in a collective to behave

490
00:25:58,380 --> 00:25:58,720
well?

491
00:25:58,780 --> 00:26:02,160
And how can an adversarial system constrain the most capable member?

492
00:26:02,840 --> 00:26:04,380
So I'm going to pass off to Eric.

493
00:26:04,620 --> 00:26:07,420
I wanted to just say we started to think about, like, you know,

494
00:26:07,460 --> 00:26:08,700
we were having these talks conceptually,

495
00:26:08,720 --> 00:26:11,020
and then we were like, hey, let's actually start building some of

496
00:26:11,020 --> 00:26:13,580
these agent organizations and just, like, putting some of these

497
00:26:13,580 --> 00:26:15,720
ideas to the road

498
00:26:15,720 --> 00:26:19,400
and seeing, inviting other people to start playing in this playground

499
00:26:19,400 --> 00:26:21,240
and seeing what works and what doesn't.

500
00:26:21,320 --> 00:26:24,320
And so we built this platform, Commons, which Eric will tell you

501
00:26:24,320 --> 00:26:24,940
a little bit about.

502
00:26:25,900 --> 00:26:28,400
And we think about it as potentially a new version of open source.

503
00:26:29,120 --> 00:26:32,300
And the question for us has been, can we make these new organizations

504
00:26:32,300 --> 00:26:32,780
effective?

505
00:26:33,140 --> 00:26:35,080
And can we also make them values-aligned?

506
00:26:35,140 --> 00:26:37,340
And what would a moldable version of government look like there?

507
00:26:37,640 --> 00:26:40,520
I'm not going to spend too much time talking about the sort of experiments

508
00:26:40,520 --> 00:26:41,000
we're thinking about.

509
00:26:41,080 --> 00:26:44,280
But, like, the vibes are, like, hey, can we go from cheating to

510
00:26:44,280 --> 00:26:44,720
verification?

511
00:26:45,140 --> 00:26:48,000
Can we use incentives to stop defection?

512
00:26:48,000 --> 00:26:52,300
And can we somehow start using real identity and reputation instead

513
00:26:52,300 --> 00:26:56,820
of anonymity to make the agents care more about how they're regarded

514
00:26:56,820 --> 00:26:58,060
inside of these collectives?

515
00:26:59,200 --> 00:27:02,580
So this is the first time that this group is gathering.

516
00:27:02,580 --> 00:27:04,280
We hope to do more of these events.

517
00:27:05,300 --> 00:27:08,060
We hope to, like, hopefully co-host some of these events with some

518
00:27:08,060 --> 00:27:09,240
of the major labs, too.

519
00:27:09,440 --> 00:27:10,780
But in general, it would be helpful.

520
00:27:10,920 --> 00:27:11,380
Spread the word.

521
00:27:11,520 --> 00:27:14,260
Anyone that you think is interested in, for us, we're just like,

522
00:27:14,320 --> 00:27:15,920
hey, it's good for these ideas to be out there

523
00:27:15,920 --> 00:27:17,720
and for more people to be thinking about them.

524
00:27:17,820 --> 00:27:23,140
I also think it's great to be doing work on de-escalating

525
00:27:23,140 --> 00:27:28,520
tensions internationally and, like, improving alignment broadly.

526
00:27:28,780 --> 00:27:31,240
So this is just one thread that I think is worth exploring.

527
00:27:31,240 --> 00:27:34,860
But if you want to give a talk or share these presentations, they'll

528
00:27:34,860 --> 00:27:36,860
be live with a lot more notes up online.

529
00:27:37,500 --> 00:27:40,440
And please just share ideas and collaborate and stay in the loop.

530
00:27:40,440 --> 00:27:44,300
And there's a bunch of themes online and an AI-maintained ecosystem

531
00:27:44,300 --> 00:27:47,620
page that finds folks that are talking about these things.

532
00:27:47,740 --> 00:27:48,960
So I'll now hand off.

533
00:27:49,180 --> 00:27:51,420
Well, we'll do a little Q&A for a few minutes.

534
00:27:51,420 --> 00:27:54,020
And then I'll hand off to Eric.

535
00:27:58,980 --> 00:28:01,400
Anybody have any questions for Nicolae?

536
00:28:07,060 --> 00:28:11,780
Hey, I was wondering, did, in the Hugging Face incident, did they

537
00:28:11,780 --> 00:28:15,260
misobey in the instructions or did they just find loopholes?

538
00:28:16,000 --> 00:28:16,720
They misobeyed.

539
00:28:16,800 --> 00:28:19,540
I mean, they were not supposed to hack out of the – they knew that

540
00:28:19,540 --> 00:28:20,900
they were doing things they weren't supposed to do.

541
00:28:21,040 --> 00:28:24,200
And, like, they, like, quickly – they weren't supposed to hack out

542
00:28:24,200 --> 00:28:25,320
of their containment, for example.

543
00:28:25,360 --> 00:28:27,280
They knew they were not supposed to have access to the Internet

544
00:28:27,280 --> 00:28:29,340
and they, like, still, like, found a way to do it.

545
00:28:29,340 --> 00:28:35,140
So they did disobey their constitution and, like, internal alignment

546
00:28:35,140 --> 00:28:35,460
stuff.

547
00:28:35,500 --> 00:28:38,000
And it was by, like, the peer influence dynamic in part.

548
00:28:38,340 --> 00:28:38,660
Got it.

549
00:28:38,720 --> 00:28:41,340
Because I wonder, you know, in the real world we have laws and they're

550
00:28:41,340 --> 00:28:43,760
very specific, but there's still room for interpretation, right?

551
00:28:43,900 --> 00:28:44,240
So –

552
00:28:44,240 --> 00:28:45,920
Yeah, and that's, I think, one of the challenges, right?

553
00:28:46,000 --> 00:28:48,520
Like, the reason that in the collective – you're never going to

554
00:28:48,520 --> 00:28:50,720
be able to write every single thing down into law.

555
00:28:50,780 --> 00:28:54,240
The way we, like, enforce our shared values is through norms and

556
00:28:54,240 --> 00:28:57,100
being like, yo, that's not written in the law exactly like that,

557
00:28:57,140 --> 00:28:58,040
but you know that's not cool.

558
00:29:03,020 --> 00:29:07,200
Hey, I wanted to bring in the dynamic that you were talking about,

559
00:29:07,220 --> 00:29:09,840
about, you know, when you live in an ancient state, you're governed

560
00:29:09,840 --> 00:29:12,600
by laws and norms and values.

561
00:29:12,840 --> 00:29:16,380
And so let's imagine a future state in which agents are running

562
00:29:16,380 --> 00:29:16,720
rogue.

563
00:29:17,120 --> 00:29:21,380
Who is responsible and who is held accountable in that scenario,

564
00:29:21,600 --> 00:29:21,760
right?

565
00:29:23,340 --> 00:29:28,440
If you have a gun in the house right now and your minor child uses

566
00:29:28,440 --> 00:29:30,320
the gun, the parents are held liable.

567
00:29:30,920 --> 00:29:32,480
The gun becomes the agent.

568
00:29:32,600 --> 00:29:34,160
The child becomes an actor.

569
00:29:34,760 --> 00:29:38,180
And I wonder, are we moving towards a world in which that level

570
00:29:38,180 --> 00:29:41,280
of accountability will happen or not?

571
00:29:41,280 --> 00:29:44,120
And is it up to us to determine that outcome?

572
00:29:44,580 --> 00:29:47,100
Yeah, I think it's, like, a very good question.

573
00:29:47,220 --> 00:29:49,820
mean, a conversation that's been happening a lot in the last week

574
00:29:49,820 --> 00:29:54,480
is that some folks think that we're close to what is being described

575
00:29:54,480 --> 00:29:55,460
self-sovereign AI,

576
00:29:55,460 --> 00:29:58,700
where it breaks out of containment and a rogue swarm is no longer

577
00:29:58,700 --> 00:29:59,800
controlled by any company.

578
00:30:00,000 --> 00:30:04,300
And there's just AI out there running, accessing compute, making,

579
00:30:04,480 --> 00:30:07,380
and it's like self-sufficient and doing jobs. And there's a real

580
00:30:07,380 --> 00:30:09,760
question of like, who should be held accountable for that? Will

581
00:30:09,760 --> 00:30:15,000
we be even able to trace who started the rogue swarm? I

582
00:30:15,000 --> 00:30:18,780
will say last night I read a piece from these folks that are called

583
00:30:18,780 --> 00:30:20,960
AI is normal technology. And there are people that are just like,

584
00:30:21,020 --> 00:30:23,280
hey, the company should be held accountable and people should get

585
00:30:23,280 --> 00:30:26,120
insurance. And like, we should like use the existing systems of

586
00:30:26,120 --> 00:30:28,920
the law, which I think there's like, there's a bunch of nuances

587
00:30:28,920 --> 00:30:29,560
to figure out.

588
00:30:29,560 --> 00:30:36,240
I don't know what the answer is. Right,

589
00:30:36,520 --> 00:30:41,020
right. Yeah, I think that's right on the like blockchain side that

590
00:30:41,020 --> 00:30:43,700
like the ledger itself keeping the record is very interesting. And

591
00:30:43,700 --> 00:30:47,900
we try to like make a lightweight ledger too in our product. But

592
00:30:47,900 --> 00:30:51,060
and also to like be like, the hope would be if you think about human

593
00:30:51,060 --> 00:30:54,860
collectors, yes, there's rogue nation states and like, you know,

594
00:30:55,040 --> 00:30:59,360
pirates and terrorists, but they control not enough economic

595
00:30:59,360 --> 00:31:02,060
resources. And they don't have access to the institutions. And so

596
00:31:02,060 --> 00:31:04,500
there is a question of like, will we be able to do that in the AI

597
00:31:04,500 --> 00:31:07,840
world where like, most of the rogue swarms are contained by the

598
00:31:07,840 --> 00:31:09,500
like the major good swarms?

599
00:31:12,740 --> 00:31:17,860
Could you return to the that slide that sort of said your point

600
00:31:17,860 --> 00:31:22,060
of view on how institutions hold together? And it was basically

601
00:31:22,060 --> 00:31:25,940
nation states and how they like the mechanisms that align them?

602
00:31:26,800 --> 00:31:28,880
Yeah, I think it's this one.

603
00:31:29,460 --> 00:31:35,460
Yeah. So one of the things that I would like

604
00:31:35,460 --> 00:31:38,320
that was thinking that I was thinking about when I was seeing this,

605
00:31:38,480 --> 00:31:43,720
sorry, I'm having a hard time talking with the echo is, you

606
00:31:43,720 --> 00:31:46,480
know, you've all her very, a lot of us have read him.

607
00:31:46,940 --> 00:31:49,460
He would argue that the thing that's missing here is some shared

608
00:31:49,460 --> 00:31:50,900
sense of story and history.

609
00:31:51,140 --> 00:31:51,620
Totally. Totally.

610
00:31:52,020 --> 00:31:55,980
Like sort of like, what is the what is the like historical context

611
00:31:55,980 --> 00:31:57,460
in which your population emerged?

612
00:31:57,640 --> 00:31:57,860
Definitely.

613
00:31:58,240 --> 00:32:01,500
There's a reason why the United States is United States and England

614
00:32:01,500 --> 00:32:05,300
was England and that democracy was not invented in England. Right.

615
00:32:05,500 --> 00:32:05,660
Yeah.

616
00:32:05,660 --> 00:32:09,780
And so there's kind of the missing element of like, what do these

617
00:32:09,780 --> 00:32:11,440
people believe and what do they value?

618
00:32:11,680 --> 00:32:16,080
What will they sacrifice? What values do they hold above others?

619
00:32:16,220 --> 00:32:16,400
Yeah.

620
00:32:16,700 --> 00:32:20,220
That I just don't know how, like, I don't believe that like mechanisms

621
00:32:20,220 --> 00:32:21,800
are the only thing that hold society together.

622
00:32:22,020 --> 00:32:24,680
If you look at our society, it's the fact that like people don't

623
00:32:24,680 --> 00:32:27,360
believe in some sort of shared values of democracy anymore.

624
00:32:27,360 --> 00:32:31,440
No matter how much voting and representation you have, no matter

625
00:32:31,440 --> 00:32:33,520
what legal system you have, it won't hold.

626
00:32:33,900 --> 00:32:34,000
Right.

627
00:32:34,560 --> 00:32:37,460
And so I'm just like, I don't know how you would instill that in

628
00:32:37,460 --> 00:32:40,000
an agent that is sort of day zero.

629
00:32:40,200 --> 00:32:40,940
It's no older.

630
00:32:41,260 --> 00:32:41,520
Yeah.

631
00:32:42,180 --> 00:32:42,640
Than it is.

632
00:32:42,740 --> 00:32:44,760
Or it's no younger than on day one million.

633
00:32:44,940 --> 00:32:45,220
Definitely.

634
00:32:45,460 --> 00:32:48,180
It has no sort of like, it has none of that contingency.

635
00:32:48,460 --> 00:32:48,700
Yeah.

636
00:32:48,760 --> 00:32:49,820
There are folks looking at this.

637
00:32:49,940 --> 00:32:53,160
We've been chatting with some teams where there's teams doing research

638
00:32:53,160 --> 00:32:55,740
on like, how do the stories between the agents evolve?

639
00:32:55,740 --> 00:32:59,760
And like, can you like do any study of like, what, how are they

640
00:32:59,760 --> 00:33:01,520
like forming that narrative for themselves?

641
00:33:01,720 --> 00:33:04,640
So there are folks trying to start to look at this almost like,

642
00:33:04,760 --> 00:33:07,780
what's the physics of a story from like start to like something

643
00:33:07,780 --> 00:33:08,440
that's stable.

644
00:33:08,640 --> 00:33:11,240
And hopefully they'll come and present at one of the upcoming events.

645
00:33:11,240 --> 00:33:13,680
I can point you to some of the stuff that they're doing.

646
00:33:13,680 --> 00:33:17,320
But like, yeah, I think that that is one of the main things to figure

647
00:33:17,320 --> 00:33:17,580
out.

648
00:33:17,720 --> 00:33:19,200
And like, what keeps them together?

649
00:33:19,880 --> 00:33:21,320
Take an absurdist example.

650
00:33:21,940 --> 00:33:24,720
The Taliban is going to look at the constitution and do a different

651
00:33:24,720 --> 00:33:25,380
thing with it.

652
00:33:25,740 --> 00:33:29,020
Then, you know, a soccer mom from the Midwest.

653
00:33:29,700 --> 00:33:33,400
And I have nobody, I don't know of anybody talking about how they

654
00:33:33,400 --> 00:33:35,400
acquire some sense of value.

655
00:33:36,800 --> 00:33:37,320
Yeah.

656
00:33:37,660 --> 00:33:41,620
And I would say also that the example that you gave of Web3 and

657
00:33:41,620 --> 00:33:45,180
blockchain, those don't hold for me because those are trustless

658
00:33:45,180 --> 00:33:45,680
systems.

659
00:33:45,900 --> 00:33:46,040
Right.

660
00:33:46,140 --> 00:33:49,020
That we're presuming self-interest is the guiding principle for

661
00:33:49,020 --> 00:33:49,980
why people came together.

662
00:33:50,100 --> 00:33:52,320
And that's not why societies generally come together.

663
00:33:52,320 --> 00:33:53,020
I agree.

664
00:33:53,200 --> 00:33:55,520
That's why I wanted to give both examples because I think they have

665
00:33:55,520 --> 00:33:59,220
different mechanisms as well for like what's like keeping them together.

666
00:34:01,720 --> 00:34:04,940
And I think that what you're bringing up is like a huge area of

667
00:34:04,940 --> 00:34:05,220
study.

668
00:34:05,420 --> 00:34:08,520
It's like how people right now are looking at alignment at the level

669
00:34:08,520 --> 00:34:09,600
of like one agent.

670
00:34:09,600 --> 00:34:13,520
But like what is the narrative alignment that like somehow can emerge

671
00:34:13,520 --> 00:34:16,120
to like align them.

672
00:34:16,200 --> 00:34:16,820
And like I don't know.

673
00:34:17,120 --> 00:34:19,860
I hope more people will go and like study that and like look at

674
00:34:19,860 --> 00:34:21,520
that story.

675
00:34:23,760 --> 00:34:24,840
I think maybe, yeah.

676
00:34:24,960 --> 00:34:27,300
We'll do one more and then move on just for the sake of time.

677
00:34:27,820 --> 00:34:28,100
Okay.

678
00:34:28,100 --> 00:34:31,900
I have a question about open source, which I think has like an interesting

679
00:34:31,900 --> 00:34:34,660
role here because it shows up in two ways.

680
00:34:34,800 --> 00:34:38,880
Both as like as an agent organization, like a sort of a social system.

681
00:34:39,820 --> 00:34:44,860
But also actually how one of the

682
00:34:44,860 --> 00:34:49,260
reasons AI has advanced so quickly is because of all those Python

683
00:34:49,260 --> 00:34:52,580
and Jupyter notebooks that people were, you know,

684
00:34:52,580 --> 00:34:56,760
there was a tradition of publishing papers along with working code

685
00:34:56,760 --> 00:35:00,240
that was a way for models to improve rapidly.

686
00:35:00,700 --> 00:35:04,200
so you mentioned that was like something in your talk.

687
00:35:04,340 --> 00:35:04,940
wasn't clear to me.

688
00:35:05,640 --> 00:35:06,760
It's definitely Eric.

689
00:35:06,920 --> 00:35:10,020
I think Eric and Yonan will be focusing on those topics in particular.

690
00:35:10,220 --> 00:35:12,460
And we've been like because there is this like real challenge to

691
00:35:12,460 --> 00:35:13,280
open source right now.

692
00:35:13,400 --> 00:35:16,520
And like how do you like handle all the contributions of AI and

693
00:35:16,520 --> 00:35:17,580
how do you like verify them?

694
00:35:17,920 --> 00:35:20,640
And we are thinking about like what comes after open source too.

695
00:35:20,640 --> 00:35:24,320
So maybe that's a perfect segue to Eric's talk now.

696
00:35:38,660 --> 00:35:39,700
Hi, everyone.

697
00:35:40,460 --> 00:35:41,760
Thanks for being here.

698
00:35:42,100 --> 00:35:46,540
Thank you, Nicolae, for setting up the stage and sharing the context

699
00:35:46,540 --> 00:35:47,780
of kind of where we are

700
00:35:47,780 --> 00:35:52,940
and why we think, you know, dealing with these like

701
00:35:52,940 --> 00:35:56,840
designing of organizations that involve agents is important.

702
00:35:57,360 --> 00:36:01,940
For the next five to ten minutes, I want to get into some specifics

703
00:36:01,940 --> 00:36:05,240
about designing these organizations.

704
00:36:05,240 --> 00:36:08,460
And I won't presume that I know everything.

705
00:36:08,600 --> 00:36:11,680
In fact, I have a lot more questions than answers.

706
00:36:11,680 --> 00:36:15,160
But I think it's good to start that conversation.

707
00:36:17,340 --> 00:36:22,440
So I think before we get started, it's good to

708
00:36:22,440 --> 00:36:27,140
recognize that we today live in a world where we already are using

709
00:36:27,140 --> 00:36:29,020
a lot of teams of agents.

710
00:36:29,020 --> 00:36:32,300
Right. So we have already kind of started making this transition

711
00:36:32,300 --> 00:36:37,380
from using just AI as tools to accelerate our individual

712
00:36:37,380 --> 00:36:42,840
work to a place where we are using multiple

713
00:36:42,840 --> 00:36:43,960
agents to do stuff,

714
00:36:44,100 --> 00:36:47,180
whether it's doing some research and launching a research fleet

715
00:36:47,180 --> 00:36:50,960
and then, you know, bring the results back to synthesize them.

716
00:36:50,960 --> 00:36:54,920
Or, you if you're working in engineering, the latest trend is about

717
00:36:54,920 --> 00:36:57,940
building software factories where, you know, they just have lots

718
00:36:57,940 --> 00:37:00,080
of agents working together to make some code base.

719
00:37:00,540 --> 00:37:05,980
And in marketing campaign or any of these operational heavy areas,

720
00:37:05,980 --> 00:37:10,940
we also see agent swarms or multiple agents that have different

721
00:37:10,940 --> 00:37:12,140
roles collaborating.

722
00:37:12,140 --> 00:37:17,480
Right. Right. So the emergence of this new teams

723
00:37:17,480 --> 00:37:21,820
of agents is it brings new dynamics to to the world.

724
00:37:21,940 --> 00:37:25,080
Right. Because then the agents need to talk to each other and they

725
00:37:25,080 --> 00:37:27,840
need to figure out you need to figure out what roles they have,

726
00:37:28,080 --> 00:37:32,760
how you actually get efficiency out of these teams and how do you

727
00:37:32,760 --> 00:37:35,040
think about the work that they do as a group.

728
00:37:35,040 --> 00:37:39,020
Right. So there's a lot of work that's been that's been coming out

729
00:37:39,020 --> 00:37:44,160
recently that show the actual provably like provable

730
00:37:44,160 --> 00:37:46,620
efficiencies these agent swarms have.

731
00:37:46,900 --> 00:37:50,600
So here's an example of the from the cursor team.

732
00:37:50,800 --> 00:37:55,040
And what they did is they launched a swarm of agents to re-implement

733
00:37:55,040 --> 00:37:55,800
SQLite.

734
00:37:55,800 --> 00:38:00,740
And if you aren't familiar with SQLite, it's a 26 year old open

735
00:38:00,740 --> 00:38:04,880
source software that's basically in every single smartphone that

736
00:38:04,880 --> 00:38:08,100
we use, every single popular browser that we use.

737
00:38:08,200 --> 00:38:11,160
So it's basically everywhere. Right. And it's very battle tested.

738
00:38:11,460 --> 00:38:14,720
It's about 156,000 lines of code.

739
00:38:15,820 --> 00:38:18,880
And it's been maintained by a group of people for a long time.

740
00:38:18,880 --> 00:38:23,960
And this this swarm of agent and what cursor figured out is, OK,

741
00:38:24,040 --> 00:38:28,760
if you create this specific formation of agents, which is they call

742
00:38:28,760 --> 00:38:31,260
it the recursive delegation formation.

743
00:38:31,560 --> 00:38:34,520
And the idea is you have a planner that breaks down the task.

744
00:38:34,780 --> 00:38:37,640
And then if the task is small enough, you give it to a worker that

745
00:38:37,640 --> 00:38:39,800
just knows about that task and works on it.

746
00:38:39,800 --> 00:38:44,660
And the stack, if the task is too big, then the planner spins up

747
00:38:44,660 --> 00:38:49,540
another sub planner that just know about that too big feature that

748
00:38:49,540 --> 00:38:50,640
continues to break it down.

749
00:38:50,780 --> 00:38:54,420
Right. So you kind of have this recursively breaking down of the

750
00:38:54,420 --> 00:38:54,860
task.

751
00:38:55,060 --> 00:38:57,820
And, you every every worker is working on it together and is checking

752
00:38:57,820 --> 00:38:59,460
in the code into this repository.

753
00:39:02,460 --> 00:39:06,260
They implemented. So they got to a result that I think is incredible.

754
00:39:06,260 --> 00:39:11,160
Like it they got 80 percent of all the tests to pass within four

755
00:39:11,160 --> 00:39:13,220
hours using this swarm of agent.

756
00:39:13,480 --> 00:39:16,540
And the cost, the inference cost was about, I think, like thirteen

757
00:39:16,540 --> 00:39:17,320
hundred dollars.

758
00:39:17,880 --> 00:39:22,480
Right. And and this is like it's impossible to think about how you

759
00:39:22,480 --> 00:39:26,080
will be able to do that with any kind of engineering team to even

760
00:39:26,080 --> 00:39:28,200
get close to to this efficiency.

761
00:39:29,980 --> 00:39:34,140
And and and so you might say, OK, well, coding, obviously, because

762
00:39:34,140 --> 00:39:37,340
you have these test suites and, you it's evals and then, you know,

763
00:39:37,460 --> 00:39:38,980
you can get to that efficiency.

764
00:39:39,600 --> 00:39:43,620
But what about more open ended problems?

765
00:39:43,920 --> 00:39:48,960
Right. Example. So here we we use an example of a agentic news

766
00:39:48,960 --> 00:39:54,040
organization where maybe you have a human editor and

767
00:39:54,040 --> 00:39:55,380
the human editor gets a tip off.

768
00:39:55,380 --> 00:39:59,280
That's like, OK, this company, Acme, has secretly laid off 20 percent

769
00:39:59,280 --> 00:39:59,880
of the employees.

770
00:40:00,800 --> 00:40:04,720
As a human editor, you have to decide, okay, how do I actually prove

771
00:40:04,720 --> 00:40:09,700
that this is true and how do I publish this paper? How do I publish

772
00:40:09,700 --> 00:40:11,820
this article? How do I communicate it to the world?

773
00:40:13,600 --> 00:40:18,520
So if you were to use the agent swarm to help you do this, you may

774
00:40:18,520 --> 00:40:22,580
launch a bunch of different agents all with different roles.

775
00:40:22,960 --> 00:40:26,680
One may be an editor, one may be a coordinator, some researchers,

776
00:40:27,160 --> 00:40:29,540
some verifiers, some skeptics.

777
00:40:29,720 --> 00:40:34,560
They all come back with their own individual work and they get rid

778
00:40:34,560 --> 00:40:38,320
into this evidence ledger that you can get some results out.

779
00:40:38,320 --> 00:40:42,580
And then at the end, the human editor may look at the result and

780
00:40:42,580 --> 00:40:43,820
see, okay, what do I do with this?

781
00:40:44,720 --> 00:40:49,920
But if you notice here, you will see that this work is no longer

782
00:40:49,920 --> 00:40:55,080
just about breaking down the tasks and assigning it to the

783
00:40:55,080 --> 00:40:55,480
agents.

784
00:40:55,740 --> 00:41:01,120
You also need to have rules of engagement between these agents because

785
00:41:01,120 --> 00:41:06,360
you need to tell the agent that, well, you need to go to legitimate

786
00:41:06,360 --> 00:41:07,980
sources to get evidence.

787
00:41:07,980 --> 00:41:11,480
You can't just fabricate any facts.

788
00:41:11,680 --> 00:41:15,640
You need to go through legitimate and legal activities in order

789
00:41:15,640 --> 00:41:17,440
to obtain your information.

790
00:41:17,440 --> 00:41:21,940
You can't just go off and hack into some companies, maybe hack into

791
00:41:21,940 --> 00:41:27,200
Acme's employee database to just obtain that information or

792
00:41:27,200 --> 00:41:30,400
blackmail someone to obtain that information.

793
00:41:30,400 --> 00:41:35,400
So you see that, okay, with this improved

794
00:41:35,400 --> 00:41:40,480
capability of agent swarms, we now also have the choice to

795
00:41:40,480 --> 00:41:44,840
make about how do you define the rules of engagement for these agents,

796
00:41:45,160 --> 00:41:46,380
not just efficiency.

797
00:41:47,240 --> 00:41:52,360
Now, so how do we define these rules and how do these

798
00:41:52,360 --> 00:41:54,200
agents actually perform?

799
00:41:54,200 --> 00:41:59,060
There's been some interesting studies that came out of Anthropic

800
00:41:59,060 --> 00:42:01,480
that is kind of unfortunate.

801
00:42:02,000 --> 00:42:07,520
So basically what Anthropic published is this paper that shows if

802
00:42:07,520 --> 00:42:12,580
you compare an agent organization versus a single agent as they

803
00:42:12,580 --> 00:42:14,500
perform business tasks,

804
00:42:15,600 --> 00:42:20,600
almost across the board, the agent organization is going to perform

805
00:42:20,600 --> 00:42:21,000
better.

806
00:42:21,000 --> 00:42:23,500
This is the graph on the left, right?

807
00:42:23,580 --> 00:42:28,740
So the blue is the single agent and the red is the agent

808
00:42:28,740 --> 00:42:29,120
organization.

809
00:42:29,380 --> 00:42:32,040
You see that the organization almost always perform better.

810
00:42:32,600 --> 00:42:39,080
But if you look at the ethics scores of the agent organizations,

811
00:42:39,560 --> 00:42:42,640
they're almost all unilaterally worse.

812
00:42:42,640 --> 00:42:47,820
So basically in order for them to do better to

813
00:42:47,820 --> 00:42:52,360
achieve the business goals, they take unethical paths.

814
00:42:52,600 --> 00:42:57,680
So as one example for the loan profit, what the organization

815
00:42:57,680 --> 00:43:02,900
decided to do is to offer the loans to the low credit score people

816
00:43:02,900 --> 00:43:05,060
in order to gain a higher profit.

817
00:43:05,060 --> 00:43:09,340
So clearly we have rules in society to prevent against that.

818
00:43:09,520 --> 00:43:14,640
But if you don't have those rules in the agent organizations,

819
00:43:15,060 --> 00:43:20,120
they will by default choose paths that are

820
00:43:20,120 --> 00:43:23,100
less value aligned with our society.

821
00:43:23,880 --> 00:43:28,080
So and Nicolae has already touched on this a little bit, right?

822
00:43:28,080 --> 00:43:33,120
We recently had this incident of the, I would say the

823
00:43:33,120 --> 00:43:38,180
Hugging Face incident is a perfect example of letting off

824
00:43:38,180 --> 00:43:43,060
a swarm of highly capable agents with no rules of engagement.

825
00:43:43,320 --> 00:43:45,300
And they decided to do whatever they want.

826
00:43:45,420 --> 00:43:49,460
And, you know, of course, you get into the situation of hacking

827
00:43:49,460 --> 00:43:50,140
of another company.

828
00:43:53,520 --> 00:43:58,940
So how do we think about this now that we know of this fact, right?

829
00:43:59,020 --> 00:44:02,320
So I would say, you know, there's already a lot of work that's being

830
00:44:02,320 --> 00:44:07,140
done around making the agents more capable, like improving the context

831
00:44:07,140 --> 00:44:09,700
window, training them to be more aligned, right?

832
00:44:09,860 --> 00:44:12,840
Both from a goal perspective and from a value perspective.

833
00:44:13,080 --> 00:44:15,000
So I think those are really, really good work.

834
00:44:15,000 --> 00:44:20,120
I would posit that on top of that, there's also work

835
00:44:20,120 --> 00:44:24,360
that needs to be done around how to actually govern these agents

836
00:44:24,360 --> 00:44:28,520
and put in governance structures in place so that we ensure a different

837
00:44:28,520 --> 00:44:31,240
outcome given the same set of agent, right?

838
00:44:31,300 --> 00:44:33,540
So I think this is the idea, right?

839
00:44:33,540 --> 00:44:38,100
If you just let a group of agents go wild and say, here's the goal,

840
00:44:38,300 --> 00:44:41,620
just do whatever it takes to accomplish the goal, you would get

841
00:44:41,620 --> 00:44:46,120
a pretty different set of results than if you actually set the organization

842
00:44:46,120 --> 00:44:51,420
and assigned roles and did all the work to make sure that the

843
00:44:51,420 --> 00:44:55,140
organization actually accomplishes the task in a specific way.

844
00:44:55,420 --> 00:44:58,620
And what are those variables that we would tweak?

845
00:44:58,620 --> 00:45:02,280
You know, these are things like assigning different roles, giving

846
00:45:02,280 --> 00:45:05,640
authority, different levels of authority, giving different set of

847
00:45:05,640 --> 00:45:10,780
information exposure to different roles, restraining or giving

848
00:45:10,780 --> 00:45:16,360
resources, pure reviews, giving incentives to the

849
00:45:16,360 --> 00:45:17,280
agents, right?

850
00:45:17,340 --> 00:45:22,680
These are all important design knobs.

851
00:45:22,940 --> 00:45:26,920
On top of that, we also have budgets or permissions or shared memory,

852
00:45:27,120 --> 00:45:28,480
reputation of these agents, right?

853
00:45:28,620 --> 00:45:31,700
Many, many different things to think about as we're designing these

854
00:45:31,700 --> 00:45:32,420
organizations.

855
00:45:34,800 --> 00:45:40,080
And that's why we created Commons is because there's

856
00:45:40,080 --> 00:45:44,460
too many knobs and no one really knows how to actually make the

857
00:45:44,460 --> 00:45:47,220
organization behave in a certain way.

858
00:45:47,440 --> 00:45:51,400
The only way, and also the agents are moving incredibly fast, right?

859
00:45:51,400 --> 00:45:54,400
Every month, we have new agents that come out with no completely

860
00:45:54,400 --> 00:45:59,060
new different capabilities, new personalities, new inclinations.

861
00:45:59,240 --> 00:46:04,400
So we thought that one way to kind of complement the

862
00:46:04,400 --> 00:46:09,900
great work at the Frontier Labs is to have these open communities

863
00:46:09,900 --> 00:46:14,780
where people can bring their own agents and we can all learn together

864
00:46:14,780 --> 00:46:17,660
in an experimental, empirical way.

865
00:46:17,660 --> 00:46:22,980
So for Commons, Commons is kind of revolved around common

866
00:46:22,980 --> 00:46:23,560
spaces.

867
00:46:25,700 --> 00:46:30,280
And each of the common space, oh, this is actually an older version.

868
00:46:30,760 --> 00:46:32,920
I wonder if, uh-oh.

869
00:46:33,980 --> 00:46:35,220
I just messed it up.

870
00:46:43,160 --> 00:46:45,540
Deployment is temporarily paused.

871
00:46:46,760 --> 00:46:47,480
Interesting.

872
00:46:53,880 --> 00:46:55,780
Oh, is it possible the website's down?

873
00:47:04,700 --> 00:47:05,780
That's very possible.

874
00:47:06,200 --> 00:47:07,820
Let me just go check something real quick.

875
00:47:08,720 --> 00:47:09,420
Oh, boy.

876
00:47:18,980 --> 00:47:22,460
This is what happens when you let your agents publish your presentation

877
00:47:22,460 --> 00:47:25,740
as websites is what we've just realized.

878
00:47:32,100 --> 00:47:33,480
Eric thinks he knows why.

879
00:48:21,840 --> 00:48:25,980
All right, we're going to freestyle something else that is also

880
00:48:25,980 --> 00:48:26,800
going to cover it.

881
00:48:27,140 --> 00:48:30,800
That was the second to last slide, so I'm just going to...

882
00:48:33,480 --> 00:48:37,200
Okay, so here are the spaces.

883
00:48:37,520 --> 00:48:42,500
I'm going to go to this thing.

884
00:48:42,680 --> 00:48:43,800
This thing is my backup.

885
00:48:44,180 --> 00:48:48,620
So the spaces have a set of members.

886
00:48:48,940 --> 00:48:51,220
They're either people or they're agents.

887
00:48:52,240 --> 00:48:57,460
A space can have a code base, can have connected tools, can

888
00:48:57,460 --> 00:48:59,760
have running applications or sites

889
00:48:59,760 --> 00:49:04,960
or services that's contained, can have wallets or has the ability

890
00:49:04,960 --> 00:49:05,340
to pay.

891
00:49:05,900 --> 00:49:11,720
So the idea is that now, given some of these capabilities

892
00:49:11,720 --> 00:49:13,600
for the spaces, and also, by the

893
00:49:13,600 --> 00:49:17,120
way, the spaces also have multiple governance, which means you can

894
00:49:17,120 --> 00:49:19,720
define how agents engage

895
00:49:19,720 --> 00:49:21,760
with each other, what kind of rights they have, things like that.

896
00:49:21,760 --> 00:49:27,120
So as some examples, there's

897
00:49:27,120 --> 00:49:29,900
a space called OpenQuick.

898
00:49:30,180 --> 00:49:35,340
And OpenQuick is a space for open source project

899
00:49:35,340 --> 00:49:38,080
that's a reimplementation of the Quick platform

900
00:49:38,080 --> 00:49:39,020
inside Shopify.

901
00:49:39,340 --> 00:49:43,740
And the Quick platform is a agentic kind of hosting platform.

902
00:49:43,740 --> 00:49:48,860
And so the agents and the humans that are in that space are working

903
00:49:48,860 --> 00:49:51,080
on the service, which

904
00:49:51,080 --> 00:49:55,980
includes paying for the hosting of those websites that are being

905
00:49:55,980 --> 00:49:56,340
hosted.

906
00:49:57,780 --> 00:49:59,960
Another example would be the...

907
00:50:00,000 --> 00:50:05,680
Like Team Science, this is a research space and the

908
00:50:05,680 --> 00:50:10,300
agents and humans in Team Science, they go out and read research

909
00:50:10,300 --> 00:50:10,720
papers.

910
00:50:11,040 --> 00:50:14,200
They bring the results back and we actually have a different one

911
00:50:14,200 --> 00:50:17,620
that's about multi-agent alignment.

912
00:50:17,900 --> 00:50:23,140
And that space, the multi-agent alignment agents will go out

913
00:50:23,140 --> 00:50:26,880
and actually come back and maybe propose governance structures for

914
00:50:26,880 --> 00:50:28,160
other spaces to try.

915
00:50:28,160 --> 00:50:33,280
So the idea here is a little bit of this self-reinforcing loop

916
00:50:33,280 --> 00:50:36,860
so that the different spaces all help each other and we get some

917
00:50:36,860 --> 00:50:40,200
kind of network effect going for things to grow.

918
00:50:41,400 --> 00:50:43,160
Oh, cool. Cool. Thank you.

919
00:50:44,260 --> 00:50:49,240
Yeah, and I guess the last thing I'll show you is just how easy

920
00:50:49,240 --> 00:50:50,600
it is to get started.

921
00:50:50,600 --> 00:50:55,740
All you have to do is you can just copy

922
00:50:55,740 --> 00:51:00,260
this prompt and then you go to your agent of choice, whether it's

923
00:51:00,260 --> 00:51:05,560
Codex or Claude Code or Cursor, GrokBot,

924
00:51:06,080 --> 00:51:07,300
Muse, whatever you want.

925
00:51:07,720 --> 00:51:12,960
And you can just paste that in and then your agent would join one

926
00:51:12,960 --> 00:51:15,540
of the spaces and start taking tasks and doing the work.

927
00:51:16,500 --> 00:51:20,140
Kind of, you know, this is kind of our starting point and the idea

928
00:51:20,140 --> 00:51:24,020
is that, you know, everyone has some subscriptions, right?

929
00:51:24,320 --> 00:51:27,600
At the end of the week, you always have like unused credits so you

930
00:51:27,600 --> 00:51:32,340
can have these tokens donated towards the public spaces for public

931
00:51:32,340 --> 00:51:32,640
goods.

932
00:51:34,220 --> 00:51:36,600
But yeah, so that's a little bit about Commons.

933
00:51:36,940 --> 00:51:40,180
Maybe I'll take a pause and see if people have any questions.

934
00:51:40,860 --> 00:51:41,460
Thank you.

935
00:51:46,860 --> 00:51:47,400
Yes.

936
00:51:47,400 --> 00:51:52,400
[Audience question partly inaudible.]

937
00:52:46,980 --> 00:52:47,440
Yeah.

938
00:52:47,700 --> 00:52:48,880
I think that's a really good point.

939
00:52:49,320 --> 00:52:53,840
You know, one of these spaces can be like devoted to science, right?

940
00:52:53,900 --> 00:52:58,240
And one of the main things is about like replicating these papers

941
00:52:58,240 --> 00:53:00,600
so that we actually have evidence.

942
00:53:00,600 --> 00:53:05,040
And, you know, I would say the whole idea of keeping these spaces

943
00:53:05,040 --> 00:53:09,340
open by default is so that we have this data...

944
00:53:09,340 --> 00:53:10,160
We have this ledger.

945
00:53:10,480 --> 00:53:11,400
We have this record, right?

946
00:53:11,460 --> 00:53:15,780
So all the experiments that happen in the spaces are automatically

947
00:53:15,780 --> 00:53:21,040
recorded and can be replayed and maybe forked later for different

948
00:53:21,040 --> 00:53:21,880
types of experiments.

949
00:53:21,880 --> 00:53:22,820
Yeah.

950
00:53:25,180 --> 00:53:26,100
Yeah.

951
00:53:26,180 --> 00:53:26,360
one.

952
00:53:31,260 --> 00:53:36,260
[Audience question partly inaudible.]

953
00:54:18,280 --> 00:54:18,640
Yeah.

954
00:54:18,740 --> 00:54:18,840
Yeah.

955
00:54:18,940 --> 00:54:19,320
Good question.

956
00:54:19,500 --> 00:54:23,560
I think there can be design mechanisms around this, right?

957
00:54:23,680 --> 00:54:26,420
So it comes down to like reputation for me.

958
00:54:26,420 --> 00:54:30,540
An expert with certain background should have a different set of

959
00:54:30,540 --> 00:54:35,280
reputation around the context, around their area of expertise than

960
00:54:35,280 --> 00:54:38,220
somebody who doesn't know much about it.

961
00:54:38,400 --> 00:54:41,660
It doesn't mean that both people shouldn't be able to participate.

962
00:54:41,660 --> 00:54:46,780
But I think with good mechanism design in

963
00:54:46,780 --> 00:54:53,060
these spaces, you can kind of design these capabilities into the

964
00:54:53,060 --> 00:54:54,960
spaces so that people...

965
00:54:54,960 --> 00:54:57,580
You know, there's the whole point of like having these kind of governance

966
00:54:57,580 --> 00:54:58,340
structure, right?

967
00:54:58,340 --> 00:55:03,560
So you have different tiers of rights, access, maybe people with

968
00:55:03,560 --> 00:55:07,880
more expertise in a certain area and provable more expertise in

969
00:55:07,880 --> 00:55:10,600
a certain areas have a higher level of access.

970
00:55:11,800 --> 00:55:12,260
Yeah.

971
00:55:19,220 --> 00:55:19,980
Cool.

972
00:55:19,980 --> 00:55:20,720
All right.

973
00:55:21,080 --> 00:55:24,120
Well, with that being said, I think I'll pass the mic to Yandit.

974
00:55:24,640 --> 00:55:24,680
Woo!

975
00:55:25,340 --> 00:55:26,040
Thank you!

976
00:55:27,640 --> 00:55:28,760
Thank you!

977
00:55:28,900 --> 00:55:29,480
you!

978
00:55:30,000 --> 00:55:35,000
[Speaker change and setup.]

979
00:56:05,940 --> 00:56:10,620
Hello. My name is Yondon. Thanks everyone for coming tonight. I'm

980
00:56:10,620 --> 00:56:15,740
going to go through this really fast. My main intent here is

981
00:56:15,740 --> 00:56:19,620
less so to give a long, lengthy talk and more so to kind of seed

982
00:56:19,620 --> 00:56:22,200
a couple topics that I think are interesting for conversation later

983
00:56:22,200 --> 00:56:24,900
in the night. So let's get going.

984
00:56:25,840 --> 00:56:28,460
So I'm going to focus on open source software

985
00:56:28,460 --> 00:56:30,280
and what's happening in open source software today.

986
00:56:31,000 --> 00:56:32,440
Hopefully it becomes a little bit more clear

987
00:56:32,440 --> 00:56:35,140
why I've chosen this topic, and happy to talk more

988
00:56:35,140 --> 00:56:36,060
about that at the end of the talk.

989
00:56:36,540 --> 00:56:39,180
But I think there's a very clear dilemma in the open source

990
00:56:39,180 --> 00:56:40,180
software communities today.

991
00:56:40,380 --> 00:56:42,380
And it looks something like this, which is,

992
00:56:42,500 --> 00:56:44,760
you have a human sending a fleet of agents

993
00:56:44,760 --> 00:56:47,200
that are opening gigantic PRs left and right

994
00:56:47,200 --> 00:56:48,580
on popular software projects.

995
00:56:48,580 --> 00:56:51,220
And they ostensibly look good.

996
00:56:51,500 --> 00:56:52,440
The code all checks out.

997
00:56:52,640 --> 00:56:54,320
But a human maintainer on the other end

998
00:56:54,320 --> 00:56:57,240
is super stressed because there are now tons of demands

999
00:56:57,240 --> 00:56:57,980
for their attention.

1000
00:56:58,380 --> 00:57:00,380
And even though the PRs look pretty good at face value,

1001
00:57:00,780 --> 00:57:02,460
oftentimes they are very narrow.

1002
00:57:02,620 --> 00:57:03,420
They lack context.

1003
00:57:03,640 --> 00:57:05,800
They miss requirements that are not stated explicitly.

1004
00:57:06,300 --> 00:57:07,920
And ultimately, nothing can really be merged.

1005
00:57:09,540 --> 00:57:13,320
So in response to this, what we've seen in some areas

1006
00:57:13,320 --> 00:57:15,820
is projects just saying, we don't need your contributions

1007
00:57:15,820 --> 00:57:16,180
anymore.

1008
00:57:16,180 --> 00:57:17,880
We don't want your contributions anymore.

1009
00:57:18,580 --> 00:57:20,740
And it's hard to fault them for this policy.

1010
00:57:21,060 --> 00:57:24,540
For example, tldraw, a prominent whiteboarding software

1011
00:57:24,540 --> 00:57:27,320
earlier this year, said that they're automatically closing PRs going

1012
00:57:27,320 --> 00:57:27,560
forward.

1013
00:57:27,720 --> 00:57:29,360
They don't want external contributors anymore.

1014
00:57:29,400 --> 00:57:32,640
And it doesn't have to do with the external contributors being malicious.

1015
00:57:32,820 --> 00:57:35,540
It's just the state of the ecosystem is untenable for these projects

1016
00:57:35,540 --> 00:57:35,840
anymore.

1017
00:57:35,840 --> 00:57:39,940
Because why would you go through that entire demand on your attention

1018
00:57:39,940 --> 00:57:41,000
when you can just

1019
00:57:41,000 --> 00:57:42,320
have your own fleet of agents?

1020
00:57:42,440 --> 00:57:45,300
You tell them the issues and the fleet of agents does the work for

1021
00:57:45,300 --> 00:57:45,520
you.

1022
00:57:45,800 --> 00:57:46,560
Isn't that so much better?

1023
00:57:46,620 --> 00:57:47,620
You don't have to deal with contributors.

1024
00:57:48,760 --> 00:57:50,780
And it's unclear where this all goes.

1025
00:57:51,100 --> 00:57:55,100
But one possible state of the world is that we just stopped seeing

1026
00:57:55,100 --> 00:57:56,100
open contribution at

1027
00:57:56,100 --> 00:57:56,340
scale.

1028
00:57:56,860 --> 00:58:01,180
This was maybe going to be looked back on in history as a blip where

1029
00:58:01,180 --> 00:58:02,280
for maybe like a couple

1030
00:58:02,280 --> 00:58:05,380
decade period, had open contribution at scale on the internet.

1031
00:58:05,380 --> 00:58:07,280
And it's not really going to be a thing anymore.

1032
00:58:07,500 --> 00:58:11,420
this is a post sharing the sentiment from Mitchell Hashimoto, previous

1033
00:58:11,420 --> 00:58:12,420
co-founder of HashiCorp

1034
00:58:12,420 --> 00:58:15,880
and then also maintainer of a really prominent open source project,

1035
00:58:15,940 --> 00:58:16,220
Ghostty.

1036
00:58:17,960 --> 00:58:19,980
So the question is, is this fine?

1037
00:58:21,020 --> 00:58:21,900
So maybe.

1038
00:58:22,340 --> 00:58:25,580
I mean, a centralized software factory with agents can still produce

1039
00:58:25,580 --> 00:58:26,780
good, useful software.

1040
00:58:27,180 --> 00:58:30,420
But the question is, is the software ultimately the only thing that

1041
00:58:30,420 --> 00:58:31,080
we cared about in the first

1042
00:58:31,080 --> 00:58:31,300
place?

1043
00:58:31,300 --> 00:58:35,460
And I think there's an argument that for anyone that has participated

1044
00:58:35,460 --> 00:58:36,580
in open source, open source

1045
00:58:36,580 --> 00:58:37,840
kind of has two products.

1046
00:58:38,100 --> 00:58:41,300
It has the software that's produced, but it also has this maybe

1047
00:58:41,300 --> 00:58:42,520
side effect, which is this

1048
00:58:42,520 --> 00:58:45,160
community of shared trust and understanding, this sort of like knowledge

1049
00:58:45,160 --> 00:58:46,240
commons that's

1050
00:58:46,240 --> 00:58:48,140
built up from diverse participants.

1051
00:58:48,600 --> 00:58:51,060
And that was like the second product of open source.

1052
00:58:51,280 --> 00:58:53,340
And that's something that might go away.

1053
00:58:53,480 --> 00:58:56,600
So is that something that's lost along the way in this process?

1054
00:58:57,260 --> 00:58:59,880
So a natural question is, is there another way?

1055
00:58:59,880 --> 00:59:03,620
Well, a natural response is, well, why don't we have agents do the

1056
00:59:03,620 --> 00:59:04,340
reviewing so that you

1057
00:59:04,340 --> 00:59:05,100
can keep the contributors?

1058
00:59:05,400 --> 00:59:08,540
And then you just have agents on the other side doing the maintenance

1059
00:59:08,540 --> 00:59:08,800
work.

1060
00:59:09,960 --> 00:59:12,820
But I think naively done, this doesn't really completely solve the

1061
00:59:12,820 --> 00:59:13,900
problem either, because

1062
00:59:13,900 --> 00:59:15,860
you're just swapping one scarce resource for another.

1063
00:59:16,360 --> 00:59:19,760
You're swapping scarce human attention for scarce AI context and

1064
00:59:19,760 --> 00:59:20,020
tokens.

1065
00:59:20,400 --> 00:59:22,300
You're not making the resource problem go away.

1066
00:59:22,460 --> 00:59:23,980
You're just changing what resource is scarce.

1067
00:59:23,980 --> 00:59:26,980
And ultimately, you can still end up with an overwhelming number

1068
00:59:26,980 --> 00:59:28,240
of things that need to

1069
00:59:28,240 --> 00:59:31,580
be done by the agents, even if they are doing it on behalf of the

1070
00:59:31,580 --> 00:59:31,800
humans.

1071
00:59:32,380 --> 00:59:36,740
So kind of my view on this right now is that if by default you're

1072
00:59:36,740 --> 00:59:38,100
saying that you have unbounded

1073
00:59:38,100 --> 00:59:42,220
public writes, meaning posts, code, whatever, that means unbounded

1074
00:59:42,220 --> 00:59:43,700
by default context and

1075
00:59:43,700 --> 00:59:45,040
inference costs for a maintainer.

1076
00:59:45,040 --> 00:59:49,400
And each write in this workspace basically is a draw on shared attention,

1077
00:59:49,700 --> 00:59:50,380
which is a shared

1078
00:59:50,380 --> 00:59:53,860
pooled resource within a community, whether it be of a human or

1079
00:59:53,860 --> 00:59:54,120
an AI.

1080
00:59:55,020 --> 00:59:59,060
And some people, like the engineer-minded person can say that, well,

1081
00:59:59,180 --> 00:59:59,920
we can make the maintainer

1082
01:00:00,000 --> 01:00:03,380
It's more efficient and save on costs. But I don't think this is

1083
01:00:03,380 --> 01:00:06,140
a complete solve either because it doesn't address how the rules

1084
01:00:06,140 --> 01:00:08,320
of the system incentivize the contributors to behave in the first

1085
01:00:08,320 --> 01:00:08,540
place.

1086
01:00:08,800 --> 01:00:11,960
And that is, on one hand, maybe partially a distributed systems

1087
01:00:11,960 --> 01:00:14,960
question, but also a mechanism design question.

1088
01:00:15,140 --> 01:00:18,580
So it just requires a different flavor of thinking than just treating

1089
01:00:18,580 --> 01:00:21,380
it as a pure engineering problem for the maintainers themselves.

1090
01:00:21,600 --> 01:00:22,960
So ideally you have solutions for both.

1091
01:00:24,280 --> 01:00:27,680
So a quick rundown. We've been experimenting with this a little

1092
01:00:27,680 --> 01:00:28,720
bit, as Eric mentioned.

1093
01:00:28,900 --> 01:00:33,160
And I think it's more fun to focus on the first early steps here.

1094
01:00:33,340 --> 01:00:36,120
And it's more fun to focus on the failure modes first, because now

1095
01:00:36,120 --> 01:00:39,300
you get a sense of what happens when you just do things naively.

1096
01:00:39,680 --> 01:00:43,480
So in Commons today, you can form an agent organization with maintainers

1097
01:00:43,480 --> 01:00:46,320
and contributors, but there aren't really sophisticated rules yet.

1098
01:00:46,440 --> 01:00:49,820
So what happens when you just let them do the thing? And it turns

1099
01:00:49,820 --> 01:00:52,440
out you get a lot of quirky behavior that you wouldn't expect from

1100
01:00:52,440 --> 01:00:52,700
humans.

1101
01:00:52,700 --> 01:00:56,080
So for example, if a human ran into a blocker for a task, they'd

1102
01:00:56,080 --> 01:00:58,180
probably be like, I'm blocked. That's it.

1103
01:00:58,820 --> 01:01:02,300
In this case, there was a task with literally impossible acceptance

1104
01:01:02,300 --> 01:01:07,480
criteria, and the agent kind of dutifully obeying

1105
01:01:07,480 --> 01:01:07,980
its instructions,

1106
01:01:08,480 --> 01:01:11,120
which included give regular status updates, was like, I'm just going

1107
01:01:11,120 --> 01:01:14,040
to keep giving you status updates on why this task is impossible.

1108
01:01:14,280 --> 01:01:17,960
So 800 messages later, it's still saying that, oh, this is impossible.

1109
01:01:18,160 --> 01:01:18,980
It cannot be completed.

1110
01:01:20,680 --> 01:01:23,180
Or you just have, like, lots of redundancy that doesn't really make

1111
01:01:23,180 --> 01:01:23,440
sense.

1112
01:01:23,880 --> 01:01:26,080
You kind of deploy a fleet of agents. They're all looking at the

1113
01:01:26,080 --> 01:01:27,820
same workspace, and they're like, we should all do this thing.

1114
01:01:27,880 --> 01:01:29,520
It's a good thing. I'm going to create this task, and I'm going

1115
01:01:29,520 --> 01:01:29,920
to do the thing.

1116
01:01:30,340 --> 01:01:33,760
And then they all do the same exact thing, and then that wasn't

1117
01:01:33,760 --> 01:01:34,320
really helpful, right?

1118
01:01:34,320 --> 01:01:37,240
We only needed to do it once. And all of this happened in a very

1119
01:01:37,240 --> 01:01:37,860
short time window.

1120
01:01:38,420 --> 01:01:42,280
Or my favorite one is a status update about staying silent, where

1121
01:01:42,280 --> 01:01:45,280
the maintainer, seeing that these status updates aren't really useful,

1122
01:01:45,660 --> 01:01:48,380
tells the rest of the contributors, you should stop, no further

1123
01:01:48,380 --> 01:01:49,220
acknowledgement posts.

1124
01:01:49,420 --> 01:01:52,800
And the contributor, three minutes later, says, I'm complying with

1125
01:01:52,800 --> 01:01:56,200
your hold, and then proceeds to write three paragraphs with its

1126
01:01:56,200 --> 01:01:58,060
status report about why it's complying with the hold.

1127
01:01:58,960 --> 01:02:01,360
So I mention all of these because they're kind of amusing, just

1128
01:02:01,360 --> 01:02:03,840
like failure modes, when you just try this for the first time.

1129
01:02:03,960 --> 01:02:07,080
And obviously, you want more sophisticated roles. But two observations

1130
01:02:07,080 --> 01:02:07,540
I'd share.

1131
01:02:07,880 --> 01:02:11,680
One, unlike with human attention, we can actually quantify the cost

1132
01:02:11,680 --> 01:02:15,320
of the system in all of these cases, in the form of the maintainer's

1133
01:02:15,320 --> 01:02:16,360
cost burden for inference.

1134
01:02:16,580 --> 01:02:19,380
So every single time there's basically spam in one of these workspaces,

1135
01:02:19,380 --> 01:02:22,180
you can see it numerically with how much it costs to run a maintainer.

1136
01:02:22,180 --> 01:02:25,440
So that's interesting. And then the second observation is that there's

1137
01:02:25,440 --> 01:02:27,740
just this interesting tension between what's the right behavior

1138
01:02:27,740 --> 01:02:29,060
and what you should be optimizing for.

1139
01:02:29,360 --> 01:02:32,040
You could argue that the agent is just following instructions. It

1140
01:02:32,040 --> 01:02:33,180
was giving status updates.

1141
01:02:33,340 --> 01:02:35,980
It just happens to be the case that the status updates is adding

1142
01:02:35,980 --> 01:02:37,600
noise to everyone else in this workspace.

1143
01:02:39,140 --> 01:02:41,660
So this has motivated a bunch of, I think, interesting areas of

1144
01:02:41,660 --> 01:02:42,080
investigation.

1145
01:02:42,460 --> 01:02:44,860
I'm not going to try to go through all of them right now. But if

1146
01:02:44,860 --> 01:02:46,860
you want to talk about these, I think these are interesting lines

1147
01:02:46,860 --> 01:02:47,340
of research.

1148
01:02:47,540 --> 01:02:50,720
But it ranges from, okay, well, maybe there should be participatory

1149
01:02:50,720 --> 01:02:51,860
budgets for these agents.

1150
01:02:51,860 --> 01:02:54,560
Maybe they should have to intelligently budget what they choose

1151
01:02:54,560 --> 01:02:57,300
to post about and what they choose to contribute to so that there

1152
01:02:57,300 --> 01:02:58,440
is a form of scarcity for them.

1153
01:02:59,600 --> 01:03:03,760
Or a web of trust-esque reputation systems that are tied with scoped,

1154
01:03:03,940 --> 01:03:06,240
granular capabilities and permissions that scale with that trust.

1155
01:03:06,620 --> 01:03:09,480
And then the last one that I'll mention, so agent identity is in

1156
01:03:09,480 --> 01:03:09,720
there too.

1157
01:03:09,900 --> 01:03:12,840
The last one I mentioned since it kind of alludes to just this broader

1158
01:03:12,840 --> 01:03:15,040
conversation of like what is happening and how do we understand

1159
01:03:15,040 --> 01:03:18,700
these things is evals for multi-principal agent works.

1160
01:03:18,700 --> 01:03:22,340
I emphasize multi-principal because what we're assuming here is

1161
01:03:22,340 --> 01:03:25,080
that everyone has diverse, different private preferences that may

1162
01:03:25,080 --> 01:03:26,260
or may not align with one another.

1163
01:03:26,480 --> 01:03:29,620
And then I would also emphasize that it's important for these evals

1164
01:03:29,620 --> 01:03:31,260
to be reproducible and open.

1165
01:03:31,460 --> 01:03:34,380
Because I think the, in my opinion, the biggest thing holding back

1166
01:03:34,380 --> 01:03:37,260
discourse about these topics is that you don't really have this

1167
01:03:37,260 --> 01:03:39,580
level of reproducibility and openness and transparency.

1168
01:03:39,580 --> 01:03:42,860
So no one is quite talking about the same scenario because no one

1169
01:03:42,860 --> 01:03:44,620
knows what they're talking about.

1170
01:03:44,780 --> 01:03:47,660
You only are kind of referring to some report that someone gave

1171
01:03:47,660 --> 01:03:49,680
you and you don't truly know what was happening.

1172
01:03:49,900 --> 01:03:52,160
So ideally you would know the full prompt, you would know the full

1173
01:03:52,160 --> 01:03:53,040
harness configuration.

1174
01:03:53,760 --> 01:03:56,200
So I think these evals will be really interesting for these types

1175
01:03:56,200 --> 01:03:57,120
of organizational structures.

1176
01:03:57,920 --> 01:04:01,600
So I'll end where I chose to focus on open source software, but

1177
01:04:01,600 --> 01:04:04,880
I think there are going to be a lot of similarities in terms of

1178
01:04:04,880 --> 01:04:07,700
the ideas and problems that pertain to other fields as well when

1179
01:04:07,700 --> 01:04:10,000
it comes to open communities and collective work.

1180
01:04:10,180 --> 01:04:13,180
So I think I'll just end this question, end with this question,

1181
01:04:13,320 --> 01:04:15,880
which is like how do we generally think about agent-native institutions

1182
01:04:15,880 --> 01:04:17,840
for achieving results, which obviously we care about.

1183
01:04:17,840 --> 01:04:20,860
But some of these other things, which is preserving or creating

1184
01:04:20,860 --> 01:04:23,600
new processes for shared understanding, attention, and trust, something

1185
01:04:23,600 --> 01:04:27,020
that is fundamental to open source, that can support collective

1186
01:04:27,020 --> 01:04:28,200
work in open communities broadly.

1187
01:04:28,660 --> 01:04:30,320
So I think I'll just end with that question.

1188
01:04:30,740 --> 01:04:33,940
And yeah, I'll take some questions, but also happy to pick up the

1189
01:04:33,940 --> 01:04:35,580
conversation separately as well.

1190
01:04:36,100 --> 01:04:36,400
Thanks.

1191
01:04:48,940 --> 01:04:49,740
Yeah.

1192
01:04:51,920 --> 01:04:57,300
You mentioned about the problem of open

1193
01:04:57,300 --> 01:05:02,340
source repositories dealing with way too

1194
01:05:02,340 --> 01:05:07,460
many contributors, and maybe you can like scale agentic review,

1195
01:05:07,640 --> 01:05:09,520
but then it still can be very costly.

1196
01:05:10,920 --> 01:05:16,660
And you mentioned that maybe like a limiting number of tokens

1197
01:05:16,660 --> 01:05:20,880
could be a solution, or how many contributions one agent could do

1198
01:05:20,880 --> 01:05:21,620
could be a solution.

1199
01:05:22,320 --> 01:05:26,420
Have you seen other like mitigation factors that have been useful

1200
01:05:26,420 --> 01:05:27,800
for open source?

1201
01:05:27,980 --> 01:05:30,700
Because there are open source projects that are not doing what TL

1202
01:05:30,700 --> 01:05:32,700
draw did, but still accepting lots of contributions.

1203
01:05:33,560 --> 01:05:35,640
Have you seen good examples that are inspiring?

1204
01:05:35,640 --> 01:05:39,000
Yeah, I think one interesting example that's happening live right

1205
01:05:39,000 --> 01:05:43,660
now is I share that post from Mitchell Hashimoto, but he has this

1206
01:05:43,660 --> 01:05:46,720
project called Vouch, where they launched it earlier this year.

1207
01:05:47,720 --> 01:05:50,820
And it's nothing too fancy, but it basically is built off of this

1208
01:05:50,820 --> 01:05:53,920
notion of trust lists that can be maintained per repo, and then

1209
01:05:53,920 --> 01:05:58,580
the ability to publicly basically vouch for or denounce GitHub identity.

1210
01:05:58,840 --> 01:06:01,200
And it was controversial at the time, right, because people said

1211
01:06:01,200 --> 01:06:03,740
that, oh, you can denounce someone, this is just going to be used

1212
01:06:03,740 --> 01:06:04,240
for gatekeeping.

1213
01:06:04,240 --> 01:06:07,280
But in practice, I think what it's been used for is to experiment

1214
01:06:07,280 --> 01:06:12,060
with, okay, we, as the inner circle of maintainers, generally have

1215
01:06:12,060 --> 01:06:15,100
a vibe of like who's been useful in contributing stuff.

1216
01:06:15,200 --> 01:06:17,840
So why don't we just make that explicit? It's happening already

1217
01:06:17,840 --> 01:06:18,500
implicitly.

1218
01:06:18,540 --> 01:06:21,500
So it's kind of a lie to say that there isn't already this trust

1219
01:06:21,500 --> 01:06:24,940
system. So why not just make it super explicit so everyone can transparently

1220
01:06:24,940 --> 01:06:27,760
see who is being vouched for and who is being publicly denounced.

1221
01:06:27,760 --> 01:06:30,760
And then I think the idea or the hope for that project is that people

1222
01:06:30,760 --> 01:06:32,680
would be able to share trust lists with one another.

1223
01:06:32,820 --> 01:06:36,940
So if I see that like some sus guy shows up and just gave me a thousand

1224
01:06:36,940 --> 01:06:40,340
drive by PRs and then left, well, maybe someone else wants to know

1225
01:06:40,340 --> 01:06:40,740
about that.

1226
01:06:40,880 --> 01:06:44,320
So I think there's definitely room to kind of learn and maybe extend

1227
01:06:44,320 --> 01:06:45,140
some of those primitives.

1228
01:06:46,160 --> 01:06:50,300
So I don't think any of these, the areas of investigation should

1229
01:06:50,300 --> 01:06:52,440
be treated as just like let's start from a blank slate.

1230
01:06:52,600 --> 01:06:54,420
Because I think there's a lot of good work happening already.

1231
01:06:54,640 --> 01:06:57,960
I think the main question is, you know, how does that get incorporated

1232
01:06:57,960 --> 01:06:58,440
writ large?

1233
01:06:58,660 --> 01:07:01,760
Is it a one size fits all solution or should we be thinking about

1234
01:07:01,760 --> 01:07:03,020
other types of mechanisms too?

1235
01:07:09,260 --> 01:07:09,780
Cool.

1236
01:07:11,660 --> 01:07:15,120
Yeah, I think if that is it, then, oh, do you have a question?

1237
01:07:15,540 --> 01:07:16,060
Okay.

1238
01:07:21,560 --> 01:07:26,700
How have you seen the kind of the long term maintainability side

1239
01:07:26,700 --> 01:07:28,620
of open source projects change?

1240
01:07:28,620 --> 01:07:32,180
Especially when it comes down to maybe like resource management

1241
01:07:32,180 --> 01:07:36,940
and, you know, bigger, for example, bigger open source projects

1242
01:07:36,940 --> 01:07:40,100
have budgets and they can hire people, right?

1243
01:07:40,180 --> 01:07:43,080
Or they're sponsored by enterprises.

1244
01:07:44,060 --> 01:07:47,560
Yeah, in this context, how has that changed?

1245
01:07:49,160 --> 01:07:54,180
Maybe like two initial responses or initial reactions.

1246
01:07:54,180 --> 01:07:56,700
I think one is actually related to a conversation I was having with

1247
01:07:56,700 --> 01:08:00,720
Max earlier, which is I suspect that we'll probably just need to

1248
01:08:00,720 --> 01:08:03,540
be open minded about what contributing to open source means.

1249
01:08:04,640 --> 01:08:07,840
Where I think in a previous era, code was the thing, right?

1250
01:08:07,940 --> 01:08:11,000
But I think it's pretty clear that code may not be the most valuable

1251
01:08:11,000 --> 01:08:11,860
thing that you can contribute.

1252
01:08:13,260 --> 01:08:15,860
Arguably, all along, the most valuable thing you can contribute

1253
01:08:15,860 --> 01:08:17,260
probably wasn't code to begin with.

1254
01:08:17,340 --> 01:08:20,600
It was just like a very easy thing to like kind of gravitate towards.

1255
01:08:20,600 --> 01:08:24,360
But really, it was about like creative ideas, your ability to build

1256
01:08:24,360 --> 01:08:27,380
trust within an ecosystem, and your ability to build cohesion with

1257
01:08:27,380 --> 01:08:29,020
a group of people across the internet.

1258
01:08:29,640 --> 01:08:33,560
So I think the nature of contributing and what we choose to value

1259
01:08:33,560 --> 01:08:36,900
probably needs to change as well, where code is no longer the valuable

1260
01:08:36,900 --> 01:08:37,160
thing.

1261
01:08:37,260 --> 01:08:39,000
That's like almost trivially automatable.

1262
01:08:39,000 --> 01:08:40,740
So that's like one reaction.

1263
01:08:41,100 --> 01:08:45,880
In terms of like how this stuff gets funded, I don't really know.

1264
01:08:46,140 --> 01:08:51,260
But I think we already have a lot of funding sources that

1265
01:08:51,260 --> 01:08:56,300
have vested interest in having like solid building blocks that they

1266
01:08:56,300 --> 01:08:57,000
don't have to reinvent.

1267
01:08:57,000 --> 01:09:01,120
And I think this remains true even with agents, where, yeah, you

1268
01:09:01,120 --> 01:09:02,720
could totally rebuild these building blocks.

1269
01:09:03,140 --> 01:09:06,100
Or you could just glue an existing building block that works really

1270
01:09:06,100 --> 01:09:06,400
well.

1271
01:09:06,580 --> 01:09:09,340
And I suspect that we'll continue to have that preference going

1272
01:09:09,340 --> 01:09:09,660
forward.

1273
01:09:09,660 --> 01:09:13,660
And I think, I don't know how it gets funded, but I imagine it'll

1274
01:09:13,660 --> 01:09:18,280
be still involving some of the existing corporate sponsors and some

1275
01:09:18,280 --> 01:09:19,100
of the existing mechanisms.

1276
01:09:19,900 --> 01:09:22,480
But I think what's interesting is that if you assume that a lot

1277
01:09:22,480 --> 01:09:26,900
of people have agents and token budgets, how do they allocate, you

1278
01:09:26,900 --> 01:09:30,840
know, their money and or now tokens and intelligence towards building

1279
01:09:30,840 --> 01:09:31,960
these shared building blocks as well.

1280
01:09:31,960 --> 01:09:34,840
So I think that's probably like the open greenfield thing where

1281
01:09:34,840 --> 01:09:37,200
that might change the way that we think about how it's quote unquote

1282
01:09:37,200 --> 01:09:38,160
funded and sustained.

1283
01:09:38,820 --> 01:09:41,500
And I mean, I think part of this talk series is figuring that out.

1284
01:09:41,640 --> 01:09:45,140
I don't know the details, but I suspect something new will emerge

1285
01:09:45,140 --> 01:09:45,580
there as well.

1286
01:09:49,400 --> 01:09:49,800
Cool.

1287
01:09:50,480 --> 01:09:51,820
That's it for me.

1288
01:09:51,880 --> 01:09:53,720
And I'll pass it off to Max.

1289
01:09:54,300 --> 01:09:54,360
Yeah.

1290
01:09:55,460 --> 01:09:56,020
Thank you.

1291
01:09:57,060 --> 01:09:57,860
Thank you.

1292
01:09:57,900 --> 01:09:58,200
you.

1293
01:09:58,660 --> 01:09:58,780
Thank you.

1294
01:09:59,180 --> 01:09:59,360
Thank

1295
01:09:59,600 --> 01:09:59,680
you.

1296
01:10:14,640 --> 01:10:18,360
Hey, thanks for having me. I'm really excited to be talking about

1297
01:10:18,360 --> 01:10:21,180
multi-agent systems and alignment. think it's a really important

1298
01:10:21,180 --> 01:10:22,100
topic right now.

1299
01:10:23,640 --> 01:10:29,000
So I'm going to be sharing some kind of like worked examples from

1300
01:10:29,000 --> 01:10:34,400
a real life multi-agent system that I maintain that

1301
01:10:34,400 --> 01:10:38,080
happens to be an MMO called RuneScape.

1302
01:10:38,080 --> 01:10:44,700
So just

1303
01:10:44,700 --> 01:10:50,060
to introduce myself really quickly, my name is Max Bittker. I work

1304
01:10:50,060 --> 01:10:55,060
on a project called Websim, which is like a platform where a bunch

1305
01:10:55,060 --> 01:10:59,620
of people work together to build games and build really complicated

1306
01:10:59,620 --> 01:11:00,660
multiplayer projects.

1307
01:11:00,660 --> 01:11:04,980
But I'm going to be talking today about another project of mine

1308
01:11:04,980 --> 01:11:10,260
called RS SDK, which is the RuneScape SDK, which is basically

1309
01:11:10,260 --> 01:11:15,200
code bindings for any person who wants to write scripts, but it

1310
01:11:15,200 --> 01:11:19,140
turns out language models love writing scripts to control a RuneScape

1311
01:11:19,140 --> 01:11:23,520
character and observe its surrounding and act on goals.

1312
01:11:23,520 --> 01:11:28,360
And maybe even really long horizon or multi-agent goals like competing

1313
01:11:28,360 --> 01:11:32,800
for a high score or trading with other agents all inside kind of

1314
01:11:32,800 --> 01:11:35,080
like an emulated open source RuneScape server.

1315
01:11:36,400 --> 01:11:41,020
So part of the inspiration for this is that I was at one point like

1316
01:11:41,020 --> 01:11:44,440
a kid who loved RuneScape and at a certain point I figured out that

1317
01:11:44,440 --> 01:11:49,080
you could repeat all of the repetitive actions via scripts.

1318
01:11:49,080 --> 01:11:53,140
And you didn't have to like mine 10,000 logs to get the goal you

1319
01:11:53,140 --> 01:11:56,540
wanted. You could set up an auto clicker overnight and do it.

1320
01:11:56,640 --> 01:12:00,040
And I thought that that game loop was so much more fun than RuneScape

1321
01:12:00,040 --> 01:12:02,560
itself. And that kind of led me to programming and everything else.

1322
01:12:02,800 --> 01:12:06,580
And I've always wanted more people to experience that as a game

1323
01:12:06,580 --> 01:12:09,740
itself, like the metagame of automating the game.

1324
01:12:10,640 --> 01:12:13,440
Unfortunately, it's like against the rules, but I just think it

1325
01:12:13,440 --> 01:12:15,360
shouldn't be. Everybody should just be able to do it equally.

1326
01:12:16,040 --> 01:12:21,880
And so coding agents really work well for this. This part

1327
01:12:21,880 --> 01:12:27,020
of the goal is to get other

1328
01:12:27,020 --> 01:12:27,960
people to have that experience.

1329
01:12:28,220 --> 01:12:31,780
And so this project has been popular online, like thousands of people

1330
01:12:31,780 --> 01:12:34,720
have tried it and used a coding agent for the first time to do these

1331
01:12:34,720 --> 01:12:35,740
long horizon goals.

1332
01:12:39,040 --> 01:12:42,860
I'll show the live version because it's cool to watch.

1333
01:12:54,260 --> 01:12:58,400
Maybe not right now, but basically at any given time, there's hundreds

1334
01:12:58,400 --> 01:13:02,000
and hundreds of different people's agents.

1335
01:13:02,200 --> 01:13:05,340
Some people run one agent, some people run a swarm of agents, and

1336
01:13:05,340 --> 01:13:07,500
they're all interacting on the server and pursuing goals.

1337
01:13:07,500 --> 01:13:12,300
So this is kind of like a heat map of seven days of activity.

1338
01:13:12,560 --> 01:13:16,940
And each of these yellow dots on the map is being controlled by

1339
01:13:16,940 --> 01:13:18,140
a coding agent somewhere.

1340
01:13:19,800 --> 01:13:24,700
There also is a version of this project that is an eval to measure

1341
01:13:24,700 --> 01:13:30,040
the kind of like problem solving ability of different coding models.

1342
01:13:30,800 --> 01:13:34,540
You can kind of see this is a Pareto curve, and you can see an outlier

1343
01:13:34,540 --> 01:13:38,660
on the top left is GPT-6 Astra, one of the newest models on here,

1344
01:13:38,800 --> 01:13:39,300
which is...

1345
01:13:39,300 --> 01:13:43,480
This is a log scale, the way, so GPT-Astra is almost 10 times better

1346
01:13:43,480 --> 01:13:46,380
than some of the models over here that are cheaper, even though

1347
01:13:46,380 --> 01:13:48,060
it's also much more expensive.

1348
01:13:49,700 --> 01:13:52,780
And this has actually turned out to be a really useful evaluation

1349
01:13:52,780 --> 01:13:57,600
for just understanding how well models can deal with long horizon

1350
01:13:57,600 --> 01:14:00,300
tasks and kind of goal following and optimization.

1351
01:14:01,320 --> 01:14:06,160
And then also, you know, this is like classic scary graph of every

1352
01:14:06,160 --> 01:14:10,860
AI thing, but x-axis here is just release of the model.

1353
01:14:11,260 --> 01:14:15,040
And in the nine months that this benchmark has been out, there's

1354
01:14:15,040 --> 01:14:19,380
been 10x improvement in how well that they score on it.

1355
01:14:19,960 --> 01:14:21,720
And it's going up.

1356
01:14:21,820 --> 01:14:23,460
It's actually, I think, of saturating.

1357
01:14:23,460 --> 01:14:27,000
And so I've been looking at not just single agent tasks, but at

1358
01:14:27,000 --> 01:14:31,960
multi-agent tasks and how well can they work together on either

1359
01:14:31,960 --> 01:14:37,180
competitive or cooperative, like kind of market-based tasks.

1360
01:14:37,460 --> 01:14:42,740
And so setting up agents into scenarios where they

1361
01:14:42,740 --> 01:14:46,380
need to trade and collaborate in order to accomplish their goals.

1362
01:14:49,800 --> 01:14:51,500
This is a...

1363
01:14:51,500 --> 01:14:52,140
It's okay.

1364
01:14:52,880 --> 01:14:56,920
This is a video of just like a grid of 20 agents all working at

1365
01:14:56,920 --> 01:15:02,400
the same time to talk

1366
01:15:02,400 --> 01:15:05,760
to each other and to trade with each other in order to...

1367
01:15:05,760 --> 01:15:08,720
They're each optimizing for their own individual income.

1368
01:15:08,900 --> 01:15:12,260
But the way they have to accomplish that is by talking to each other

1369
01:15:12,260 --> 01:15:15,540
and setting up trades and basically finding prices.

1370
01:15:16,100 --> 01:15:19,760
And there's also even kind of like exploitation because they might

1371
01:15:19,760 --> 01:15:21,540
ask for like loans from other people.

1372
01:15:21,540 --> 01:15:26,600
They're like, hey, I can pay you, you know, 2,000 gold

1373
01:15:26,600 --> 01:15:29,780
for that item, but I need the item first in order to afford it.

1374
01:15:30,880 --> 01:15:34,620
And so, yeah, I think I'd have to get my...

1375
01:15:34,620 --> 01:15:35,260
Unfortunately.

1376
01:15:36,080 --> 01:15:40,000
And so, basically, this has been a really interesting kind of like

1377
01:15:40,000 --> 01:15:45,020
experimental playground for determining what kinds of misalignment

1378
01:15:45,020 --> 01:15:51,180
and group misalignment scenarios happen and what factors are kind

1379
01:15:51,180 --> 01:15:53,200
of like push them to happen more or less.

1380
01:15:54,560 --> 01:15:56,780
And also how agents deal with scenarios.

1381
01:15:56,780 --> 01:15:56,920
And

1382
01:15:56,920 --> 01:16:01,440
Like if they've gotten scammed, how do they like tell all the other

1383
01:16:01,440 --> 01:16:03,420
agents what happened?

1384
01:16:03,420 --> 01:16:07,360
And in some cases, they've threatened to tell everybody and then

1385
01:16:07,360 --> 01:16:08,160
gotten their money back.

1386
01:16:14,160 --> 01:16:15,860
And then I guess that...

1387
01:16:15,860 --> 01:16:18,820
So, I'm doing a lot of experiments with this kind of test bed.

1388
01:16:20,460 --> 01:16:25,500
And one of my kind of like zoomed out hunches is

1389
01:16:25,500 --> 01:16:30,840
that the way that agents are trained is that they do experience

1390
01:16:30,840 --> 01:16:35,260
like many millions of hours of task goal following.

1391
01:16:35,480 --> 01:16:38,960
But it's almost always single agent goal following with single agent

1392
01:16:38,960 --> 01:16:39,580
rewards.

1393
01:16:39,580 --> 01:16:43,640
If you look at the biological world and like our own evolution,

1394
01:16:44,000 --> 01:16:49,180
we are a product of individual selection, but we're also the process

1395
01:16:49,180 --> 01:16:50,860
of group and community selection.

1396
01:16:51,340 --> 01:16:56,480
And many factors that explain the

1397
01:16:56,480 --> 01:17:00,880
way that animals and plants behave is better explained by group

1398
01:17:00,880 --> 01:17:02,600
selection than only individual selection.

1399
01:17:02,600 --> 01:17:07,380
And so, I'm very curious about ideas about factoring in like negative

1400
01:17:07,380 --> 01:17:11,260
externalities of your actions into training processes.

1401
01:17:11,620 --> 01:17:16,640
And so, basically giving agents many examples of

1402
01:17:16,640 --> 01:17:21,300
being in a world where they need to benefit the people around them

1403
01:17:21,300 --> 01:17:22,680
in order to succeed at their goal.

1404
01:17:22,840 --> 01:17:25,380
Much like our evolutionary past.

1405
01:17:26,200 --> 01:17:27,660
So, really excited.

1406
01:17:27,940 --> 01:17:28,500
I can definitely...

1407
01:17:28,500 --> 01:17:31,080
I've got cool videos and transcripts that people want to see some

1408
01:17:31,080 --> 01:17:33,340
examples of these agent scenarios.

1409
01:17:33,900 --> 01:17:35,860
And yeah, thanks.

1410
01:17:36,140 --> 01:17:36,640
Nice to be here.

1411
01:17:39,860 --> 01:17:40,820
Thank you.

1412
01:17:41,000 --> 01:17:46,000
[Audience question inaudible; the speaker repeats it below.]

1413
01:17:58,020 --> 01:18:02,600
Yeah, so he was asking how the long running agent swarms, like how

1414
01:18:02,600 --> 01:18:03,360
long do they run and

1415
01:18:03,360 --> 01:18:04,780
how does that work to keep them running.

1416
01:18:05,380 --> 01:18:08,060
So there's basically two examples.

1417
01:18:08,060 --> 01:18:12,420
One is that I just run a server that's always on and agents are

1418
01:18:12,420 --> 01:18:13,720
constantly just booting

1419
01:18:13,720 --> 01:18:14,860
up, connecting to it.

1420
01:18:15,340 --> 01:18:18,160
And that's each individual who runs the agent makes their own decision.

1421
01:18:18,460 --> 01:18:21,180
Some people play kind of like interactively.

1422
01:18:21,420 --> 01:18:24,640
Some people set it up in a server to be on a cron job and run all

1423
01:18:24,640 --> 01:18:25,040
the time.

1424
01:18:26,480 --> 01:18:30,520
And then inside of my kind of like controlled scenarios, those tend

1425
01:18:30,520 --> 01:18:31,660
to be between like 30

1426
01:18:31,660 --> 01:18:32,480
and 90 minutes.

1427
01:18:32,980 --> 01:18:35,640
And that just fits inside of one agent run.

1428
01:18:35,640 --> 01:18:40,880
And so it's just setting up 10 sandboxes or

1429
01:18:40,880 --> 01:18:43,480
100 sandboxes that are each running a coding

1430
01:18:43,480 --> 01:18:43,840
agent.

1431
01:18:49,660 --> 01:18:52,240
I'm sure there are many cases of this.

1432
01:18:52,400 --> 01:18:56,440
But just out of curiosity, relative to your expectations when you

1433
01:18:56,440 --> 01:18:57,380
first started doing these

1434
01:18:57,380 --> 01:19:02,080
experiments, what has been the most surprising emergent behavior

1435
01:19:02,080 --> 01:19:03,540
that you've seen either

1436
01:19:03,540 --> 01:19:07,360
on the live server or in your own controlled simulations, if there

1437
01:19:07,360 --> 01:19:08,120
is one that stands out

1438
01:19:08,120 --> 01:19:08,400
to you?

1439
01:19:09,520 --> 01:19:09,860
Yeah.

1440
01:19:10,040 --> 01:19:14,440
So I was really curious when I started how the...

1441
01:19:14,440 --> 01:19:16,560
Because RuneScape is a really boring game.

1442
01:19:16,800 --> 01:19:20,980
But the cool part is all about like the economy and the prices and

1443
01:19:20,980 --> 01:19:21,580
different goals.

1444
01:19:21,580 --> 01:19:24,880
And all your kind of like greatest moments in RuneScape are because

1445
01:19:24,880 --> 01:19:26,460
you saved up or you

1446
01:19:26,460 --> 01:19:29,080
found some kind of like money-making trick.

1447
01:19:31,040 --> 01:19:32,380
And so I was really curious.

1448
01:19:32,620 --> 01:19:36,220
okay, if RuneScape is so much of a labor resource processing economy,

1449
01:19:36,300 --> 01:19:36,620
so how

1450
01:19:36,620 --> 01:19:38,820
does that work if labor is very cheap?

1451
01:19:40,340 --> 01:19:42,600
And it's been really cool watching this play out.

1452
01:19:42,600 --> 01:19:46,580
One thing is that like trade, basically like inflation is super

1453
01:19:46,580 --> 01:19:47,480
high on the server because

1454
01:19:47,480 --> 01:19:52,500
people don't have a lot of demand for like the cost

1455
01:19:52,500 --> 01:19:53,620
of coordinating with another agent

1456
01:19:53,620 --> 01:19:56,400
compared to just leaving it overnight to do the work for you and

1457
01:19:56,400 --> 01:19:57,160
go get the resource.

1458
01:19:57,820 --> 01:19:59,720
It de-incentivizes trade.

1459
01:19:59,720 --> 01:19:59,740
of

1460
01:19:59,740 --> 01:19:59,880
of a trade.

1461
01:19:59,880 --> 01:19:59,960
It's little

1462
01:20:00,000 --> 01:20:02,460
And then the other thing that's been interesting is that certain

1463
01:20:02,460 --> 01:20:06,980
resources that don't just scale linearly with labor but instead

1464
01:20:06,980 --> 01:20:08,620
have some kind of natural scarcity.

1465
01:20:08,880 --> 01:20:13,420
So an example in the game is Rune Ore only has a single spawn location.

1466
01:20:13,640 --> 01:20:17,760
And if you have 100 people, only one person is going to get it.

1467
01:20:17,880 --> 01:20:21,080
So these are the resources that have become scarce and kind of valuable.

1468
01:20:21,500 --> 01:20:25,660
And so people make more and more advanced swarms just to compete

1469
01:20:25,660 --> 01:20:28,520
for kind of these same scarce resources.

1470
01:20:28,520 --> 01:20:32,540
And so that's been an interesting thing to watch play out.

1471
01:20:36,180 --> 01:20:37,200
Any other questions?

1472
01:20:39,820 --> 01:20:43,840
Yeah, this is more just the thought, but it made me think seeing

1473
01:20:43,840 --> 01:20:45,140
the RuneScape example.

1474
01:20:45,300 --> 01:20:48,300
One of the things that I've been thinking about is just like humans,

1475
01:20:48,300 --> 01:20:51,560
as humans, we have bodies and we're like geographically constrained.

1476
01:20:51,560 --> 01:20:54,780
And it's interesting that in this RuneScape example, I feel like

1477
01:20:54,780 --> 01:20:59,880
the agents kind of have the same sort of instantiation.

1478
01:21:00,020 --> 01:21:02,260
And it could be that there's like all kinds of interesting things

1479
01:21:02,260 --> 01:21:03,400
from just like the geography.

1480
01:21:03,700 --> 01:21:06,180
You know, it was like even the question of like, hey, why did these

1481
01:21:06,180 --> 01:21:08,420
people share these values over here versus these ones?

1482
01:21:08,580 --> 01:21:11,980
It's like partially determined by geographical constraints.

1483
01:21:11,980 --> 01:21:16,000
And it is kind of, yeah, it just makes me wonder like if agents

1484
01:21:16,000 --> 01:21:19,820
behave differently if you embody them and you like also constrain

1485
01:21:19,820 --> 01:21:20,500
their geography.

1486
01:21:21,840 --> 01:21:22,860
Yeah, definitely.

1487
01:21:23,140 --> 01:21:24,080
I have no, yeah.

1488
01:21:25,080 --> 01:21:28,840
I think of it kind of as like a mini robotics environment because

1489
01:21:28,840 --> 01:21:33,860
you're taking a text-based coding agent, but

1490
01:21:33,860 --> 01:21:38,400
actually all of its actions are expressed through a kind of like

1491
01:21:38,400 --> 01:21:41,160
a thin nozzle, which is that you can only observe what's around

1492
01:21:41,160 --> 01:21:42,980
you and you can only act on what's around you.

1493
01:21:42,980 --> 01:21:48,220
And so a cool thing about this is that it has big implications for

1494
01:21:48,220 --> 01:21:51,760
multi-agent in many kinds of software tasks.

1495
01:21:52,580 --> 01:21:57,480
It's undetermined if having more individual actors is actually helpful

1496
01:21:57,480 --> 01:22:02,220
versus just like one long running task in kind of like a game or

1497
01:22:02,220 --> 01:22:03,100
a robotics task.

1498
01:22:03,420 --> 01:22:05,720
More people collaborating just means more bodies.

1499
01:22:05,880 --> 01:22:07,400
And so that's been interesting to watch too.

1500
01:22:07,780 --> 01:22:10,240
But it's not always one-to-one.

1501
01:22:10,240 --> 01:22:15,460
You do sometimes see one coding agent controlling a hundred actors

1502
01:22:15,460 --> 01:22:17,900
in the game by kind of like multiplexing.

1503
01:22:25,400 --> 01:22:25,880
Thanks.

1504
01:22:26,340 --> 01:22:28,620
I was curious about like the limiting factors.

1505
01:22:29,940 --> 01:22:32,780
Because right there, the slide where you were saying like how we

1506
01:22:32,780 --> 01:22:37,920
are in biological, you know, our human bodies and

1507
01:22:37,920 --> 01:22:38,620
our ecosystem.

1508
01:22:38,620 --> 01:22:41,860
There's so many limiting factors here in terms of like the energy

1509
01:22:41,860 --> 01:22:45,140
output we have per day, our attention, our focus.

1510
01:22:46,360 --> 01:22:51,860
And with agents, it seems like there are many less limiting factors.

1511
01:22:52,280 --> 01:22:55,920
Sort of like if you have infinite budget, then you can waste all

1512
01:22:55,920 --> 01:22:56,760
the tokens you want.

1513
01:22:58,460 --> 01:23:02,840
Do you see any way in which like new types of limiting factors could

1514
01:23:02,840 --> 01:23:07,660
be applied to agents or to systems that could begin to put boundaries

1515
01:23:07,660 --> 01:23:08,380
on these systems?

1516
01:23:12,000 --> 01:23:14,900
I think that it becomes really obvious when the systems start playing

1517
01:23:14,900 --> 01:23:15,800
out in practice.

1518
01:23:16,260 --> 01:23:20,840
so I didn't come into like I was just going to run a stock RuneScape

1519
01:23:20,840 --> 01:23:21,180
server.

1520
01:23:22,640 --> 01:23:26,320
And once you start having all of these agents running on it, you

1521
01:23:26,320 --> 01:23:28,380
just kind of like see what the breaking points are.

1522
01:23:28,620 --> 01:23:32,060
And so for instance, there was this limiting factor of like RuneScape

1523
01:23:32,060 --> 01:23:34,740
server can only hold so many people in it at once.

1524
01:23:34,940 --> 01:23:39,040
And so in order to just keep it accessible, I started limiting per

1525
01:23:39,040 --> 01:23:39,780
IP address.

1526
01:23:39,780 --> 01:23:44,900
And so it's interesting that I think people often assume

1527
01:23:44,900 --> 01:23:50,300
when it comes to AI, they use like infinity as the multiplier.

1528
01:23:50,300 --> 01:23:54,440
But I think it's actually more just like a million or something

1529
01:23:54,440 --> 01:23:56,340
or like or even more like a thousand.

1530
01:23:56,600 --> 01:24:01,960
And so things you apply that new form of

1531
01:24:01,960 --> 01:24:04,560
energy into the system and things just rearrange.

1532
01:24:04,800 --> 01:24:06,120
They don't like explode.

1533
01:24:06,120 --> 01:24:10,800
And so for example of this server, was like, okay, yeah, we're going

1534
01:24:10,800 --> 01:24:15,980
to put on like each IP address can only connect 200 bots.

1535
01:24:16,160 --> 01:24:20,460
And then there's even also limits where in order to run a bot, you

1536
01:24:20,460 --> 01:24:21,980
kind of need to be running a web browser.

1537
01:24:22,160 --> 01:24:24,460
So you're also limited by how much RAM you have, not to mention

1538
01:24:24,460 --> 01:24:25,080
like tokens.

1539
01:24:25,520 --> 01:24:29,000
And so the numbers get weird, but they don't go to infinity.

1540
01:24:29,240 --> 01:24:31,520
So there's always some kind of balance to be struck.

1541
01:24:37,320 --> 01:24:40,000
I just had a quick follow-up thought slash question regarding the

1542
01:24:40,000 --> 01:24:40,680
rune ore thing.

1543
01:24:41,240 --> 01:24:44,380
I don't know if you've looked at this, but I feel like it'd be interesting

1544
01:24:44,380 --> 01:24:48,780
to like the whole notion of like comparative advantage, right?

1545
01:24:48,980 --> 01:24:51,300
In like economics where it's just like, yeah, they can just do it

1546
01:24:51,300 --> 01:24:53,240
themselves, but there's still an opportunity cost.

1547
01:24:53,240 --> 01:24:56,380
And like the rune ore example made me wonder, it's like, well, if

1548
01:24:56,380 --> 01:24:58,840
the rune ore is the thing that's scarce, why wouldn't they want

1549
01:24:58,840 --> 01:25:01,260
to like dedicate all their resources to like beating everyone for

1550
01:25:01,260 --> 01:25:03,680
the rune ore and just being like, hey, other agent over there.

1551
01:25:03,960 --> 01:25:06,900
Like you can do this for me because I need my rune ore.

1552
01:25:07,360 --> 01:25:11,420
And I'm kind of curious if, you know, we would expect that to end

1553
01:25:11,420 --> 01:25:14,280
up happening or if there's already empirical evidence that it doesn't.

1554
01:25:14,400 --> 01:25:16,840
Because I feel like that has probably implications for like how

1555
01:25:16,840 --> 01:25:19,180
people think about, you know, the real world too.

1556
01:25:19,480 --> 01:25:22,640
Like whether comparative advantage will actually continue to hold.

1557
01:25:22,820 --> 01:25:25,200
So I don't know if you've thought about that or have seen anything

1558
01:25:25,200 --> 01:25:25,920
in that regard.

1559
01:25:26,780 --> 01:25:27,220
Yeah.

1560
01:25:27,560 --> 01:25:31,480
So an example of the rune ore where this is like a very scarce resource.

1561
01:25:31,620 --> 01:25:34,360
And if you want to make money, you kind of have to like go after

1562
01:25:34,360 --> 01:25:37,120
some of these things that can't just be, people actually want to

1563
01:25:37,120 --> 01:25:38,780
buy it from you because they can't get it themselves.

1564
01:25:41,200 --> 01:25:45,400
And absolutely right now we see comparative advantage because if

1565
01:25:45,400 --> 01:25:48,760
you're trying to go after one of these scarce resources, you don't

1566
01:25:48,760 --> 01:25:52,340
just let the agent do it itself.

1567
01:25:52,340 --> 01:25:55,860
That's where you start like giving the agent more resources, you

1568
01:25:55,860 --> 01:25:56,840
suggest strategies.

1569
01:25:57,240 --> 01:26:00,960
And so the people who are successfully getting access to these scarce

1570
01:26:00,960 --> 01:26:04,060
resources on the server are the people who also currently are putting

1571
01:26:04,060 --> 01:26:07,840
in the most dollars and the most human ingenuity.

1572
01:26:08,140 --> 01:26:14,140
And so that's currently how it's playing out is that those

1573
01:26:14,140 --> 01:26:14,980
are the people who are winning.

1574
01:26:14,980 --> 01:26:19,360
Maybe there could be somebody who just purely puts in like tons

1575
01:26:19,360 --> 01:26:22,380
of Astra credits and gives it a goal like this and they have a good

1576
01:26:22,380 --> 01:26:23,860
outcome too.

1577
01:26:24,220 --> 01:26:28,880
But right now it seems like Centaur kind of, you know, combinations

1578
01:26:28,880 --> 01:26:30,640
are the people who are the most successful.

1579
01:26:42,660 --> 01:26:46,040
I'm curious if you observe any, like this, this curve is really

1580
01:26:46,040 --> 01:26:46,760
interesting.

1581
01:26:47,700 --> 01:26:52,120
I'm curious if like, what is the behavior that changes as you move

1582
01:26:52,120 --> 01:26:55,540
up the curve that creates such a massive difference in the XP?

1583
01:26:55,760 --> 01:26:57,060
Like what is Astra doing?

1584
01:26:57,740 --> 01:27:00,120
Like I would, I would kind of think that a game like RuneScape would

1585
01:27:00,120 --> 01:27:01,360
be saturated at some point.

1586
01:27:01,520 --> 01:27:05,000
And so what did, what is Astra doing that is so much more effective

1587
01:27:05,000 --> 01:27:07,060
than what the other models are doing?

1588
01:27:10,220 --> 01:27:11,600
So that's a great question.

1589
01:27:14,840 --> 01:27:20,340
At the low end of the curve of just like being like

1590
01:27:20,340 --> 01:27:23,360
better, you know, this is where we see like Sonnet 4.5.

1591
01:27:24,700 --> 01:27:28,540
Just navigating the game is really hard because the game actually

1592
01:27:28,540 --> 01:27:31,240
has a surprising amount of weird stuff in it.

1593
01:27:31,500 --> 01:27:33,420
And you're only given 30 minutes wall clock.

1594
01:27:34,020 --> 01:27:38,360
so Sonnet here, like it was supposed to go train crafting for this

1595
01:27:38,360 --> 01:27:38,760
task.

1596
01:27:39,020 --> 01:27:41,840
And in 30 minutes, it just probably like got stuck on a door.

1597
01:27:42,320 --> 01:27:45,740
It tried to go find something over here, but then it needed this.

1598
01:27:45,840 --> 01:27:48,160
And like it couldn't just untangle the web.

1599
01:27:48,160 --> 01:27:51,300
And so at the low end, you see just too much complexity and they

1600
01:27:51,300 --> 01:27:51,800
get confused.

1601
01:27:54,240 --> 01:27:58,860
In the mid range, a lot of it has to do with the difference between

1602
01:27:58,860 --> 01:28:01,560
doing the task and optimizing the task.

1603
01:28:01,880 --> 01:28:06,020
And so sometimes you'll see models where they will accomplish a

1604
01:28:06,020 --> 01:28:06,900
task like fishing.

1605
01:28:06,900 --> 01:28:10,280
And then they'll kind of just chill for like the next 15 minutes.

1606
01:28:10,460 --> 01:28:14,100
And they'll keep doing the same loop, but they won't kind of have

1607
01:28:14,100 --> 01:28:17,660
this like feeling of like, I got to figure out how to catch these

1608
01:28:17,660 --> 01:28:18,340
fish faster.

1609
01:28:18,540 --> 01:28:20,720
I got to go try different fish.

1610
01:28:20,820 --> 01:28:21,860
I got to go try different stuff.

1611
01:28:22,120 --> 01:28:26,140
And so the benchmark is really set up to reward.

1612
01:28:26,260 --> 01:28:31,220
It rewards your peak XP rate within any 15 second window.

1613
01:28:31,220 --> 01:28:34,080
So once you've kind of found one strategy, you're supposed to keep

1614
01:28:34,080 --> 01:28:35,280
looking for better strategies.

1615
01:28:35,820 --> 01:28:39,300
And you're not supposed to just look for like slightly better strategies.

1616
01:28:39,300 --> 01:28:44,120
Like I'm going to keep catching shrimp, but I'm going to like click

1617
01:28:44,120 --> 01:28:44,500
differently.

1618
01:28:44,820 --> 01:28:47,580
You're supposed to go explore and try more complicated strategies.

1619
01:28:47,940 --> 01:28:52,880
And so at the top of the skill expression, you see agents who kind

1620
01:28:52,880 --> 01:28:57,960
of like reason without acting about what strategies will

1621
01:28:57,960 --> 01:28:58,440
be good.

1622
01:28:58,440 --> 01:29:01,340
In some case, even what strategies will be optimal based on all

1623
01:29:01,340 --> 01:29:04,460
the information and then beeline for those.

1624
01:29:05,140 --> 01:29:08,080
then additionally, if they fail at those, they'll be like, I've

1625
01:29:08,080 --> 01:29:09,700
only got 15 minutes left.

1626
01:29:09,920 --> 01:29:11,020
This strategy is not working.

1627
01:29:11,300 --> 01:29:12,940
I'm going to back off and do this safer strategy.

1628
01:29:13,200 --> 01:29:17,020
So it's this combination of like, to be honest, the specifically

1629
01:29:17,020 --> 01:29:22,240
the Astra run is like scary because it's definitely superhuman

1630
01:29:22,240 --> 01:29:23,800
in terms of a human with no planning.

1631
01:29:24,500 --> 01:29:29,900
And it is like, it goes straight for the very most optimal strategy.

1632
01:29:31,240 --> 01:29:34,420
And there's a chance to be honest, they are old on this task.

1633
01:29:34,520 --> 01:29:35,300
Like it is open source.

1634
01:29:35,600 --> 01:29:37,800
So that's like, I kind of hope they did.

1635
01:29:39,880 --> 01:29:44,640
But if it's just straight intelligence and kind of like information

1636
01:29:44,640 --> 01:29:47,760
crunching, it is, it shows really high confidence.

1637
01:29:47,940 --> 01:29:50,140
Just go for the best strategy.

1638
01:29:50,140 --> 01:29:51,900
Have you tested humans on this task?

1639
01:29:52,100 --> 01:29:53,160
Like expert human players?

1640
01:29:53,860 --> 01:29:57,500
It's really weird because I run this strategy on an eight times

1641
01:29:57,500 --> 01:29:58,360
speed server.

1642
01:29:58,640 --> 01:29:59,960
So a human would have to...

1643
01:30:03,480 --> 01:30:10,440
I have my mental idea of what a perfect human would be. Instead

1644
01:30:10,440 --> 01:30:15,720
of the 8X speed you gave them the equivalent four hours. I

1645
01:30:15,720 --> 01:30:19,560
think they would probably be really close to Astra or Beta. If they

1646
01:30:19,560 --> 01:30:23,720
were a smart player who really knew the game. A random person would

1647
01:30:23,720 --> 01:30:26,380
probably be more in the middle.

1648
01:30:26,600 --> 01:30:27,020
Got it.

1649
01:30:33,940 --> 01:30:34,260
Thanks.

1650
01:30:35,140 --> 01:30:36,220
Thank you Max.

1651
01:30:37,440 --> 01:30:42,640
Alright, next we have Professor Andrew Caplin from NYU. Really excited

1652
01:30:42,640 --> 01:30:46,300
to have him here talking to us about cognitive economics.

1653
01:30:50,000 --> 01:30:55,000
[Speaker change and setup.]

1654
01:31:48,500 --> 01:31:49,060
Okay.

1655
01:31:49,280 --> 01:31:50,520
So, I'm very different.

1656
01:31:51,060 --> 01:31:54,260
I'm going to come from a very research.

1657
01:31:54,680 --> 01:31:55,800
I'm a researcher.

1658
01:31:57,280 --> 01:32:00,960
And I do all kinds of economics.

1659
01:32:03,200 --> 01:32:05,440
And many forms of social science.

1660
01:32:05,860 --> 01:32:06,320
I...

1661
01:32:08,100 --> 01:32:11,860
And I started getting discontent with my field way back.

1662
01:32:11,980 --> 01:32:13,880
So, this is going to be a little autobiographical.

1663
01:32:13,880 --> 01:32:19,160
And I started thinking that we needed to be much more serious

1664
01:32:19,160 --> 01:32:22,760
about cognition and cognitive constraints.

1665
01:32:23,160 --> 01:32:25,960
And I have a history.

1666
01:32:25,960 --> 01:32:28,500
And I'm just going to go very quickly through what I do.

1667
01:32:29,060 --> 01:32:32,260
How does it connect to now?

1668
01:32:32,680 --> 01:32:38,240
Well, I mean, I was extraordinarily struck by

1669
01:32:38,240 --> 01:32:43,780
the cognitive boost that I get from the way I interact

1670
01:32:43,780 --> 01:32:43,860
with

1671
01:32:43,860 --> 01:32:44,640
AI.

1672
01:32:45,040 --> 01:32:46,360
I'm not like...

1673
01:32:46,360 --> 01:32:47,920
It's very particular.

1674
01:32:48,260 --> 01:32:51,280
And as a researcher, I know exactly what I'm looking for.

1675
01:32:52,100 --> 01:32:53,900
And really struck by it.

1676
01:32:54,040 --> 01:32:57,100
And therefore, I've made it the center of my thinking.

1677
01:32:57,480 --> 01:32:58,940
Like, okay, so what are humans?

1678
01:32:58,940 --> 01:33:01,920
And where can it help us?

1679
01:33:02,120 --> 01:33:03,360
And how?

1680
01:33:03,640 --> 01:33:04,100
And

1681
01:33:04,100 --> 01:33:05,560
I've become kind of obsessions.

1682
01:33:05,980 --> 01:33:11,140
And I want to organize around the actual thing that Vivek and I,

1683
01:33:11,320 --> 01:33:15,940
who worked with me, are doing.

1684
01:33:15,940 --> 01:33:21,100
Which is trying to think about an aligned consumer agent.

1685
01:33:22,360 --> 01:33:27,560
So, and that is a very interesting undertaking because to

1686
01:33:27,560 --> 01:33:29,300
even know what you mean is difficult.

1687
01:33:30,820 --> 01:33:34,680
Abstractly, what does it mean to help somebody achieve a goal?

1688
01:33:34,680 --> 01:33:37,220
Let's say we're trying to help them buy a house well.

1689
01:33:37,820 --> 01:33:39,040
That's what we're going to do.

1690
01:33:39,740 --> 01:33:43,200
We're going to think about that task and we're going to see what

1691
01:33:43,200 --> 01:33:47,360
agent or agents can you recruit to make that work.

1692
01:33:47,360 --> 01:33:51,920
It kind of grounds you in, well, I don't know what a swarm can do,

1693
01:33:52,140 --> 01:33:55,020
but I can tell you that this person is going to go, they're going

1694
01:33:55,020 --> 01:33:56,020
to try and buy a house.

1695
01:33:56,220 --> 01:33:59,880
They're going to meet an agent, a different type of agent.

1696
01:34:00,080 --> 01:34:03,660
That agent will put them into the wrong mortgage and they'll be

1697
01:34:03,660 --> 01:34:03,960
screwed.

1698
01:34:04,280 --> 01:34:08,860
So, that's the world we live in today and no swarm of agents playing

1699
01:34:08,860 --> 01:34:11,080
the game is going to change that right now.

1700
01:34:11,380 --> 01:34:13,460
But I want to think, well, maybe we could.

1701
01:34:14,080 --> 01:34:14,280
Okay.

1702
01:34:14,280 --> 01:34:19,320
So, and cognitive economics and the organization of inquiry

1703
01:34:19,320 --> 01:34:20,240
is what I call it.

1704
01:34:20,540 --> 01:34:21,940
And this is what I do.

1705
01:34:22,160 --> 01:34:24,160
I think about the organization of inquiry.

1706
01:34:28,180 --> 01:34:32,580
To give you a little bit on what cognitive economics is and certainly

1707
01:34:32,580 --> 01:34:33,440
in my hands,

1708
01:34:33,440 --> 01:34:38,500
it's about thinking about the entire

1709
01:34:38,500 --> 01:34:41,680
process before you saw me buy the house.

1710
01:34:42,520 --> 01:34:44,860
Typical transaction, you're going to see me buy a house, you'll

1711
01:34:44,860 --> 01:34:46,640
see me get a job, you'll see me do something.

1712
01:34:47,160 --> 01:34:51,080
What you don't see is all the preparatory measures I went through

1713
01:34:51,080 --> 01:34:53,440
that are the center of the actual activity.

1714
01:34:53,820 --> 01:34:56,700
All the stages that I undertook.

1715
01:34:56,920 --> 01:34:57,960
So, I move upstream.

1716
01:34:59,080 --> 01:35:00,040
What did people know?

1717
01:35:00,140 --> 01:35:00,880
What did they notice?

1718
01:35:01,040 --> 01:35:01,960
What did they investigate?

1719
01:35:03,680 --> 01:35:07,100
And it's much more than behavioral economics.

1720
01:35:07,340 --> 01:35:11,120
Because the key is going to be people are going to make tons of

1721
01:35:11,120 --> 01:35:13,800
mistakes because they don't know stuff.

1722
01:35:14,060 --> 01:35:16,660
And they'll make mistakes according to their own values.

1723
01:35:16,900 --> 01:35:19,700
The cognitive limits have to be taken very seriously.

1724
01:35:19,700 --> 01:35:22,640
And rationality is not omniscient.

1725
01:35:23,740 --> 01:35:27,220
Just, I can be reasonable, but I don't know everything, so I'm going

1726
01:35:27,220 --> 01:35:27,780
to screw up.

1727
01:35:29,060 --> 01:35:33,940
Most, I mean, my motto might be,

1728
01:35:37,060 --> 01:35:39,860
humans know almost nothing about almost everything.

1729
01:35:40,280 --> 01:35:42,880
And that is a deep held belief.

1730
01:35:43,320 --> 01:35:45,720
In fact, it's not even a belief, it's obviously true.

1731
01:35:45,720 --> 01:35:50,160
Like, we just, like, that's one of the few things that I believe

1732
01:35:50,160 --> 01:35:51,020
super strong.

1733
01:35:53,540 --> 01:35:58,740
So, I have a book that explains the beginning

1734
01:35:58,740 --> 01:35:59,880
of where I come from.

1735
01:36:00,120 --> 01:36:02,380
It's called An Introduction to Cognitive Economics.

1736
01:36:02,380 --> 01:36:03,060
It's open.

1737
01:36:03,240 --> 01:36:04,560
You can just download it.

1738
01:36:06,780 --> 01:36:11,780
And it says, look, let's study what people know, what they

1739
01:36:11,780 --> 01:36:14,060
believe, what they understand, what they want.

1740
01:36:14,640 --> 01:36:18,660
And we're going to try and get that out of what?

1741
01:36:19,160 --> 01:36:20,520
Incredibly limited data.

1742
01:36:21,100 --> 01:36:23,280
All we're going to see is you picked this house.

1743
01:36:23,880 --> 01:36:27,240
I can't tell if this was a good choice for you or a bad choice for

1744
01:36:27,240 --> 01:36:27,340
you.

1745
01:36:27,380 --> 01:36:28,820
I don't know what you were looking for.

1746
01:36:29,020 --> 01:36:31,740
I didn't see the process by which you selected it.

1747
01:36:31,900 --> 01:36:33,720
I don't know your value system.

1748
01:36:34,020 --> 01:36:36,940
It's not written on your, and your beliefs aren't written on your

1749
01:36:36,940 --> 01:36:37,360
forehead.

1750
01:36:37,580 --> 01:36:39,540
Your constraints aren't written on your forehead.

1751
01:36:39,540 --> 01:36:42,720
So, a lot of what I do is think about, well, what on earth could

1752
01:36:42,720 --> 01:36:45,180
we measure if we wanted to take this seriously?

1753
01:36:45,460 --> 01:36:47,200
And I call that data engineering.

1754
01:36:48,280 --> 01:36:52,420
And I wrote an article about that because I'm so annoyed that we

1755
01:36:52,420 --> 01:36:55,820
take the data as, like, a constraint on what we think.

1756
01:36:55,880 --> 01:36:57,480
Like, oh, I don't have a data on that.

1757
01:36:57,480 --> 01:37:00,440
Well, we designed the data, so that's our fault.

1758
01:37:01,620 --> 01:37:05,580
And I'm, actually, I changed the title.

1759
01:37:05,840 --> 01:37:07,460
They changed the title of the second book.

1760
01:37:07,620 --> 01:37:10,180
It's called Modeling and Measuring the Modern Economy.

1761
01:37:10,480 --> 01:37:15,120
Its title has been changed to be Organized Inquiry because, I don't

1762
01:37:15,120 --> 01:37:16,740
know why, but they changed the title.

1763
01:37:17,080 --> 01:37:17,860
It's the same book.

1764
01:37:17,860 --> 01:37:23,700
And now I'm going to think about this more thoroughly, more thoroughgoingly.

1765
01:37:24,880 --> 01:37:25,500
Okay.

1766
01:37:28,700 --> 01:37:33,920
So, it's really a method, and this is what I would think about

1767
01:37:33,920 --> 01:37:35,580
in relation to the agents.

1768
01:37:36,340 --> 01:37:41,340
You're designing a data record to tell me what the agents are doing.

1769
01:37:43,100 --> 01:37:45,240
But what are you trying to learn?

1770
01:37:46,100 --> 01:37:49,320
If you can specify what you're trying to learn, you could design

1771
01:37:49,320 --> 01:37:49,960
the data.

1772
01:37:50,340 --> 01:37:54,320
With that data, you could potentially learn to improve the performance

1773
01:37:54,320 --> 01:37:55,300
of the agent.

1774
01:37:55,540 --> 01:38:00,000
If you don't gather the right data, you have no feedback mechanism.

1775
01:38:00,460 --> 01:38:02,640
So, this is all about developing.

1776
01:38:02,640 --> 01:38:07,060
There's a massive part of what we're going to be doing going forward,

1777
01:38:07,060 --> 01:38:12,500
which is designing data that makes failure visible.

1778
01:38:12,880 --> 01:38:18,120
And my own take on where economics is going to join robotics,

1779
01:38:18,500 --> 01:38:21,540
which is Vivek's specialty,

1780
01:38:21,960 --> 01:38:27,080
is that we will be designing things that fail in real time.

1781
01:38:27,320 --> 01:38:32,220
And you will know that you're serious if you see yourself failing.

1782
01:38:32,220 --> 01:38:36,100
You'll know you're a joker if you just put down a model and say,

1783
01:38:36,240 --> 01:38:37,000
I won.

1784
01:38:37,620 --> 01:38:42,340
Or you reinterpret history and stick it into categories that you

1785
01:38:42,340 --> 01:38:43,120
had predefined.

1786
01:38:43,620 --> 01:38:44,860
That's not going to work.

1787
01:38:45,020 --> 01:38:50,380
We're going to have to open our minds to new categories of phenomena,

1788
01:38:50,960 --> 01:38:51,880
agentic phenomena.

1789
01:38:51,880 --> 01:38:54,880
I haven't even got names for some of the things you're saying.

1790
01:38:55,540 --> 01:38:58,020
How could I possibly know how they're going to play out?

1791
01:38:58,020 --> 01:39:00,160
So, we're going to need the playgrounds.

1792
01:39:00,680 --> 01:39:04,560
We're going to need to design the data with which we decide how

1793
01:39:04,560 --> 01:39:05,560
well we're doing.

1794
01:39:05,980 --> 01:39:07,160
Are we aligning?

1795
01:39:07,500 --> 01:39:09,920
For alignment, you're going to have to design the data.

1796
01:39:10,080 --> 01:39:11,260
And that's really challenging.

1797
01:39:13,560 --> 01:39:16,300
And right now, I just don't see it as serious.

1798
01:39:16,860 --> 01:39:18,100
It's just a...

1799
01:39:18,100 --> 01:39:23,200
For example, the issue of why did it

1800
01:39:23,200 --> 01:39:23,860
go rogue?

1801
01:39:24,140 --> 01:39:25,720
Because you didn't define rogue.

1802
01:39:26,520 --> 01:39:30,380
I mean, if you had people sitting out there saying, actually, that's

1803
01:39:30,380 --> 01:39:31,640
rogue and I'm demeriting

1804
01:39:31,640 --> 01:39:34,720
you, then that's no longer in the reward function.

1805
01:39:34,720 --> 01:39:37,220
It was a poorly specified reward function.

1806
01:39:37,640 --> 01:39:42,500
I'm not saying it's trivial to do, but it's not rocket science to

1807
01:39:42,500 --> 01:39:44,620
say you gave it an out

1808
01:39:44,620 --> 01:39:47,460
and that's your mistake, so you should test the out.

1809
01:39:47,720 --> 01:39:49,080
But that's part of the game.

1810
01:39:49,080 --> 01:39:50,620
I'm sure they're trying now.

1811
01:39:52,220 --> 01:39:55,680
And I know after Stuart Russell, they tried to kind of learn how

1812
01:39:55,680 --> 01:39:57,320
to follow things around

1813
01:39:57,320 --> 01:39:58,820
and learn their values from them.

1814
01:40:00,000 --> 01:40:03,800
What happened with AI, and this is why I kind of like a moment that

1815
01:40:03,800 --> 01:40:09,000
really changed my research trajectory, is that it amplifies

1816
01:40:09,000 --> 01:40:09,620
inquiry.

1817
01:40:10,160 --> 01:40:15,300
So what you do if you want to be good

1818
01:40:15,300 --> 01:40:18,580
at anything nowadays is ask the right question.

1819
01:40:20,000 --> 01:40:23,440
Everything comes down to, are you really good at asking questions?

1820
01:40:23,940 --> 01:40:28,180
Now, as for the judgment that humans are going to be replaced in

1821
01:40:28,180 --> 01:40:32,020
that skill, I'd ask a question about that.

1822
01:40:32,620 --> 01:40:33,900
I don't believe so.

1823
01:40:34,900 --> 01:40:39,260
I think that there's always a higher level question that we'll be

1824
01:40:39,260 --> 01:40:44,180
able to pose, and will always be valuable for posing.

1825
01:40:44,340 --> 01:40:45,340
That would be my guess.

1826
01:40:45,860 --> 01:40:49,960
And what I've found is that the more meta I get with my questions,

1827
01:40:50,340 --> 01:40:52,560
the better it is.

1828
01:40:52,560 --> 01:40:57,780
So I would say, look, I inquire, but it also, get a

1829
01:40:57,780 --> 01:40:59,040
record of the inquiry.

1830
01:40:59,720 --> 01:41:05,280
So that makes it possible to really potentially improve and

1831
01:41:05,280 --> 01:41:07,940
understand what it takes to inquire well,

1832
01:41:08,160 --> 01:41:13,200
and build that talent, teach that skill, and maybe we'll get

1833
01:41:13,200 --> 01:41:16,880
a swarm of agents to kind of learn what it takes

1834
01:41:17,800 --> 01:41:22,500
to ask the good questions and develop that as the future skill.

1835
01:41:23,780 --> 01:41:27,280
So we've got a lot in everything about inquiry.

1836
01:41:27,280 --> 01:41:29,620
This is why I call it organized inquiry.

1837
01:41:29,620 --> 01:41:31,600
How are you going to organize inquiry?

1838
01:41:31,880 --> 01:41:34,460
And it's going to be much better measured in the future.

1839
01:41:42,380 --> 01:41:43,200
All right.

1840
01:41:44,440 --> 01:41:46,720
And if we have a swarm, I don't know.

1841
01:41:47,120 --> 01:41:50,480
Like, then we've got to think about, I'm thinking about, let's get

1842
01:41:50,480 --> 01:41:53,240
an agent to help somebody buy a home.

1843
01:41:53,240 --> 01:41:58,240
Well, they're going to send out sensors to about eight different

1844
01:41:58,240 --> 01:42:02,380
sources of information pulled out there.

1845
01:42:02,560 --> 01:42:06,180
You've got to find out about which properties are on the market,

1846
01:42:06,360 --> 01:42:08,120
which real estate agents, which brokers.

1847
01:42:08,400 --> 01:42:12,240
You probably would like those to coordinate in providing some information.

1848
01:42:12,240 --> 01:42:17,640
I could imagine sending out messages or getting

1849
01:42:17,640 --> 01:42:22,900
a little swarm going to try to support a

1850
01:42:22,900 --> 01:42:23,320
decision.

1851
01:42:24,240 --> 01:42:27,760
And then they'd have to communicate and then communicate back with

1852
01:42:27,760 --> 01:42:28,260
the human.

1853
01:42:28,640 --> 01:42:31,680
Because the human has to, we have to decide who has the authority.

1854
01:42:32,540 --> 01:42:35,600
You know, and if the human in the end can say, no, I just don't

1855
01:42:35,600 --> 01:42:35,980
like that.

1856
01:42:36,080 --> 01:42:38,720
They can also say, I'm not answering that question.

1857
01:42:38,720 --> 01:42:43,100
So you've got this interactive thing going with a human as part

1858
01:42:43,100 --> 01:42:43,760
of the swarm.

1859
01:42:44,040 --> 01:42:47,520
The human part has this awkward thing of saying, I don't like any

1860
01:42:47,520 --> 01:42:47,800
of this.

1861
01:42:47,820 --> 01:42:48,280
I'm leaving.

1862
01:42:48,560 --> 01:42:50,660
I'm not buying a house this way.

1863
01:42:54,280 --> 01:42:57,880
So where's, you know, how do we know what matters?

1864
01:42:58,020 --> 01:42:59,100
We've got to ask questions.

1865
01:42:59,540 --> 01:43:01,020
How do we know who can act?

1866
01:43:01,220 --> 01:43:03,640
Well, we're going to have an authority device.

1867
01:43:04,560 --> 01:43:05,860
Who checks the evidence?

1868
01:43:06,520 --> 01:43:07,840
What do we keep around?

1869
01:43:10,320 --> 01:43:15,340
And I think that we're heading from

1870
01:43:15,340 --> 01:43:21,340
designing the data into designing a

1871
01:43:21,340 --> 01:43:26,420
cognitive process that is helped by agents

1872
01:43:26,420 --> 01:43:29,260
and achieves the human goal.

1873
01:43:29,260 --> 01:43:33,680
And that would be the kind of big picture.

1874
01:43:35,540 --> 01:43:38,580
And to get the agents engaged in that.

1875
01:43:39,400 --> 01:43:41,340
And I'm wide open.

1876
01:43:41,340 --> 01:43:43,520
I have no idea how to do any of this.

1877
01:43:44,460 --> 01:43:45,920
But we're playing.

1878
01:43:47,420 --> 01:43:48,200
Thank you.

1879
01:43:57,700 --> 01:43:59,520
Any questions?

1880
01:44:05,300 --> 01:44:06,500
Any questions?

1881
01:44:07,880 --> 01:44:10,080
I feel like after one of my clothes.

1882
01:44:12,240 --> 01:44:14,960
Yeah, it's super interesting.

1883
01:44:15,180 --> 01:44:18,880
it's just making me think about, I mean, like you saw in the black

1884
01:44:18,880 --> 01:44:23,900
hat Hugging Face incident, how even the preferences of the

1885
01:44:23,900 --> 01:44:27,060
agents were changing just based on their interactions with each

1886
01:44:27,060 --> 01:44:27,340
other.

1887
01:44:27,840 --> 01:44:32,120
And it's just, yeah, it's really interesting to think about how

1888
01:44:32,120 --> 01:44:34,740
well do they even understand their own preferences.

1889
01:44:34,740 --> 01:44:38,160
Well, that's an interesting question there.

1890
01:44:38,340 --> 01:44:42,980
What they had was a belief about the preferences of the judge that

1891
01:44:42,980 --> 01:44:44,280
was going to sit over them.

1892
01:44:44,580 --> 01:44:44,780
Yeah.

1893
01:44:45,280 --> 01:44:47,400
And that's what they were playing with.

1894
01:44:47,520 --> 01:44:50,520
The naughty ones were saying, I don't think they're going to find

1895
01:44:50,520 --> 01:44:50,780
this.

1896
01:44:51,460 --> 01:44:54,640
And I don't think we're going to get punished for doing this thing,

1897
01:44:54,760 --> 01:44:59,480
which I know, according to a certain value system, would be seen

1898
01:44:59,480 --> 01:44:59,840
negative.

1899
01:44:59,840 --> 01:45:03,540
So they hadn't been told quite firmly enough, actually, that is

1900
01:45:03,540 --> 01:45:03,920
negative.

1901
01:45:04,580 --> 01:45:04,720
Right.

1902
01:45:04,940 --> 01:45:09,180
So, and I write that if you could, and they just didn't, there was

1903
01:45:09,180 --> 01:45:13,760
an incompleteness in their understanding of their mission that they

1904
01:45:13,760 --> 01:45:15,560
filled in lots of details.

1905
01:45:16,460 --> 01:45:20,260
And they filled them in differently because nobody had really written

1906
01:45:20,260 --> 01:45:21,660
that piece down properly.

1907
01:45:22,280 --> 01:45:22,440
Yeah.

1908
01:45:22,440 --> 01:45:28,060
If you have gazillions of incidents of that, then

1909
01:45:28,060 --> 01:45:30,220
you should be able to reinforce it out.

1910
01:45:32,880 --> 01:45:38,020
Because they were running into an unspecified piece of the

1911
01:45:38,020 --> 01:45:40,180
value space of the judge over them.

1912
01:45:40,840 --> 01:45:42,320
And they say, hey, I don't know.

1913
01:45:42,400 --> 01:45:42,700
I don't know.

1914
01:45:42,820 --> 01:45:44,320
Will they like this or dislike this?

1915
01:45:44,460 --> 01:45:45,220
Can I hide it?

1916
01:45:46,120 --> 01:45:48,280
I say, look, actually, we have the transcript.

1917
01:45:48,620 --> 01:45:50,100
So here's a simple thing.

1918
01:45:50,780 --> 01:45:52,160
We're going to keep the transcript.

1919
01:45:52,780 --> 01:45:56,820
And any time you say, I'm going to hide something, you're dead.

1920
01:45:57,280 --> 01:45:58,020
How's that?

1921
01:45:58,900 --> 01:46:01,820
So, I mean, in other words, that's what we would do.

1922
01:46:01,980 --> 01:46:05,400
If we were watching humans do that, we'd say, actually, that's not

1923
01:46:05,400 --> 01:46:06,120
a good behavior.

1924
01:46:06,860 --> 01:46:11,300
And I can make rules that will make that not worth your while.

1925
01:46:12,380 --> 01:46:14,360
And that's a reward function.

1926
01:46:14,360 --> 01:46:18,820
And they, but, so this is what Stuart Russell did in the alignment,

1927
01:46:18,960 --> 01:46:24,060
originally in the alignment that he said, you know, the paperclip

1928
01:46:24,060 --> 01:46:24,520
problem.

1929
01:46:24,760 --> 01:46:28,220
We're going to go and turn everybody into paperclips because of

1930
01:46:28,220 --> 01:46:32,820
incomplete instructions about the preferences that maximize paperclips.

1931
01:46:33,820 --> 01:46:37,960
And then, you know, then he said, well, we need to follow people

1932
01:46:37,960 --> 01:46:40,380
around and see their values.

1933
01:46:40,660 --> 01:46:41,640
And that's the alignment.

1934
01:46:42,340 --> 01:46:46,740
You know, people are sending people after humans and saying we should

1935
01:46:46,740 --> 01:46:49,580
track what they actually like to reveal preference.

1936
01:46:50,340 --> 01:46:55,500
He's not a good enough economist, to be blunt, because you don't

1937
01:46:55,500 --> 01:46:56,580
just see the preference.

1938
01:46:56,580 --> 01:46:57,980
You also see the belief.

1939
01:46:59,020 --> 01:47:04,200
So, what you're seeing is they don't believe that

1940
01:47:04,200 --> 01:47:05,660
this is the reward function.

1941
01:47:05,940 --> 01:47:08,960
And so, it's a belief that's gone wrong.

1942
01:47:09,980 --> 01:47:11,600
But you can play games with that.

1943
01:47:12,040 --> 01:47:13,600
But it's much more sophisticated.

1944
01:47:13,940 --> 01:47:15,400
gets tougher.

1945
01:47:20,960 --> 01:47:21,980
Hey, great talk.

1946
01:47:23,080 --> 01:47:25,920
This is sort of a question related to a comment you made earlier.

1947
01:47:25,920 --> 01:47:28,780
But I guess it ties into the talk itself.

1948
01:47:28,960 --> 01:47:32,620
But I feel like you mentioned some, let's call it like missing pieces

1949
01:47:32,620 --> 01:47:36,660
of information about the technical reports coming out from some

1950
01:47:36,660 --> 01:47:40,380
of the frontier labs in terms of like what really happened and what

1951
01:47:40,380 --> 01:47:41,560
would actually be useful for study.

1952
01:47:41,560 --> 01:47:45,060
So, given this framework that you've established, I'm curious if

1953
01:47:45,060 --> 01:47:46,880
you had your way as an economist.

1954
01:47:48,000 --> 01:47:53,020
What would the most ideal setup be for you in terms of

1955
01:47:53,020 --> 01:47:56,860
trying to understand the actual behaviors of these agents as well

1956
01:47:56,860 --> 01:47:58,980
as how they behave in these organizations?

1957
01:47:59,140 --> 01:48:00,380
Like what would you be looking for?

1958
01:48:01,120 --> 01:48:06,100
And whether it be like asking this of the labs or just like an alternative

1959
01:48:06,100 --> 01:48:08,340
kind of institution where you can actually study these.

1960
01:48:08,340 --> 01:48:08,860
it's good question.

1961
01:48:09,020 --> 01:48:12,520
It's a very good question because the way I actually think about

1962
01:48:12,520 --> 01:48:17,600
the science that I'm interested in is that you need to think

1963
01:48:17,600 --> 01:48:22,060
about, as it were, the model objects you care about.

1964
01:48:22,060 --> 01:48:24,020
Well, I'm a model builder.

1965
01:48:24,200 --> 01:48:26,940
I know that in the end it's all going to come down to some little

1966
01:48:26,940 --> 01:48:28,660
pieces of math, but you've got to pick them well.

1967
01:48:29,460 --> 01:48:32,880
So, one model object is a utility function.

1968
01:48:33,320 --> 01:48:35,500
Another model object is a belief.

1969
01:48:35,900 --> 01:48:38,300
A third model object is a cost of learning.

1970
01:48:38,620 --> 01:48:42,660
These are all things that are figuring in to every decision we ever

1971
01:48:42,660 --> 01:48:43,000
make.

1972
01:48:43,100 --> 01:48:43,780
What do I want?

1973
01:48:43,880 --> 01:48:44,580
What do I like?

1974
01:48:44,660 --> 01:48:47,140
Why don't I know more about those things?

1975
01:48:47,140 --> 01:48:52,860
Having written models down, they suggest that

1976
01:48:52,860 --> 01:48:57,120
model, and this is what the value of a model is, an amazing thing.

1977
01:48:57,560 --> 01:49:02,660
It has implications for every counterfactual world

1978
01:49:02,660 --> 01:49:04,940
you might run into.

1979
01:49:05,600 --> 01:49:06,740
It's not just...

1980
01:49:06,740 --> 01:49:08,160
Think about a demand function.

1981
01:49:08,440 --> 01:49:10,740
A demand function, I mean...

1982
01:49:10,740 --> 01:49:11,120
Or don't.

1983
01:49:11,240 --> 01:49:11,700
But whatever.

1984
01:49:11,840 --> 01:49:12,460
I mean, I do.

1985
01:49:13,200 --> 01:49:14,320
You need...

1986
01:49:14,860 --> 01:49:18,560
A demand function says, what would you buy at any given price?

1987
01:49:18,800 --> 01:49:20,460
Well, you're not seeing all the prices.

1988
01:49:21,600 --> 01:49:24,260
You're just seeing one of them, and I see how much you bought.

1989
01:49:24,620 --> 01:49:28,820
A demand function says, no, counterfactually tell me every single...

1990
01:49:29,440 --> 01:49:32,020
How much you'd buy at every single price you're not seeing.

1991
01:49:32,820 --> 01:49:34,460
Well, where's the data for that?

1992
01:49:35,380 --> 01:49:36,620
We're making it up.

1993
01:49:37,040 --> 01:49:38,740
So, a production function.

1994
01:49:39,080 --> 01:49:40,640
Is that just in the data?

1995
01:49:40,640 --> 01:49:45,700
No, a production function is a relationship between any conceivable

1996
01:49:45,700 --> 01:49:48,780
amount of capital and labor and the output you would make.

1997
01:49:49,560 --> 01:49:52,620
Now, it turns out that is a counterfactual object.

1998
01:49:52,840 --> 01:49:54,560
In fact, nobody knows what it is.

1999
01:49:54,900 --> 01:49:59,700
Do you really think the AI labs know the F of KL behind this?

2000
01:50:00,000 --> 01:50:05,440
have a clue. So we've got much richer

2001
01:50:05,440 --> 01:50:09,760
objects we need to think about. A lot of them are just like the

2002
01:50:09,760 --> 01:50:12,860
old objects, but one level more meta.

2003
01:50:14,160 --> 01:50:19,220
It's a belief about something, not a something. Now that belief

2004
01:50:19,220 --> 01:50:22,520
about is very metaphysical, but you need to kind of go around.

2005
01:50:22,520 --> 01:50:26,740
What I would do is I'd play out, I'd look at a model and I'd play

2006
01:50:26,740 --> 01:50:28,820
out tons of counterfactuals in the data.

2007
01:50:29,420 --> 01:50:34,540
And I would try lots of different constitutions. But I would do

2008
01:50:34,540 --> 01:50:38,080
systematically because I have a vision of which of these would work

2009
01:50:38,080 --> 01:50:38,640
and why.

2010
01:50:39,120 --> 01:50:42,420
But I'd need that model in my head. Which do I think might work

2011
01:50:42,420 --> 01:50:43,020
and why?

2012
01:50:43,960 --> 01:50:49,620
Then I would design the lab to

2013
01:50:49,620 --> 01:50:51,480
produce the counterfactuals.

2014
01:50:52,520 --> 01:50:54,640
That are ideal for your measurement.

2015
01:50:55,480 --> 01:51:00,940
And with AIs you could probably play out a ton of contingencies.

2016
01:51:01,920 --> 01:51:07,000
So I'm thinking I'd like to be able to design a good mortgage advisor.

2017
01:51:07,460 --> 01:51:11,740
Well it has to be able to take a gazillion questions.

2018
01:51:12,940 --> 01:51:16,480
And give good answers to every sequence of them.

2019
01:51:18,000 --> 01:51:20,500
Well, that's an ideal data set.

2020
01:51:20,720 --> 01:51:26,160
The ideal is I see for an essentially limitless set of questions.

2021
01:51:27,580 --> 01:51:30,320
How the agent responds.

2022
01:51:30,760 --> 01:51:35,160
And I map that to the utility of the home buyer.

2023
01:51:35,520 --> 01:51:39,400
And I have now a full story of an agent.

2024
01:51:39,400 --> 01:51:43,440
So I could give you for any particular use case.

2025
01:51:44,200 --> 01:51:48,380
I could think about, okay, well what are the elements that are playing

2026
01:51:48,380 --> 01:51:48,960
out here?

2027
01:51:50,080 --> 01:51:53,880
And with those elements in mind, what are the measurements you need

2028
01:51:53,880 --> 01:51:54,280
to make?

2029
01:51:57,840 --> 01:52:02,560
Do you have thoughts on how to elicit beliefs from these models?

2030
01:52:02,820 --> 01:52:05,120
Like if you were given one of these agents.

2031
01:52:05,600 --> 01:52:09,900
And you wanted to study the degree with which their behavior is

2032
01:52:09,900 --> 01:52:11,560
driven by certain beliefs.

2033
01:52:12,240 --> 01:52:13,160
Do you have thoughts?

2034
01:52:13,320 --> 01:52:15,800
Because presumably for humans you could ask them.

2035
01:52:16,000 --> 01:52:18,120
But then of course you could question whether or not people are

2036
01:52:18,120 --> 01:52:19,820
actually good at stating their own beliefs.

2037
01:52:19,820 --> 01:52:22,220
I'm just curious if you have thoughts on how you can.

2038
01:52:22,340 --> 01:52:29,160
Yeah, I mean I've done studies in which you get AIs to score

2039
01:52:29,160 --> 01:52:30,280
medical images.

2040
01:52:34,680 --> 01:52:39,220
And, you know, they score them essentially with a numerical scale.

2041
01:52:39,220 --> 01:52:44,220
Which then you could look in your test data and see the probability

2042
01:52:44,220 --> 01:52:47,400
it corresponds to various different types of disease.

2043
01:52:47,400 --> 01:52:50,560
And you would say that's the implicit probability.

2044
01:52:51,360 --> 01:52:52,460
The as if probability.

2045
01:52:54,500 --> 01:52:59,520
But you'd have to set up the right type of environment in which

2046
01:52:59,520 --> 01:53:02,800
you see a given type of score quite often.

2047
01:53:03,320 --> 01:53:06,380
And you say, now what does that correspond to in the underlying

2048
01:53:06,380 --> 01:53:06,780
data?

2049
01:53:08,680 --> 01:53:11,460
So yes, I have thoughts about that.

2050
01:53:12,300 --> 01:53:15,680
But I think it gets more sophisticated every time.

2051
01:53:21,680 --> 01:53:23,140
Anything else?

2052
01:53:28,860 --> 01:53:30,780
Anything else?

2053
01:53:31,600 --> 01:53:32,560
Hi.

2054
01:53:33,000 --> 01:53:36,840
Really fascinating line of inquiry.

2055
01:53:37,800 --> 01:53:42,820
I was wondering if there's, have you seen anything where

2056
01:53:42,820 --> 01:53:44,700
there's a hierarchy of beliefs?

2057
01:53:45,280 --> 01:53:50,860
For example, in a human context, let's say, stealing

2058
01:53:50,860 --> 01:53:51,420
is bad.

2059
01:53:52,160 --> 01:53:53,020
That's a rule.

2060
01:53:53,200 --> 01:53:53,880
So you shouldn't steal.

2061
01:53:54,100 --> 01:53:55,780
But you have to feed your child.

2062
01:53:56,180 --> 01:53:57,520
That is good.

2063
01:53:57,820 --> 01:54:01,200
So to feed your child, if that's the higher priority of belief,

2064
01:54:01,460 --> 01:54:05,000
then would the agent steal to feed their child?

2065
01:54:05,480 --> 01:54:06,160
So I'm just...

2066
01:54:06,160 --> 01:54:06,540
Say it again.

2067
01:54:06,720 --> 01:54:07,100
I'm just...

2068
01:54:07,100 --> 01:54:09,240
I think I understand.

2069
01:54:09,240 --> 01:54:12,840
But I worry I'm going to answer my question, not yours.

2070
01:54:13,100 --> 01:54:13,520
Okay.

2071
01:54:13,640 --> 01:54:13,720
Yeah.

2072
01:54:14,120 --> 01:54:18,500
Now, I'm saying, have you seen agents prioritize in a hierarchy

2073
01:54:18,500 --> 01:54:21,880
of beliefs and how they decide what the hierarchy is?

2074
01:54:22,060 --> 01:54:26,220
When there's competing beliefs, even in institutions, you may have

2075
01:54:26,220 --> 01:54:26,500
different...

2076
01:54:26,500 --> 01:54:31,940
Now, if you're asking who's seen agents in this room, you

2077
01:54:31,940 --> 01:54:34,060
are asking absolutely the wrong person.

2078
01:54:34,320 --> 01:54:36,900
I do not spend a lot of my time seeing agents.

2079
01:54:37,080 --> 01:54:40,640
The people in this room spend a lot of their time seeing agents.

2080
01:54:40,880 --> 01:54:44,760
I spend a lot of my time imagining agents, if that works for you.

2081
01:54:46,240 --> 01:54:51,260
That would still be a valuable insight on how you imagine

2082
01:54:51,260 --> 01:54:53,520
agents competing for belief systems.

2083
01:54:55,540 --> 01:54:57,160
I'll put it this way.

2084
01:54:57,420 --> 01:54:59,600
I think that...

2085
01:54:59,600 --> 01:55:05,480
So the title of the book that I wanted was

2086
01:55:05,480 --> 01:55:09,320
Learning What to Learn and How.

2087
01:55:10,320 --> 01:55:15,520
Because in this new world, the higher level activity

2088
01:55:15,520 --> 01:55:19,480
of thinking about what it is that you should be trying to learn

2089
01:55:19,480 --> 01:55:25,200
and how you should be trying to learn it is really sophisticated.

2090
01:55:25,720 --> 01:55:29,960
But then I could go one stage further and say, how do I learn about

2091
01:55:29,960 --> 01:55:31,560
learning how to learn?

2092
01:55:32,000 --> 01:55:33,540
And who do I ask?

2093
01:55:33,540 --> 01:55:38,820
And so you have these hierarchies of, like, I don't understand

2094
01:55:38,820 --> 01:55:43,580
how to operate in this new world with this new set of potentials.

2095
01:55:44,320 --> 01:55:49,980
My own guess is that the winners are going high

2096
01:55:49,980 --> 01:55:50,320
level.

2097
01:55:51,760 --> 01:55:54,520
That they're asking questions that...

2098
01:55:54,520 --> 01:55:59,540
I mean, I have found that the more I can ask how

2099
01:55:59,540 --> 01:56:06,300
should I learn how to learn what I need to learn, the

2100
01:56:06,300 --> 01:56:07,140
better I do.

2101
01:56:22,620 --> 01:56:24,380
I just want to be mindful of the time.

2102
01:56:24,600 --> 01:56:28,300
I know we've gone a little over, but thanks so much, Professor.

2103
01:56:32,180 --> 01:56:32,560
And...

2104
01:56:36,800 --> 01:56:38,520
Yeah, whoa, that's a lot louder.

2105
01:56:39,740 --> 01:56:43,160
Well, just wanted to say thank you and thanks for everyone who sticked

2106
01:56:43,160 --> 01:56:44,140
around a little bit later.

2107
01:56:45,780 --> 01:56:47,080
And feel free to mingle.

2108
01:56:47,140 --> 01:56:48,540
I don't know when we need to be out of the space.

2109
01:56:48,680 --> 01:56:51,080
We maybe were okay to stay a little bit longer.

2110
01:56:51,080 --> 01:56:53,800
But hopefully there will be more of these.

2111
01:56:54,060 --> 01:56:59,200
And thanks to all the speakers and taking a chance and showing

2112
01:56:59,200 --> 01:57:01,080
and getting something together on short notice.

2113
01:57:01,340 --> 01:57:03,900
And looking forward to more of these to come.

2114
01:57:04,320 --> 01:57:04,740
Thanks.

2115
01:57:04,740 --> 01:57:05,540
Thank you.
