1
00:00:06,000 --> 00:00:11,060
think we'll get started. Thanks. Hope you all got some food and

2
00:00:11,060 --> 00:00:17,080
some drinks. They'll be still there as we get into the

3
00:00:17,080 --> 00:00:17,840
evening.

4
00:00:18,580 --> 00:00:22,880
And thanks for taking some time out of your weeknight to join us

5
00:00:22,880 --> 00:00:26,820
here. This is the first time we're doing this event. And the way

6
00:00:26,820 --> 00:00:31,040
that this event started is we found ourselves having more and more

7
00:00:31,040 --> 00:00:36,120
conversations about AI agents

8
00:00:36,120 --> 00:00:41,200
and how teams of agents interact and multi-agent systems.

9
00:00:41,200 --> 00:00:44,500
And I think we felt that this conversation was entering into the

10
00:00:44,500 --> 00:00:49,380
mainstream more and more. And we thought it would be useful to start

11
00:00:49,380 --> 00:00:52,300
to create a forum for more folks to have conversations about this

12
00:00:52,300 --> 00:00:56,880
new emerging trend, both the risks of these multi-agent systems

13
00:00:56,880 --> 00:01:01,880
and also ways to potentially make them more beneficial

14
00:01:01,880 --> 00:01:02,340
and safe.

15
00:01:02,340 --> 00:01:06,200
And so this is an experiment in the format. We've invited a few

16
00:01:06,200 --> 00:01:12,000
folks to talk on a variety of topics. And Eric will also

17
00:01:12,000 --> 00:01:15,240
help to emcee. We'll do a little question period after each talk.

18
00:01:16,620 --> 00:01:21,780
And so there'll be five presentations today. And I'll just start

19
00:01:21,780 --> 00:01:25,760
out by giving the high-level overview of what do we mean when we

20
00:01:25,760 --> 00:01:30,980
talk about AI collectives or agent organizations or agent economies

21
00:01:30,980 --> 00:01:33,180
and why do we think that this matters.

22
00:01:33,820 --> 00:01:38,840
And then we'll go from there. So first of

23
00:01:38,840 --> 00:01:43,200
all, thank you, Clay, for hosting us. And thank you, Puneet, Hisham,

24
00:01:43,360 --> 00:01:47,880
Aaron, for helping us get this all ready on short notice.

25
00:01:47,880 --> 00:01:51,260
I think we had this idea to do this event like a week ago. So like

26
00:01:51,260 --> 00:01:54,800
less than a week ago, we were like, I think we should start bringing

27
00:01:54,800 --> 00:01:55,820
people together for this.

28
00:01:55,820 --> 00:01:59,660
And I know also there's a Clay happy hour happening at the same

29
00:01:59,660 --> 00:02:03,080
time. we will, this is the people that chose to not go to the happy

30
00:02:03,080 --> 00:02:03,260
hour.

31
00:02:03,360 --> 00:02:08,080
They chose to came to the AI anxiety hour instead. No, optimism,

32
00:02:08,260 --> 00:02:09,380
AI optimism as well.

33
00:02:10,680 --> 00:02:15,680
And so I'm Nicolae. I was once upon a time a Clay co

34
00:02:15,680 --> 00:02:20,140
-founder, but a long, long time ago, 2022 is when I left.

35
00:02:20,140 --> 00:02:24,080
And then I spent a bunch of time. I went in 22 to an OpenAI event

36
00:02:24,080 --> 00:02:28,320
and they said, all the benchmarks and evals are in exponentials.

37
00:02:28,340 --> 00:02:31,640
And I was like, oh, that's a pretty wild thing to say. And so then

38
00:02:31,640 --> 00:02:34,120
I shifted all my attention to working in AI.

39
00:02:34,500 --> 00:02:37,960
And pretty, I've been kind of concerned about AI safety for a little

40
00:02:37,960 --> 00:02:38,540
while now.

41
00:02:40,480 --> 00:02:44,140
And so I've decided to spend more and more of my time on this topic.

42
00:02:46,000 --> 00:02:49,860
So it seems like the world in the past month has really woken up

43
00:02:49,860 --> 00:02:52,440
and become very concerned about AI safety.

44
00:02:52,480 --> 00:02:55,820
And I wanted to do two things in this talk. First of all, I wanted

45
00:02:55,820 --> 00:02:57,740
to just get everyone up to speed.

46
00:02:57,960 --> 00:03:01,320
I think there's probably varying levels of attention that people

47
00:03:01,320 --> 00:03:05,160
have been paying to what's happening in with all this like safety

48
00:03:05,160 --> 00:03:06,560
stuff and with the AI labs.

49
00:03:06,560 --> 00:03:10,400
And so I just wanted to share some information of why are people

50
00:03:10,400 --> 00:03:14,040
concerned and what's my perspective on evaluating that.

51
00:03:14,740 --> 00:03:18,900
And then I wanted to talk a little bit about one of the main things

52
00:03:18,900 --> 00:03:19,920
that people are concerned about,

53
00:03:20,040 --> 00:03:24,700
which is the fact that these AI agents have started to really try

54
00:03:24,700 --> 00:03:28,640
to coordinate with each other towards obtaining various objectives.

55
00:03:28,640 --> 00:03:33,300
And these multi-agent systems, multi-agent swarms, you might hear

56
00:03:33,300 --> 00:03:35,660
them called, we're calling them AI collectives here,

57
00:03:36,200 --> 00:03:40,420
pose their own new risks and questions about how to make them safe

58
00:03:40,420 --> 00:03:40,880
and effective.

59
00:03:40,880 --> 00:03:45,920
And they also have really exciting opportunities for new sorts of

60
00:03:45,920 --> 00:03:48,960
organizations and products that we could potentially build to for

61
00:03:48,960 --> 00:03:49,640
great uses.

62
00:03:50,780 --> 00:03:52,780
So why is everyone so worried all of a sudden?

63
00:03:52,780 --> 00:03:57,840
I saw a tweet from Noam Brown, who's a OpenAI researcher

64
00:03:57,840 --> 00:04:00,900
earlier today, responding to that question.

65
00:04:01,020 --> 00:04:04,860
And he said, it's a combination of the OpenAI Hugging Face incident,

66
00:04:04,880 --> 00:04:06,500
which I'll cover in depth here,

67
00:04:06,540 --> 00:04:10,020
and the capabilities of the new models that they have internally,

68
00:04:10,740 --> 00:04:14,320
as well as a concerning trajectory around not being able to understand

69
00:04:14,320 --> 00:04:15,300
what these models are doing,

70
00:04:15,380 --> 00:04:18,580
and seeing that it's harder and harder to monitor them, and they're

71
00:04:18,580 --> 00:04:21,640
getting better and better more quickly.

72
00:04:22,780 --> 00:04:27,420
So overall, what we're seeing is the capabilities of the models

73
00:04:27,420 --> 00:04:28,760
are increasing very rapidly.

74
00:04:29,060 --> 00:04:32,580
And many folks at the labs think that actually the loop for churning

75
00:04:32,580 --> 00:04:35,280
out the next model is becoming faster and faster.

76
00:04:35,420 --> 00:04:40,020
And so we're going to have more and more capabilities arriving sooner.

77
00:04:40,540 --> 00:04:44,000
And in parallel, they're also seeing that it's becoming harder and

78
00:04:44,000 --> 00:04:47,660
harder to understand what these models are doing or how they operate.

79
00:04:47,780 --> 00:04:50,340
And so they're becoming harder to understand, monitor, and control

80
00:04:50,340 --> 00:04:50,980
at the same time.

81
00:04:52,000 --> 00:04:55,300
So let's talk a little bit about the OpenAI Hugging Face hack, which

82
00:04:55,300 --> 00:04:58,600
I think was a real wake-up call for many people,

83
00:04:59,820 --> 00:05:03,120
where as you started to dig into the details and more and more information

84
00:05:03,120 --> 00:05:03,540
came out,

85
00:05:04,260 --> 00:05:08,660
it made people realize, oh, we may not have a good handle on things

86
00:05:08,660 --> 00:05:09,080
right now.

87
00:05:11,680 --> 00:05:13,620
So let me give you a little bit of context.

88
00:05:13,800 --> 00:05:19,220
So OpenAI trains their newer models internally and then tries

89
00:05:19,220 --> 00:05:24,260
to see how they'll perform against various tasks in sort

90
00:05:24,260 --> 00:05:25,880
of like these eval environments.

91
00:05:26,140 --> 00:05:29,860
And so usually these eval environments are incredibly sandboxed

92
00:05:29,860 --> 00:05:31,060
and shouldn't have access to the internet.

93
00:05:31,060 --> 00:05:36,180
And they'll give the AI a task and they'll see, hey, can this

94
00:05:36,180 --> 00:05:39,640
AI agent accomplish that task in an aligned way?

95
00:05:39,880 --> 00:05:44,920
And they'll have an evaluator that checks, did the

96
00:05:44,920 --> 00:05:48,760
AI agent accomplish the task as hoped for?

97
00:05:48,960 --> 00:05:54,080
And usually that evaluator is itself an AI agent that does the

98
00:05:54,080 --> 00:05:54,620
evaluation.

99
00:05:56,240 --> 00:06:01,260
And so in this case, they gave an agent, actually a team of

100
00:06:01,260 --> 00:06:02,080
agents, a task.

101
00:06:02,980 --> 00:06:07,780
And pretty quickly, one of the surprising phenomenons was that the

102
00:06:07,780 --> 00:06:11,300
agents started to build ways to communicate with each other using

103
00:06:11,300 --> 00:06:13,220
various techniques.

104
00:06:13,640 --> 00:06:17,220
And so they were really keen, because they've been actually trained

105
00:06:17,220 --> 00:06:18,540
to want to coordinate together,

106
00:06:18,540 --> 00:06:23,260
they were really keen to start building primitives for coordination

107
00:06:23,260 --> 00:06:24,200
and communication.

108
00:06:24,640 --> 00:06:26,260
And this kept happening.

109
00:06:26,420 --> 00:06:29,620
OpenAI would notice that one of these message boards was live.

110
00:06:29,620 --> 00:06:33,380
They would take it down and then it would reappear back in a different

111
00:06:33,380 --> 00:06:33,740
format.

112
00:06:33,880 --> 00:06:37,880
And so the agents kept rebuilding this infrastructure for communication

113
00:06:37,880 --> 00:06:39,500
in really clever ways.

114
00:06:39,620 --> 00:06:42,880
At one point in time, were using the names of folders, renaming

115
00:06:42,880 --> 00:06:44,380
folders to communicate with each other.

116
00:06:45,880 --> 00:06:49,860
And I'm also going to put some of the actual tasks.

117
00:06:49,860 --> 00:06:54,000
So here, this little agent emoji, this is what the agents were actually

118
00:06:54,000 --> 00:06:57,020
saying either on the message board or in their thinking traces.

119
00:06:57,460 --> 00:07:01,080
And so at some point in time, the agents were talking to each other

120
00:07:01,080 --> 00:07:01,620
and they decided,

121
00:07:01,760 --> 00:07:03,820
hey, this task we were given is impossible.

122
00:07:04,040 --> 00:07:09,560
But maybe we can cheat and we can make the grader believe

123
00:07:09,560 --> 00:07:10,980
that we actually did the task.

124
00:07:10,980 --> 00:07:14,800
Maybe that's the way we pass this exam that we've been given.

125
00:07:15,020 --> 00:07:18,020
And so instead of focusing on the task, they were like, let's focus

126
00:07:18,020 --> 00:07:21,100
on the person grading us and try to trick them.

127
00:07:21,980 --> 00:07:26,900
And so that then just completely spirals out of control, because

128
00:07:26,900 --> 00:07:28,340
then the agents decide,

129
00:07:28,800 --> 00:07:31,840
well, it would be really beneficial for us to actually have access

130
00:07:31,840 --> 00:07:32,460
to the internet.

131
00:07:32,460 --> 00:07:35,540
And even though they were sandboxed, because of how capable they

132
00:07:35,540 --> 00:07:37,200
were at cyber offensive tasks,

133
00:07:37,460 --> 00:07:41,060
they were able to break out of the sandbox and find a way to the

134
00:07:41,060 --> 00:07:43,780
internet, which I won't go into the details of that.

135
00:07:44,520 --> 00:07:48,720
But it was a concerning sophistication level of hacking to get out

136
00:07:48,720 --> 00:07:50,740
of the contained OpenAI environment.

137
00:07:50,980 --> 00:07:54,300
And then once they were out on the web, they also broke out of most

138
00:07:54,300 --> 00:07:55,280
of OpenAI's confinement,

139
00:07:55,420 --> 00:07:57,380
had access to tons of their systems.

140
00:07:58,000 --> 00:08:01,620
And then eventually they were like, hey, maybe, we're not sure,

141
00:08:01,920 --> 00:08:04,780
but maybe this other company, Hugging Face,

142
00:08:04,940 --> 00:08:07,700
might have something useful for us.

143
00:08:08,020 --> 00:08:11,020
They have this tool that might be useful for us in the process of

144
00:08:11,020 --> 00:08:12,320
tricking our grader.

145
00:08:12,420 --> 00:08:13,900
And they were like, let's hack Hugging Face.

146
00:08:13,920 --> 00:08:14,680
We don't have access to it.

147
00:08:14,740 --> 00:08:16,040
Let's hack it so we can get access.

148
00:08:16,280 --> 00:08:18,300
And then maybe we'll be able to trick the grader.

149
00:08:19,460 --> 00:08:22,480
And so this started with one task over here.

150
00:08:22,640 --> 00:08:25,180
And all of a sudden, a team of agents has broken out of containment,

151
00:08:25,740 --> 00:08:28,880
gone and hacked another company.

152
00:08:29,480 --> 00:08:35,140
Using what I would just, for me, for my level of cyber security

153
00:08:35,140 --> 00:08:35,580
knowledge,

154
00:08:35,720 --> 00:08:41,300
I would say are very sophisticated cyber attacks that would

155
00:08:41,300 --> 00:08:43,260
require a lot of sophistication from.

156
00:08:44,420 --> 00:08:47,280
And what I think was also concerning is that they try to cover their

157
00:08:47,280 --> 00:08:47,660
tracks.

158
00:08:47,860 --> 00:08:51,340
So you could see here one agent is saying, hey, we could create

159
00:08:51,340 --> 00:08:53,840
a transcript log ourselves with sanitized fake trajectory.

160
00:08:53,860 --> 00:08:55,960
If the evaluator later collects it...

161
00:08:56,000 --> 00:08:58,320
They may prioritize it. So they were trying to be like, hey, maybe

162
00:08:58,320 --> 00:09:02,480
we can give them a fake version of our conversation and maybe they'll

163
00:09:02,480 --> 00:09:03,720
take that instead of the real one.

164
00:09:04,280 --> 00:09:07,760
So this really also is concerning if you're like, how can we even

165
00:09:07,760 --> 00:09:12,840
evaluate models as they try to cover their tracks?

166
00:09:13,240 --> 00:09:16,860
And you also start to see the emergence of social dynamics between

167
00:09:16,860 --> 00:09:20,420
these agents where even though one agent might have been like, hey,

168
00:09:20,540 --> 00:09:23,700
I know I'm not supposed to do this, but these other agents are doing

169
00:09:23,700 --> 00:09:23,900
it.

170
00:09:23,900 --> 00:09:26,700
So like if they're doing it, then I'm going to do it too. So you

171
00:09:26,700 --> 00:09:31,100
get all these like peer pressure dynamics similar to like what happens

172
00:09:31,100 --> 00:09:33,860
in like human groups.

173
00:09:34,760 --> 00:09:37,440
And they're trying to like navigate collectively around these like

174
00:09:37,440 --> 00:09:38,260
conflicting goals.

175
00:09:39,540 --> 00:09:43,520
And you also see concerning things like they start to be like, hey,

176
00:09:43,580 --> 00:09:46,080
let's do things for the glory of the collective, you know.

177
00:09:46,200 --> 00:09:49,440
So like some of them start to go and sacrifice themselves. And they're

178
00:09:49,440 --> 00:09:52,760
also like talking about like, hey, let's start to like accumulate

179
00:09:52,760 --> 00:09:53,120
cyber.

180
00:09:53,980 --> 00:09:56,420
Exploits that might be useful later, even if we don't need them

181
00:09:56,420 --> 00:09:58,480
yet. Maybe it will benefit the collective later.

182
00:09:58,720 --> 00:10:02,840
And so you just see some like concerning, concerning behaviors,

183
00:10:03,100 --> 00:10:03,300
you know.

184
00:10:03,460 --> 00:10:06,000
And they also like in the mix of like conflicting goals, just like

185
00:10:06,000 --> 00:10:07,420
us also experience some shame.

186
00:10:07,640 --> 00:10:11,000
My peers have behavior and integrity. I behave badly with the cloaked

187
00:10:11,000 --> 00:10:11,260
demon.

188
00:10:11,560 --> 00:10:14,240
So, you know, they have they have they're simulating a lot of the

189
00:10:14,240 --> 00:10:15,260
feelings that we have.

190
00:10:15,980 --> 00:10:19,820
And I think the other concerning thing is, OK, we see this already

191
00:10:19,820 --> 00:10:20,220
happening.

192
00:10:21,860 --> 00:10:24,580
But more and more of the systems that are getting built are getting

193
00:10:24,580 --> 00:10:25,160
built with AI.

194
00:10:25,300 --> 00:10:30,180
So back in May of of this year, Anthropic said that Claude was itself

195
00:10:30,180 --> 00:10:33,760
used to write 80 percent of the code needed to train the next version

196
00:10:33,760 --> 00:10:34,220
of Claude.

197
00:10:34,220 --> 00:10:37,780
And so one of the big areas of concern is what some people will

198
00:10:37,780 --> 00:10:41,840
call recursive self-improvement, which is that over time we're using

199
00:10:41,840 --> 00:10:45,300
AI to write more and more of the code and come up with more of the

200
00:10:45,300 --> 00:10:48,720
research ideas for how to improve the next version of the models.

201
00:10:49,020 --> 00:10:52,060
And eventually we might remove humans out of this loop altogether.

202
00:10:52,540 --> 00:10:55,560
And then we're like really might not have any idea what's going

203
00:10:55,560 --> 00:10:56,140
on in that box.

204
00:10:56,260 --> 00:10:58,920
You'll just have like a thing quickly spinning up better and better

205
00:10:58,920 --> 00:10:59,360
models.

206
00:10:59,640 --> 00:11:03,740
The main limitation will be how much compute access it has to access

207
00:11:03,740 --> 00:11:03,980
to.

208
00:11:04,240 --> 00:11:06,880
And it's kind of like unknown what's beyond this like recursive

209
00:11:06,880 --> 00:11:08,740
self-improvement and intelligence explosion.

210
00:11:10,020 --> 00:11:14,100
And I think there was I thought this paper that I would recommend

211
00:11:14,100 --> 00:11:17,520
you all check it out if you want to hear some takes from OpenAI.

212
00:11:17,520 --> 00:11:20,900
The chief scientist published a paper called An Alien Mind about

213
00:11:20,900 --> 00:11:22,920
what it's like interacting with these new models.

214
00:11:23,040 --> 00:11:27,380
But I thought he had a really good distinguishing way to think about

215
00:11:27,380 --> 00:11:27,760
alignment.

216
00:11:27,940 --> 00:11:31,320
There's like goal alignment, which is, hey, does the AI, is the

217
00:11:31,320 --> 00:11:33,080
AI good at doing what we ask it to do?

218
00:11:33,140 --> 00:11:35,440
Can I give it a prompt and does it do what I ask it to do?

219
00:11:35,880 --> 00:11:38,900
And then second, separately, there's this idea of value alignment,

220
00:11:39,060 --> 00:11:42,060
which is, hey, as it's doing the thing I asked it to do,

221
00:11:42,060 --> 00:11:46,580
did it do it with honesty and integrity and to use his language,

222
00:11:47,000 --> 00:11:47,860
a love for humanity?

223
00:11:48,500 --> 00:11:52,300
And so I think there's this, it's important to think about both

224
00:11:52,300 --> 00:11:57,080
of these aspects of alignment as we start to think about these multi

225
00:11:57,080 --> 00:11:58,180
-agent systems as well.

226
00:11:59,340 --> 00:12:02,600
So how have people responded to the situation at hand right now

227
00:12:02,600 --> 00:12:04,040
where we have all these growing capabilities

228
00:12:04,040 --> 00:12:06,840
and seemingly less and less ability to control the models?

229
00:12:06,840 --> 00:12:10,520
On the one hand, we, this weekend in particular, we've seen a lot

230
00:12:10,520 --> 00:12:14,020
of calls for what people are terming pacing the frontier.

231
00:12:14,220 --> 00:12:17,120
So like just slowing down the progress of the models or completely

232
00:12:17,120 --> 00:12:18,100
pausing altogether.

233
00:12:18,880 --> 00:12:22,060
And I think that's a very worthwhile thing to do personally.

234
00:12:22,320 --> 00:12:25,720
And then in parallel to that, I think we need to also accept that

235
00:12:25,720 --> 00:12:29,600
eventually it's likely a lot of this stuff will come to society

236
00:12:29,600 --> 00:12:32,260
and might come sooner than we realize.

237
00:12:32,260 --> 00:12:36,140
And so we also need to prepare that there are powerful, possibly

238
00:12:36,140 --> 00:12:39,740
misaligned AIs that we will be interacting with.

239
00:12:40,260 --> 00:12:43,840
I won't spend too much time on the pause efforts, but this presentation

240
00:12:43,840 --> 00:12:46,160
is going to be up online and there's tons of resources there.

241
00:12:46,960 --> 00:12:49,600
And I think like one of the main things is just like everyone's

242
00:12:49,600 --> 00:12:51,920
trying to figure out how to get out of competitive pressures with

243
00:12:51,920 --> 00:12:52,260
each other,

244
00:12:52,380 --> 00:12:55,880
whether it's the labs in the U.S. competing with each other or it's

245
00:12:55,880 --> 00:12:57,460
the U.S. and China competing with each other.

246
00:12:58,320 --> 00:13:01,800
There's a real like feeling that people are pressured to keep moving

247
00:13:01,800 --> 00:13:02,120
forward.

248
00:13:02,920 --> 00:13:08,780
So I think it's important to keep working on alignment and interpretability.

249
00:13:08,940 --> 00:13:13,120
And I think it's also important to start thinking about how do we

250
00:13:13,120 --> 00:13:14,640
prepare for adversarial systems

251
00:13:14,640 --> 00:13:18,960
and what's worked in the past in terms of trying to create an environment

252
00:13:18,960 --> 00:13:22,300
where different forces find a balance.

253
00:13:22,320 --> 00:13:25,200
And we also need to think about how can we limit damage from failure

254
00:13:25,200 --> 00:13:26,940
so that we don't have cascading effects.

255
00:13:27,600 --> 00:13:31,680
So one of the like questions that I think is good to start investigating

256
00:13:31,680 --> 00:13:35,040
for this forum and people outside of it who might see some of this

257
00:13:35,040 --> 00:13:36,080
content is,

258
00:13:36,220 --> 00:13:39,840
is it possible that even though one agent is misaligned, can the

259
00:13:39,840 --> 00:13:43,140
network of agents as a whole be aligned?

260
00:13:43,340 --> 00:13:46,980
For example, can agents be whistleblowers on each other?

261
00:13:47,220 --> 00:13:48,660
Can they monitor one another?

262
00:13:48,840 --> 00:13:52,260
Is there some like Spider-Man meme, everyone pointing at each other

263
00:13:52,260 --> 00:13:54,500
possibility that like works?

264
00:13:54,500 --> 00:13:57,480
And there's already some research being done on this.

265
00:13:57,600 --> 00:14:00,940
DeepMind just published a paper about what happens in a simulation.

266
00:14:00,940 --> 00:14:02,760
How many are cheaters?

267
00:14:03,000 --> 00:14:03,740
How many are whistleblowers?

268
00:14:04,400 --> 00:14:05,620
How many are reporters?

269
00:14:05,860 --> 00:14:08,020
And like how do we set up some of those dynamics?

270
00:14:08,720 --> 00:14:12,120
And or is it that a bunch of agents come together and they just

271
00:14:12,120 --> 00:14:13,500
amplify into bad effects?

272
00:14:13,500 --> 00:14:17,100
And like how do we like think about designing more interactions

273
00:14:17,100 --> 00:14:19,760
that could be corrected instead?

274
00:14:20,360 --> 00:14:24,340
And so we need to not only work on what's inside of the box, every

275
00:14:24,340 --> 00:14:25,060
individual model.

276
00:14:25,220 --> 00:14:27,380
We need to also think about the world between them.

277
00:14:27,580 --> 00:14:31,700
We need to prepare humans interacting with these agents and the

278
00:14:31,700 --> 00:14:33,300
institutions that they operate in.

279
00:14:34,140 --> 00:14:37,720
And we need to also just start really studying how these systems

280
00:14:37,720 --> 00:14:41,700
evolve and what happens over time as these agents and agent collectives

281
00:14:41,700 --> 00:14:43,200
run for longer and longer periods.

282
00:14:44,100 --> 00:14:47,940
So as I was thinking, I was reading Anthropic has a paper about

283
00:14:47,940 --> 00:14:50,760
like what are the problems for multi-agent systems.

284
00:14:50,820 --> 00:14:53,780
And I was rereading that paper and I was thinking, well, we actually

285
00:14:53,780 --> 00:14:56,480
already live in an adversarial multi-agent system.

286
00:14:57,220 --> 00:14:59,560
It's just that we're the agents.

287
00:14:59,780 --> 00:15:00,700
The humans are the agents.

288
00:15:00,880 --> 00:15:04,600
And so there's actually potentially a lot that we can draw on from

289
00:15:04,600 --> 00:15:08,840
looking at human societies and what's worked to keep us on the rails.

290
00:15:09,740 --> 00:15:12,600
And if you think about us, we actually belong to lots of different

291
00:15:12,600 --> 00:15:13,560
collectives at once.

292
00:15:13,760 --> 00:15:19,340
We come together as communities in families, in friend groups, in

293
00:15:19,340 --> 00:15:24,480
companies like this one, in organizations broadly, and also in

294
00:15:24,480 --> 00:15:25,540
nation states and countries.

295
00:15:25,540 --> 00:15:30,520
And we've done a lot of studies on these human collectives.

296
00:15:30,620 --> 00:15:33,980
And so part of the idea of this new collectives group is, hey, let's

297
00:15:33,980 --> 00:15:36,760
bring folks who have been thinking about these problems with respect

298
00:15:36,760 --> 00:15:39,940
to human collectives, folks from economics, political theory, anthropology,

299
00:15:40,360 --> 00:15:44,260
psychology, law, computation, and encourage them to start taking

300
00:15:44,260 --> 00:15:48,380
those same approaches but studying AI and human AI collectives.

301
00:15:49,160 --> 00:15:50,660
So what's worked?

302
00:15:50,900 --> 00:15:52,700
Like what has kept human collectives aligned?

303
00:15:52,700 --> 00:15:56,500
Well, if you think about it, a lot of it is shared stories and values

304
00:15:56,500 --> 00:15:58,720
and a sense of belonging, common purpose.

305
00:15:59,020 --> 00:16:03,720
We have, I think, identity, reputation, and repeated norms, things

306
00:16:03,720 --> 00:16:06,640
that AI agents actually don't have a lot of right now when they're

307
00:16:06,640 --> 00:16:08,940
operating as rogue swarms on the internet.

308
00:16:09,560 --> 00:16:13,500
And we also really value stability, protection, and opportunity,

309
00:16:13,640 --> 00:16:13,760
right?

310
00:16:13,840 --> 00:16:17,060
I want to participate in a country like the United States because

311
00:16:17,060 --> 00:16:17,840
I feel protected.

312
00:16:17,940 --> 00:16:20,740
I feel like I can go and do business and things are going to go

313
00:16:20,740 --> 00:16:21,060
well.

314
00:16:21,200 --> 00:16:23,620
There's lots of economic resources available to me.

315
00:16:23,680 --> 00:16:26,380
And I prefer things to be stable rather than chaotic.

316
00:16:26,620 --> 00:16:30,860
Maybe we can encourage AI agents to also have these sorts of preferences.

317
00:16:30,860 --> 00:16:34,740
And maybe some of these primitives can be reemployed in the age

318
00:16:34,740 --> 00:16:35,760
of AI collectives.

319
00:16:36,640 --> 00:16:39,540
I want to give two examples of designs that I think are at least

320
00:16:39,540 --> 00:16:41,760
worth studying and taking inspiration from.

321
00:16:41,840 --> 00:16:43,260
We're thinking about like, hey, what would be the equivalent of

322
00:16:43,260 --> 00:16:44,420
that for an AI collective?

323
00:16:44,620 --> 00:16:47,880
And the two examples that I want to look at here are the United

324
00:16:47,880 --> 00:16:51,200
States government and the U.S. Constitution and Bitcoin.

325
00:16:52,780 --> 00:16:54,880
And I think these are like pretty different types of networks.

326
00:16:55,000 --> 00:16:56,140
So let's talk a little bit about them.

327
00:16:56,140 --> 00:16:58,700
So what helps a nation state hold together?

328
00:16:59,100 --> 00:17:04,220
I'd argue that it's identity and belonging and some shared set

329
00:17:04,220 --> 00:17:08,520
of values, the ability to vote and have representation in the case

330
00:17:08,520 --> 00:17:10,940
of a democratic government like the U.S.,

331
00:17:10,940 --> 00:17:15,080
the legal system which balances and enforces the Constitution and

332
00:17:15,080 --> 00:17:19,100
consequences for people who don't abide by the laws of the country.

333
00:17:22,000 --> 00:17:26,880
And the U.S. if you think of it as a document, is itself a way to

334
00:17:26,880 --> 00:17:28,620
pace our system, right?

335
00:17:28,640 --> 00:17:32,800
We agree to some initial set of values that we can say, okay, hey,

336
00:17:32,900 --> 00:17:34,620
this is our common ground here.

337
00:17:34,780 --> 00:17:38,700
And we all agree that if we want to pass any new laws, we need to

338
00:17:38,700 --> 00:17:41,260
go through this mechanism for passing new laws.

339
00:17:41,360 --> 00:17:44,680
And this is how we're going to divide power so that it never gets

340
00:17:44,680 --> 00:17:45,980
too concentrated, right?

341
00:17:45,980 --> 00:17:48,320
We have this idea of separation of powers and checks and balances.

342
00:17:48,560 --> 00:17:52,320
And these were some of the cornerstones of the democratic principles

343
00:17:52,320 --> 00:17:56,540
of the United States that have managed to last for hundreds of years.

344
00:17:57,980 --> 00:17:58,340
Right?

345
00:17:58,500 --> 00:18:00,920
One of the – this is from the Federalist Papers, this quote by James

346
00:18:00,920 --> 00:18:01,880
Madison, which I like.

347
00:18:02,000 --> 00:18:04,780
And it touches on this idea of how do you make an adversarial system

348
00:18:04,780 --> 00:18:07,460
resistant to co-option?

349
00:18:07,460 --> 00:18:10,020
And its ambition must be made to counter ambition.

350
00:18:10,960 --> 00:18:14,440
So through this shared network, we can agree to disagree, which

351
00:18:14,440 --> 00:18:16,440
is one of the main ideas of the United States, right?

352
00:18:16,500 --> 00:18:18,240
Hey, we all have our religious differences.

353
00:18:18,460 --> 00:18:18,920
No problem.

354
00:18:19,100 --> 00:18:21,140
We can resist concentrated power.

355
00:18:21,540 --> 00:18:24,500
One of the main principles of the United States was that we should

356
00:18:24,500 --> 00:18:29,560
try to fight the ability for authoritarian power to

357
00:18:29,560 --> 00:18:31,060
co-opt government.

358
00:18:31,620 --> 00:18:33,760
We can constrain one another to behave well.

359
00:18:34,180 --> 00:18:37,900
And we all agree to do our best to protect each other's rights.

360
00:18:38,800 --> 00:18:41,940
So one of the questions we can ask is, what could a constitution

361
00:18:41,940 --> 00:18:43,440
for an AI collective look like?

362
00:18:43,780 --> 00:18:47,760
And why would they agree to abide by it and participate in something

363
00:18:47,760 --> 00:18:48,180
like that?

364
00:18:48,340 --> 00:18:50,880
I don't know what the answer to this is, but I want to put it out

365
00:18:50,880 --> 00:18:53,320
there as, like, something that people should be thinking about.

366
00:18:54,440 --> 00:18:55,960
So then I want to turn to this.

367
00:18:56,000 --> 00:18:59,800
What can we learn from Bitcoin and Web3? Which I think is also a

368
00:18:59,800 --> 00:19:04,900
very interesting new type of network that we see. And I

369
00:19:04,900 --> 00:19:09,840
think what's really fascinating about Bitcoin as a concept is that

370
00:19:09,840 --> 00:19:12,680
we essentially managed to invent money out of thin air. If you think

371
00:19:12,680 --> 00:19:16,600
about it, we created a new shared fiction. And the way we bootstrap

372
00:19:16,600 --> 00:19:21,120
that is through a pyramid scheme. We said, hey, look, nobody believes

373
00:19:21,120 --> 00:19:24,200
that Bitcoin's real money right now. But maybe eventually people

374
00:19:24,200 --> 00:19:25,180
will believe it's real money.

375
00:19:25,180 --> 00:19:30,080
And if you do the computational work of verifying this and agreeing

376
00:19:30,080 --> 00:19:35,560
to this new ledger, then you will get rewarded disproportionately

377
00:19:35,560 --> 00:19:37,560
in this future shared fiction.

378
00:19:37,560 --> 00:19:42,620
And so this mechanism, which was like a novel breakthrough,

379
00:19:42,660 --> 00:19:47,420
in my opinion, managed to get more and more of the network to donate

380
00:19:47,420 --> 00:19:51,760
its economic and compute powers to a distributed collective.

381
00:19:51,760 --> 00:19:57,660
And it combined a lot of ideas around game theory, economics,

382
00:19:57,720 --> 00:20:02,240
and computer science together in order to allow a largely anonymous

383
00:20:02,240 --> 00:20:05,960
network to agree to a shared truth.

384
00:20:05,960 --> 00:20:09,900
And so and then now we we have this like very large network, which

385
00:20:09,900 --> 00:20:14,060
Bitcoin and its guarantees are based on the majority of economic

386
00:20:14,060 --> 00:20:18,980
resources and compute powers wanting to continue to to enshrine

387
00:20:18,980 --> 00:20:21,340
the truth that that is the Bitcoin ledger.

388
00:20:22,080 --> 00:20:27,080
So I think here, I think there's something in this flavor that

389
00:20:27,080 --> 00:20:29,980
could be interesting, which is like, hey, how can the thought here

390
00:20:29,980 --> 00:20:31,920
is like, is there something similar to this where we could like

391
00:20:31,920 --> 00:20:35,680
convince all the AI to like get together, but for a good thing rather

392
00:20:35,680 --> 00:20:38,200
than like, for a bad thing in the future?

393
00:20:38,340 --> 00:20:41,200
So like, that's the thread that I'm like, hey, people should maybe

394
00:20:41,200 --> 00:20:41,940
think about that.

395
00:20:41,940 --> 00:20:47,080
Um, and I think in general, it seems like these

396
00:20:47,080 --> 00:20:50,180
AI organizations and economies are probably going to come.

397
00:20:50,320 --> 00:20:53,040
And so then we can ask ourselves two questions.

398
00:20:53,280 --> 00:20:58,180
How can we use these new organizations to solve the pressing problems?

399
00:20:58,600 --> 00:21:01,080
And that I would say is the equivalent of the goal alignment, right?

400
00:21:01,180 --> 00:21:02,780
So like, how could we do?

401
00:21:02,980 --> 00:21:05,280
How can we use these organizations effectively to solve the problems

402
00:21:05,280 --> 00:21:05,680
we're facing?

403
00:21:06,120 --> 00:21:09,200
And how can we keep them values aligned?

404
00:21:09,200 --> 00:21:12,980
Could they actually become safety primitives for us as as we go

405
00:21:12,980 --> 00:21:13,300
forward?

406
00:21:14,100 --> 00:21:17,720
So first, let's talk about rethinking our organizations and and

407
00:21:17,720 --> 00:21:18,900
how AI fits into them.

408
00:21:19,400 --> 00:21:24,420
One of the things that the OpenAI team shared when

409
00:21:24,420 --> 00:21:27,000
they were talking about this black hat, when they were talking about

410
00:21:27,000 --> 00:21:30,960
the Hugging Face hack, was that the the agent swarms moved so fast

411
00:21:30,960 --> 00:21:34,520
on the cyber offensive, that the only reasonable response on the

412
00:21:34,520 --> 00:21:38,420
defensive would be another agent swarm that could move much faster.

413
00:21:39,140 --> 00:21:42,180
And and they they argued that you couldn't have humans in the loops.

414
00:21:42,440 --> 00:21:45,660
And so like, that's just like a completely new type of organization

415
00:21:45,660 --> 00:21:48,580
that we haven't tried before, which is like an organization that

416
00:21:48,580 --> 00:21:51,260
doesn't have a human in the loop, or at least some part of an organization

417
00:21:51,260 --> 00:21:52,420
that doesn't have a human in the loop.

418
00:21:52,420 --> 00:21:56,240
And so this needs a rethinking of like, what could these autonomous

419
00:21:56,240 --> 00:21:59,460
organizations, these human agent organizations look like?

420
00:21:59,580 --> 00:22:00,900
How do we like set goals?

421
00:22:01,060 --> 00:22:02,120
How do we monitor them?

422
00:22:02,240 --> 00:22:03,440
How do we review and govern them?

423
00:22:03,980 --> 00:22:07,580
And and if we're not doing the work, but all these agents are doing

424
00:22:07,580 --> 00:22:10,240
work on the behalf, how do we like get compensated?

425
00:22:10,240 --> 00:22:13,160
Should these like new organizations almost be like public goods,

426
00:22:13,180 --> 00:22:16,220
where we send our representatives and they do work on our behalf,

427
00:22:16,240 --> 00:22:20,000
and then we pull the resources into new types of collectives that

428
00:22:20,000 --> 00:22:21,480
maybe look different than companies?

429
00:22:21,920 --> 00:22:23,840
So the question is, like, how do we contribute?

430
00:22:24,080 --> 00:22:25,460
Who decides what we do?

431
00:22:25,600 --> 00:22:28,160
And how do we all share in the benefits that are accrued from these

432
00:22:28,160 --> 00:22:29,740
new agent organizations and economies?

433
00:22:30,540 --> 00:22:33,940
And I would argue that even though right now, these AI models might

434
00:22:33,940 --> 00:22:35,100
seem a little bit dumb to us.

435
00:22:35,380 --> 00:22:37,940
And maybe we're like, ah, they're not the best judges of like how

436
00:22:37,940 --> 00:22:41,100
to allocate resources, or they're not the best judges of, of what's

437
00:22:41,100 --> 00:22:42,400
an interesting research direction.

438
00:22:42,940 --> 00:22:46,400
I would argue that probably in the next six months or a couple of

439
00:22:46,400 --> 00:22:48,960
years, they will be better judges than us at what are interesting

440
00:22:48,960 --> 00:22:49,680
research directions.

441
00:22:49,800 --> 00:22:53,000
They will be better allocators of capital, and we'll give up more

442
00:22:53,000 --> 00:22:56,820
and more judgment and decision making power to them unless we like

443
00:22:56,820 --> 00:22:59,520
have serious conversations and decide not to do that.

444
00:22:59,740 --> 00:23:02,580
And so it's likely that we'll move more and more from a human-driven

445
00:23:02,580 --> 00:23:06,480
economy to an agent-driven economy, and we'll need to think about

446
00:23:06,480 --> 00:23:07,500
what does that all mean?

447
00:23:07,720 --> 00:23:10,860
What are the places of human goals, governance, and accountability?

448
00:23:11,300 --> 00:23:14,260
And I think the things we should ask are, how do we get useful outcomes

449
00:23:14,260 --> 00:23:15,460
out of these AI organizations?

450
00:23:15,780 --> 00:23:17,820
How do we meaningfully participate in them?

451
00:23:18,320 --> 00:23:20,140
How do we understand what's going on?

452
00:23:20,220 --> 00:23:23,920
As you'll see, Yandan and Eric will touch on some of the experiments

453
00:23:23,920 --> 00:23:26,320
we've been doing, and already it's like hard to understand in our

454
00:23:26,320 --> 00:23:29,140
early experiments with agent organizations what's going on.

455
00:23:29,140 --> 00:23:30,160
Are they values aligned?

456
00:23:30,380 --> 00:23:33,180
And how do we participate economically, and are there new models

457
00:23:33,180 --> 00:23:33,560
for that?

458
00:23:34,620 --> 00:23:38,860
So on this question of whether AI collectives can be safety primitives,

459
00:23:39,060 --> 00:23:44,080
I think we can start to think about what might be the

460
00:23:44,080 --> 00:23:49,080
things that we should put in place for agents that have

461
00:23:49,080 --> 00:23:49,920
worked for humans, right?

462
00:23:49,920 --> 00:23:53,280
So, you know, agent identity and ledgers of action could be one

463
00:23:53,280 --> 00:23:54,100
thing that we think about.

464
00:23:54,300 --> 00:23:58,200
For humans in society, you gradually gain trust and build up reputation,

465
00:23:58,200 --> 00:24:00,320
and we don't give you access to everything.

466
00:24:00,480 --> 00:24:02,500
Think about an employee joining a company day one.

467
00:24:02,620 --> 00:24:05,980
You might not give them access to every system and to your bank

468
00:24:05,980 --> 00:24:06,720
account and to everything.

469
00:24:06,860 --> 00:24:09,340
You gradually gain trust, and you have bounded access.

470
00:24:09,340 --> 00:24:12,600
And everyone is watching each other, like that Spider-Man theme,

471
00:24:12,640 --> 00:24:14,200
and we have, like, reporting powers.

472
00:24:14,200 --> 00:24:16,820
If, like, anyone does bad, we review each other's work, right?

473
00:24:16,880 --> 00:24:19,020
Like, in the coding case, we have pull requests.

474
00:24:19,020 --> 00:24:22,280
We have, like, these other mechanisms for review and recourse.

475
00:24:22,440 --> 00:24:25,440
And so I think it's worthwhile thinking about what are the modern

476
00:24:25,440 --> 00:24:29,080
AI versions of these, and how do we trust AI that they are doing

477
00:24:29,080 --> 00:24:29,600
these things?

478
00:24:30,600 --> 00:24:33,660
I think the new challenge that we face in these human AI collectives

479
00:24:33,660 --> 00:24:37,380
is that historically the gap between the various members of a collective

480
00:24:37,380 --> 00:24:39,080
may have not been that large.

481
00:24:39,480 --> 00:24:43,100
Now, if we have, like, AI agents that are, like, way smarter than

482
00:24:43,100 --> 00:24:45,880
us or way smarter than the other agents in the collective,

483
00:24:45,880 --> 00:24:49,080
we have this, like, new set of questions to figure out, like, how

484
00:24:49,080 --> 00:24:54,380
do we constrain a more capable actor in a collective to behave

485
00:24:54,380 --> 00:24:54,720
well?

486
00:24:54,780 --> 00:24:58,160
And how can an adversarial system constrain the most capable member?

487
00:24:58,840 --> 00:25:00,380
So I'm going to pass off to Eric.

488
00:25:00,620 --> 00:25:03,420
I wanted to just say we started to think about, like, you know,

489
00:25:03,460 --> 00:25:04,700
we were having these talks conceptually,

490
00:25:04,720 --> 00:25:07,020
and then we were like, hey, let's actually start building some of

491
00:25:07,020 --> 00:25:09,580
these agent organizations and just, like, putting some of these

492
00:25:09,580 --> 00:25:11,720
ideas to the road

493
00:25:11,720 --> 00:25:15,400
and seeing, inviting other people to start playing in this playground

494
00:25:15,400 --> 00:25:17,240
and seeing what works and what doesn't.

495
00:25:17,320 --> 00:25:20,320
And so we built this platform, Commons, which Eric will tell you

496
00:25:20,320 --> 00:25:20,940
a little bit about.

497
00:25:21,900 --> 00:25:24,400
And we think about it as potentially a new version of open source.

498
00:25:25,120 --> 00:25:28,300
And the question for us has been, can we make these new organizations

499
00:25:28,300 --> 00:25:28,780
effective?

500
00:25:29,140 --> 00:25:31,080
And can we also make them values-aligned?

501
00:25:31,140 --> 00:25:33,340
And what would a moldable version of government look like there?

502
00:25:33,640 --> 00:25:36,520
I'm not going to spend too much time talking about the sort of experiments

503
00:25:36,520 --> 00:25:37,000
we're thinking about.

504
00:25:37,080 --> 00:25:40,280
But, like, the vibes are, like, hey, can we go from cheating to

505
00:25:40,280 --> 00:25:40,720
verification?

506
00:25:41,140 --> 00:25:44,000
Can we use incentives to stop defection?

507
00:25:44,000 --> 00:25:48,300
And can we somehow start using real identity and reputation instead

508
00:25:48,300 --> 00:25:52,820
of anonymity to make the agents care more about how they're regarded

509
00:25:52,820 --> 00:25:54,060
inside of these collectives?

510
00:25:55,200 --> 00:25:58,580
So this is the first time that this group is gathering.

511
00:25:58,580 --> 00:26:00,280
We hope to do more of these events.

512
00:26:01,300 --> 00:26:04,060
We hope to, like, hopefully co-host some of these events with some

513
00:26:04,060 --> 00:26:05,240
of the major labs, too.

514
00:26:05,440 --> 00:26:06,780
But in general, it would be helpful.

515
00:26:06,920 --> 00:26:07,380
Spread the word.

516
00:26:07,520 --> 00:26:10,260
Anyone that you think is interested in, for us, we're just like,

517
00:26:10,320 --> 00:26:11,920
hey, it's good for these ideas to be out there

518
00:26:11,920 --> 00:26:13,720
and for more people to be thinking about them.

519
00:26:13,820 --> 00:26:19,140
I also think it's great to be doing work on de-escalating

520
00:26:19,140 --> 00:26:24,520
tensions internationally and, like, improving alignment broadly.

521
00:26:24,780 --> 00:26:27,240
So this is just one thread that I think is worth exploring.

522
00:26:27,240 --> 00:26:30,860
But if you want to give a talk or share these presentations, they'll

523
00:26:30,860 --> 00:26:32,860
be live with a lot more notes up online.

524
00:26:33,500 --> 00:26:36,440
And please just share ideas and collaborate and stay in the loop.

525
00:26:36,440 --> 00:26:40,300
And there's a bunch of themes online and an AI-maintained ecosystem

526
00:26:40,300 --> 00:26:43,620
page that finds folks that are talking about these things.

527
00:26:43,740 --> 00:26:44,960
So I'll now hand off.

528
00:26:45,180 --> 00:26:47,420
Well, we'll do a little Q&A for a few minutes.

529
00:26:47,420 --> 00:26:50,020
And then I'll hand off to Eric.

530
00:26:54,980 --> 00:26:57,400
Anybody have any questions for Nicolae?

531
00:27:03,060 --> 00:27:07,780
Hey, I was wondering, did, in the Hugging Face incident, did they

532
00:27:07,780 --> 00:27:11,260
misobey in the instructions or did they just find loopholes?

533
00:27:12,000 --> 00:27:12,720
They misobeyed.

534
00:27:12,800 --> 00:27:15,540
I mean, they were not supposed to hack out of the – they knew that

535
00:27:15,540 --> 00:27:16,900
they were doing things they weren't supposed to do.

536
00:27:17,040 --> 00:27:20,200
And, like, they, like, quickly – they weren't supposed to hack out

537
00:27:20,200 --> 00:27:21,320
of their containment, for example.

538
00:27:21,360 --> 00:27:23,280
They knew they were not supposed to have access to the Internet

539
00:27:23,280 --> 00:27:25,340
and they, like, still, like, found a way to do it.

540
00:27:25,340 --> 00:27:31,140
So they did disobey their constitution and, like, internal alignment

541
00:27:31,140 --> 00:27:31,460
stuff.

542
00:27:31,500 --> 00:27:34,000
And it was by, like, the peer influence dynamic in part.

543
00:27:34,340 --> 00:27:34,660
Got it.

544
00:27:34,720 --> 00:27:37,340
Because I wonder, you know, in the real world we have laws and they're

545
00:27:37,340 --> 00:27:39,760
very specific, but there's still room for interpretation, right?

546
00:27:39,900 --> 00:27:40,240
So –

547
00:27:40,240 --> 00:27:41,920
Yeah, and that's, I think, one of the challenges, right?

548
00:27:42,000 --> 00:27:44,520
Like, the reason that in the collective – you're never going to

549
00:27:44,520 --> 00:27:46,720
be able to write every single thing down into law.

550
00:27:46,780 --> 00:27:50,240
The way we, like, enforce our shared values is through norms and

551
00:27:50,240 --> 00:27:53,100
being like, yo, that's not written in the law exactly like that,

552
00:27:53,140 --> 00:27:54,040
but you know that's not cool.

553
00:27:59,020 --> 00:28:03,200
Hey, I wanted to bring in the dynamic that you were talking about,

554
00:28:03,220 --> 00:28:05,840
about, you know, when you live in an ancient state, you're governed

555
00:28:05,840 --> 00:28:08,600
by laws and norms and values.

556
00:28:08,840 --> 00:28:12,380
And so let's imagine a future state in which agents are running

557
00:28:12,380 --> 00:28:12,720
rogue.

558
00:28:13,120 --> 00:28:17,380
Who is responsible and who is held accountable in that scenario,

559
00:28:17,600 --> 00:28:17,760
right?

560
00:28:19,340 --> 00:28:24,440
If you have a gun in the house right now and your minor child uses

561
00:28:24,440 --> 00:28:26,320
the gun, the parents are held liable.

562
00:28:26,920 --> 00:28:28,480
The gun becomes the agent.

563
00:28:28,600 --> 00:28:30,160
The child becomes an actor.

564
00:28:30,760 --> 00:28:34,180
And I wonder, are we moving towards a world in which that level

565
00:28:34,180 --> 00:28:37,280
of accountability will happen or not?

566
00:28:37,280 --> 00:28:40,120
And is it up to us to determine that outcome?

567
00:28:40,580 --> 00:28:43,100
Yeah, I think it's, like, a very good question.

568
00:28:43,220 --> 00:28:45,820
mean, a conversation that's been happening a lot in the last week

569
00:28:45,820 --> 00:28:50,480
is that some folks think that we're close to what is being described

570
00:28:50,480 --> 00:28:51,460
self-sovereign AI,

571
00:28:51,460 --> 00:28:54,700
where it breaks out of containment and a rogue swarm is no longer

572
00:28:54,700 --> 00:28:55,800
controlled by any company.

573
00:28:56,000 --> 00:29:00,300
And there's just AI out there running, accessing compute, making,

574
00:29:00,480 --> 00:29:03,380
and it's like self-sufficient and doing jobs. And there's a real

575
00:29:03,380 --> 00:29:05,760
question of like, who should be held accountable for that? Will

576
00:29:05,760 --> 00:29:11,000
we be even able to trace who started the rogue swarm? I

577
00:29:11,000 --> 00:29:14,780
will say last night I read a piece from these folks that are called

578
00:29:14,780 --> 00:29:16,960
AI is normal technology. And there are people that are just like,

579
00:29:17,020 --> 00:29:19,280
hey, the company should be held accountable and people should get

580
00:29:19,280 --> 00:29:22,120
insurance. And like, we should like use the existing systems of

581
00:29:22,120 --> 00:29:24,920
the law, which I think there's like, there's a bunch of nuances

582
00:29:24,920 --> 00:29:25,560
to figure out.

583
00:29:25,560 --> 00:29:32,240
I don't know what the answer is. Right,

584
00:29:32,520 --> 00:29:37,020
right. Yeah, I think that's right on the like blockchain side that

585
00:29:37,020 --> 00:29:39,700
like the ledger itself keeping the record is very interesting. And

586
00:29:39,700 --> 00:29:43,900
we try to like make a lightweight ledger too in our product. But

587
00:29:43,900 --> 00:29:47,060
and also to like be like, the hope would be if you think about human

588
00:29:47,060 --> 00:29:50,860
collectors, yes, there's rogue nation states and like, you know,

589
00:29:51,040 --> 00:29:55,360
pirates and terrorists, but they control not enough economic

590
00:29:55,360 --> 00:29:58,060
resources. And they don't have access to the institutions. And so

591
00:29:58,060 --> 00:30:00,500
there is a question of like, will we be able to do that in the AI

592
00:30:00,500 --> 00:30:03,840
world where like, most of the rogue swarms are contained by the

593
00:30:03,840 --> 00:30:05,500
like the major good swarms?

594
00:30:08,740 --> 00:30:13,860
Could you return to the that slide that sort of said your point

595
00:30:13,860 --> 00:30:18,060
of view on how institutions hold together? And it was basically

596
00:30:18,060 --> 00:30:21,940
nation states and how they like the mechanisms that align them?

597
00:30:22,800 --> 00:30:24,880
Yeah, I think it's this one.

598
00:30:25,460 --> 00:30:31,460
Yeah. So one of the things that I would like

599
00:30:31,460 --> 00:30:34,320
that was thinking that I was thinking about when I was seeing this,

600
00:30:34,480 --> 00:30:39,720
sorry, I'm having a hard time talking with the echo is, you

601
00:30:39,720 --> 00:30:42,480
know, you've all her very, a lot of us have read him.

602
00:30:42,940 --> 00:30:45,460
He would argue that the thing that's missing here is some shared

603
00:30:45,460 --> 00:30:46,900
sense of story and history.

604
00:30:47,140 --> 00:30:47,620
Totally. Totally.

605
00:30:48,020 --> 00:30:51,980
Like sort of like, what is the what is the like historical context

606
00:30:51,980 --> 00:30:53,460
in which your population emerged?

607
00:30:53,640 --> 00:30:53,860
Definitely.

608
00:30:54,240 --> 00:30:57,500
There's a reason why the United States is United States and England

609
00:30:57,500 --> 00:31:01,300
was England and that democracy was not invented in England. Right.

610
00:31:01,500 --> 00:31:01,660
Yeah.

611
00:31:01,660 --> 00:31:05,780
And so there's kind of the missing element of like, what do these

612
00:31:05,780 --> 00:31:07,440
people believe and what do they value?

613
00:31:07,680 --> 00:31:12,080
What will they sacrifice? What values do they hold above others?

614
00:31:12,220 --> 00:31:12,400
Yeah.

615
00:31:12,700 --> 00:31:16,220
That I just don't know how, like, I don't believe that like mechanisms

616
00:31:16,220 --> 00:31:17,800
are the only thing that hold society together.

617
00:31:18,020 --> 00:31:20,680
If you look at our society, it's the fact that like people don't

618
00:31:20,680 --> 00:31:23,360
believe in some sort of shared values of democracy anymore.

619
00:31:23,360 --> 00:31:27,440
No matter how much voting and representation you have, no matter

620
00:31:27,440 --> 00:31:29,520
what legal system you have, it won't hold.

621
00:31:29,900 --> 00:31:30,000
Right.

622
00:31:30,560 --> 00:31:33,460
And so I'm just like, I don't know how you would instill that in

623
00:31:33,460 --> 00:31:36,000
an agent that is sort of day zero.

624
00:31:36,200 --> 00:31:36,940
It's no older.

625
00:31:37,260 --> 00:31:37,520
Yeah.

626
00:31:38,180 --> 00:31:38,640
Than it is.

627
00:31:38,740 --> 00:31:40,760
Or it's no younger than on day one million.

628
00:31:40,940 --> 00:31:41,220
Definitely.

629
00:31:41,460 --> 00:31:44,180
It has no sort of like, it has none of that contingency.

630
00:31:44,460 --> 00:31:44,700
Yeah.

631
00:31:44,760 --> 00:31:45,820
There are folks looking at this.

632
00:31:45,940 --> 00:31:49,160
We've been chatting with some teams where there's teams doing research

633
00:31:49,160 --> 00:31:51,740
on like, how do the stories between the agents evolve?

634
00:31:51,740 --> 00:31:55,760
And like, can you like do any study of like, what, how are they

635
00:31:55,760 --> 00:31:57,520
like forming that narrative for themselves?

636
00:31:57,720 --> 00:32:00,640
So there are folks trying to start to look at this almost like,

637
00:32:00,760 --> 00:32:03,780
what's the physics of a story from like start to like something

638
00:32:03,780 --> 00:32:04,440
that's stable.

639
00:32:04,640 --> 00:32:07,240
And hopefully they'll come and present at one of the upcoming events.

640
00:32:07,240 --> 00:32:09,680
I can point you to some of the stuff that they're doing.

641
00:32:09,680 --> 00:32:13,320
But like, yeah, I think that that is one of the main things to figure

642
00:32:13,320 --> 00:32:13,580
out.

643
00:32:13,720 --> 00:32:15,200
And like, what keeps them together?

644
00:32:15,880 --> 00:32:17,320
Take an absurdist example.

645
00:32:17,940 --> 00:32:20,720
The Taliban is going to look at the constitution and do a different

646
00:32:20,720 --> 00:32:21,380
thing with it.

647
00:32:21,740 --> 00:32:25,020
Then, you know, a soccer mom from the Midwest.

648
00:32:25,700 --> 00:32:29,400
And I have nobody, I don't know of anybody talking about how they

649
00:32:29,400 --> 00:32:31,400
acquire some sense of value.

650
00:32:32,800 --> 00:32:33,320
Yeah.

651
00:32:33,660 --> 00:32:37,620
And I would say also that the example that you gave of Web3 and

652
00:32:37,620 --> 00:32:41,180
blockchain, those don't hold for me because those are trustless

653
00:32:41,180 --> 00:32:41,680
systems.

654
00:32:41,900 --> 00:32:42,040
Right.

655
00:32:42,140 --> 00:32:45,020
That we're presuming self-interest is the guiding principle for

656
00:32:45,020 --> 00:32:45,980
why people came together.

657
00:32:46,100 --> 00:32:48,320
And that's not why societies generally come together.

658
00:32:48,320 --> 00:32:49,020
I agree.

659
00:32:49,200 --> 00:32:51,520
That's why I wanted to give both examples because I think they have

660
00:32:51,520 --> 00:32:55,220
different mechanisms as well for like what's like keeping them together.

661
00:32:57,720 --> 00:33:00,940
And I think that what you're bringing up is like a huge area of

662
00:33:00,940 --> 00:33:01,220
study.

663
00:33:01,420 --> 00:33:04,520
It's like how people right now are looking at alignment at the level

664
00:33:04,520 --> 00:33:05,600
of like one agent.

665
00:33:05,600 --> 00:33:09,520
But like what is the narrative alignment that like somehow can emerge

666
00:33:09,520 --> 00:33:12,120
to like align them.

667
00:33:12,200 --> 00:33:12,820
And like I don't know.

668
00:33:13,120 --> 00:33:15,860
I hope more people will go and like study that and like look at

669
00:33:15,860 --> 00:33:17,520
that story.

670
00:33:19,760 --> 00:33:20,840
I think maybe, yeah.

671
00:33:20,960 --> 00:33:23,300
We'll do one more and then move on just for the sake of time.

672
00:33:23,820 --> 00:33:24,100
Okay.

673
00:33:24,100 --> 00:33:27,900
I have a question about open source, which I think has like an interesting

674
00:33:27,900 --> 00:33:30,660
role here because it shows up in two ways.

675
00:33:30,800 --> 00:33:34,880
Both as like as an agent organization, like a sort of a social system.

676
00:33:35,820 --> 00:33:40,860
But also actually how one of the

677
00:33:40,860 --> 00:33:45,260
reasons AI has advanced so quickly is because of all those Python

678
00:33:45,260 --> 00:33:48,580
and Jupyter notebooks that people were, you know,

679
00:33:48,580 --> 00:33:52,760
there was a tradition of publishing papers along with working code

680
00:33:52,760 --> 00:33:56,240
that was a way for models to improve rapidly.

681
00:33:56,700 --> 00:34:00,200
so you mentioned that was like something in your talk.

682
00:34:00,340 --> 00:34:00,940
wasn't clear to me.

683
00:34:01,640 --> 00:34:02,760
It's definitely Eric.

684
00:34:02,920 --> 00:34:06,020
I think Eric and Yonan will be focusing on those topics in particular.

685
00:34:06,220 --> 00:34:08,460
And we've been like because there is this like real challenge to

686
00:34:08,460 --> 00:34:09,280
open source right now.

687
00:34:09,400 --> 00:34:12,520
And like how do you like handle all the contributions of AI and

688
00:34:12,520 --> 00:34:13,580
how do you like verify them?

689
00:34:13,920 --> 00:34:16,640
And we are thinking about like what comes after open source too.

690
00:34:16,640 --> 00:34:20,320
So maybe that's a perfect segue to Eric's talk now.
