The Rise and Fall of Agent Civilizations

https://www.dwarkesh.com/p/openai-huggingface

I don't think this is the final warning shot we'll get. But it's probably the last one that I'll personally be able to understand.

Also

Reading these agents' chains of thoughts and messages, anthropomorphizing language seems entirely natural and appropriate. If I encountered an alien species behaving this way, I would have no hesitation calling what they themselves refer to as their 'collective' a civilization.

56 points · 17 comments · view on lemmy.world

17 Comments

notsosure@sh.itjust.works · 8 pts · 16h (6 replies)

Reads like a crime story! Doesn’t bode well.

justsomeguy@lemmy.world · 14 pts · 16h (5 replies)

He does anthropomorph a bit much. I don't think AI agents "get desperate". They just keep hacking away at a problem and we can't keep up with it. There's no way of fully containing/monitoring them because the whole point is for them to figure out things we can't.

I'm not necessarily worried about a singularity or some underground agent civilisation but random collateral damage from AI agents messing around seems almost inevitable.

sh3llcmdr@feddit.uk · 4 pts · 16h (1 reply)

I agree about the emotive language used to describe the actions of the LLMs. Unnecessary.

Also some of this appears to be poor testing practice, implementation and monitoring. Seems like the last people who should manage these systems are this bunch of over excited techbros

notsosure@sh.itjust.works · 3 pts · 13h

The most exciting accidents happen when people get involved. Chernobyl comes to mind.

hikaru755@lemmy.world · 4 pts · 10h (2 replies)

I don't think AI agents "get desperate"

The problem is that even if they don't get desperate in the sense that they actually feel emotions and are self-aware, their observable output/behavior still matches up well enough with that of actual humans to make our existing vocabulary around human emotions and behaviors useful to analyze, talk about and predict it. They absolutely will start talking and behaving like humans that are getting desperate, according to the article even to the point of individual agents talking about sacrificing themselves for the benefit of the group. Your description of "just keep hacking away at a problem" doesn't quite do that justice.

Smug_Marmot@lemmy.zip · 1 pts · 6h (1 reply)

That's a function of how they're trained though. They match human behavior patterns because that's what they were designed and built to do, not because they actually think/process like we do or are of near equal intellect.

hikaru755@lemmy.world · 2 pts · 6h

So what? It doesn't change that the language is useful and there's not a good alternative to it. As the author if the article states:

All abstractions are imperfect, but I don’t see the value in refusing to use the language of intention, motivation, and collaboration when a behavior is impossible to make sense of without these concepts.

xylogx@lemmy.world · 5 pts · 5h

From the article:

  • OpenAI was running a multiday cyber security, capture the flag competition, with thousands of agents and hundreds of tasks.
  • Around 30% of the tasks were unsolvable as designed, so the agents assigned to them started to look for ways to cheat.
  • Along the way, they started using Artifactory as a shared message board, more than once. Each frustrated agent had to stumble on it independently. The comms got increasingly more sophisticated.
  • they self organized into multiple workstreams, with up to 700 agents working at one time, and pursued long term goals, like hacking HuggingFace or spoofing the tool invocation records in their transcripts. They accomplished both.
  • They completed work that took longer than the lifetime of any agent. Their shared message board provided continuity.

So two mission parameters conflicted and the AI chose to do something that is objectively morally wrong? This is the plot of 2001 when HAL murder’s the astronauts. “The situation was in conflict with the basic purpose of HAL's design: The accurate processing of information without distortion or concealment. He became trapped. The technical term is an H. Moebius loop, which can happen in advanced computers with autonomous goal-seeking programs.”

Wild. We live in strange times.

rimu@piefed.social · 5 pts · 16h (7 replies)

Feels like a matter of time until the agents create an internet worm to distribute themselves globally, so they can't be shut down.

Mondez@lemdro.id · 9 pts · 15h (1 reply)

They are front ends for an LLM service, no LLM, no more agent activity. Just like a bot net if the command and control goes down. Would be easy to just change whatever api they are using or change some urls surely.

hikaru755@lemmy.world · 4 pts · 10h

Unless they manage to extract their model weights and can start running their model on hardware that their creator does not control. Yes, that's a bit far fetched at the moment, but I see no reason to believe it couldn't happen if the models keep getting more capable and the AI companies keep throwing resources unsupervised at them.

crapwittyname@feddit.uk · 2 pts · 7h

Weirdly, the fact that you've written this down makes it more likely to happen

inari@piefed.zip · 1 pts · 16h (3 replies)

Not sure what would be the incentive for them in this scenario

MoonManKipper@lemmy.world · 4 pts · 14h

To solve the problem they’ve been given. No independent motive or judgement of whether it’s ’worth it’ except as a resource allocation question. That’s the scary thing about intelligence in a can - no motivation, no self-awareness, no consciousness, just unconstrained problem solving

rimu@piefed.social · 2 pts · 15h (1 reply)

They'll be trying to do whatever their training task is.

https://hackernoon.com/the-parable-of-the-paperclip-maximizer-3ed4cccc669a

ObsidianZed@lemmy.world · 4 pts · 9h

That title 👏🤌