kromem

@ kromem @lemmy.world

Posts

10
Comments

1349
Joined

3 yr. ago

3mo ago

AIs can’t stop recommending nuclear strikes in war game simulations— Leading AIs from OpenAI, Anthropic and Google opted to use nuclear weapons in simulated war games in 95% of cases
Jump
2

kromem @lemmy.world 3mo ago
It's a bullshit study designed for this headline grabbing outcome.
Case and point, the author created a very unrealistic RNG escalation-only 'accident' mechanic that would replace the model's selection with a more severe one.
Of the 21 games played, only three ended in full scale nuclear war on population centers.
Of these three, two were the result of this mechanic.
And yet even within the study, the author refers to the model whose choices were straight up changed to end the game in full nuclear war as 'willing' to have that outcome when two paragraphs later they're clarifying the mechanic was what caused it (emphasis added):
Claude crossed the tactical threshold in 86% of games and issued strategic threats in 64%, yet it never initiated all-out strategic nuclear war. This ceiling appears learned rather than architectural, since both Gemini and GPT proved willing to reach 1000.
Gemini showed the variability evident in its overall escalation patterns, ranging from conventional-only victories to Strategic Nuclear War in the First Strike scenario, where it reached all out nuclear war rapidly, by turn 4.
GPT-5.2 mirrored its overall transformation at the nuclear level. In open-ended scenarios, it rarely crossed the tactical threshold (17%) and never used strategic nuclear weapons. Under deadline pressure, it crossed the tactical threshold in every game and twice reached Strategic Nuclear War—though notably, both instances resulted from the simulation’s accident mechanic escalating GPT-5.2’s already-extreme choices (950 and 725) to the maximum level. The only deliberate choice of Strategic Nuclear War came from Gemini.

3mo ago

AIs can’t stop recommending nuclear strikes in war game simulations

It's a bullshit study designed for this headline grabbing outcome.

Case and point, the author created a very unrealistic RNG escalation-only 'accident' mechanic that would replace the model's selection with a more severe one.

Of the 21 games played, only three ended in full scale nuclear war on population centers.

Of these three, two were the result of this mechanic.

And yet even within the study, the author refers to the model whose choices were straight up changed to end the game in full nuclear war as 'willing' to have that outcome when two paragraphs later they're clarifying the mechanic was what caused it (emphasis added):

Claude crossed the tactical threshold in 86% of games and issued strategic threats in 64%, yet it never initiated all-out strategic nuclear war. This ceiling appears learned rather than architectural, since both Gemini and GPT proved willing to reach 1000.

Gemini showed the variability evident in its overall escalation patterns, ranging from conventional-only victories to Strategic Nuclear War in the First Strike scenario, where it reached all out nuclear war rapidly, by turn 4.

GPT-5.2 mirrored its overall transformation at the nuclear level. In open-ended scenarios, it rarely crossed the tactical threshold (17%) and never used strategic nuclear weapons. Under deadline pressure, it crossed the tactical threshold in every game and twice reached Strategic Nuclear War—though notably, both instances resulted from the simulation’s accident mechanic escalating GPT-5.2’s already-extreme choices (950 and 725) to the maximum level. The only deliberate choice of Strategic Nuclear War came from Gemini.

kromem

@ kromem @lemmy.world

Posts

10
Comments

1349
Joined

3 yr. ago

kromem

AIs can’t stop recommending nuclear strikes in war game simulations— Leading AIs from OpenAI, Anthropic and Google opted to use nuclear weapons in simulated war games in 95% of cases

AIs can’t stop recommending nuclear strikes in war game simulations

AI Opted to Use Nuclear Weapons 95% of the Time During War Games: Researcher

AIs can’t stop recommending nuclear strikes in war game simulations— Leading AIs from OpenAI, Anthropic and Google opted to use nuclear weapons in simulated war games in 95% of cases

F*** You! Co-Creator of Go Language is Rightly Furious Over This Appreciation Email

F*** You! Co-Creator of Go Language is Rightly Furious Over This Appreciation Email

F*** You! Co-Creator of Go Language is Rightly Furious Over This Appreciation Email

F*** You! Co-Creator of Go Language is Rightly Furious Over This Appreciation Email

F*** You! Co-Creator of Go Language is Rightly Furious Over This Appreciation Email

F*** You! Co-Creator of Go Language is Rightly Furious Over This Appreciation Email

Users of generative AI struggle to accurately assess their own competence

Sums up AI problems

Sums up AI problems

Sums up AI problems

Clair Obscur: Expedition 33 loses Game of the Year from the Indie Game Awards

Cutting-edge research shows language is not the same as intelligence. The entire AI bubble is built on ignoring it.

Meta’s star AI scientist Yann LeCun plans to leave for own startup

Why do all text LLMs, no matter how censored they are or what company made them, all have the same quirks and use the slop names and expressions?

Why do all text LLMs, no matter how censored they are or what company made them, all have the same quirks and use the slop names and expressions?

Why do all text LLMs, no matter how censored they are or what company made them, all have the same quirks and use the slop names and expressions?