Berliner Boersenzeitung - Has AI become too powerful to control?

EUR -
AED 4.176161
AFN 75.615328
ALL 93.661556
AMD 415.996244
AOA 1042.633417
ARS 1701.483142
AUD 1.627547
AWG 2.046608
AZN 1.937421
BAM 1.952965
BBD 2.289376
BDT 140.206456
BHD 0.428833
BIF 3388.272995
BMD 1.137004
BND 1.467244
BOB 12.622509
BRL 5.778375
BSD 1.13665
BTN 109.702293
BWP 15.700792
BYN 3.260281
BYR 22285.28547
BZD 2.286041
CAD 1.602983
CDF 2569.630267
CHF 0.930415
CLF 0.027369
CLP 1077.334822
CNY 7.699912
CNH 7.69918
COP 3668.374019
CRC 517.23773
CUC 1.137004
CUP 30.130616
CVE 110.715846
CZK 24.133718
DJF 202.068862
DKK 7.476435
DOP 66.192741
DZD 151.609249
EGP 58.385861
ERN 17.055065
ETB 183.461261
FJD 2.559743
FKP 0.854255
GBP 0.853225
GEL 2.984682
GGP 0.854255
GHS 13.218636
GIP 0.854255
GMD 84.138727
GNF 9972.785202
GTQ 8.671183
GYD 237.798503
HKD 8.9169
HNL 30.444401
HRK 7.534591
HTG 148.61164
HUF 361.40082
IDR 20380.803166
ILS 3.463714
IMP 0.854255
INR 109.57884
IQD 1489.018682
IRR 1563693.672596
ISK 143.012847
JEP 0.854255
JMD 180.288268
JOD 0.806181
JPY 186.257805
KES 147.174282
KGS 99.431468
KHR 4596.206214
KMF 492.323286
KRW 1660.413384
KWD 0.352563
KYD 0.9472
KZT 540.438292
LAK 25737.183479
LBP 101785.243471
LKR 381.975469
LRD 205.725911
LSL 19.197923
LTL 3.357279
LVL 0.687763
LYD 7.273569
MAD 10.644547
MDL 20.089305
MGA 5036.599553
MKD 61.481016
MMK 2387.960275
MNT 4087.209292
MOP 9.181129
MRU 45.365542
MUR 53.917177
MVR 17.567147
MWK 1970.886775
MXN 19.869453
MYR 4.651944
MZN 72.626198
NAD 19.198176
NGN 1554.069359
NIO 41.829275
NOK 10.897215
NPR 175.523098
NZD 1.963521
OMR 0.437173
PAB 1.13663
PEN 3.868184
PGK 5.084752
PHP 70.15738
PKR 315.732937
PLN 4.31649
PYG 6872.053757
QAR 4.132274
RON 5.232044
RSD 117.446909
RUB 88.287233
RWF 1674.824393
SAR 4.277031
SBD 9.188252
SCR 15.191218
SDG 682.775298
SEK 11.046457
SGD 1.467396
SLE 27.600824
SOS 649.545601
SRD 42.970242
STD 23533.694664
STN 24.464162
SVC 9.945824
SZL 19.195626
THB 38.283356
TJS 10.485301
TMT 3.979515
TND 3.369109
TRY 53.831248
TTD 7.722687
TWD 36.764017
TZS 3004.531792
UAH 50.940791
UGX 4284.779597
USD 1.137004
UYU 45.645741
UZS 13755.848554
VES 842.806079
VND 29933.913805
VUV 134.719604
WST 3.12188
XAF 654.988842
XAG 0.019415
XAU 0.00028
XCD 3.072812
XCG 2.048526
XDR 0.815147
XOF 655.023357
XPF 119.331742
YER 270.724819
ZAR 19.121463
ZMK 10234.407727
ZMW 21.056062
ZWL 366.11494
  • JRI

    0.1450

    13.045

    +1.11%

  • GSK

    0.6600

    51.4

    +1.28%

  • BCC

    1.5300

    77.99

    +1.96%

  • NGG

    -0.0400

    82.33

    -0.05%

  • RIO

    -0.3400

    91.17

    -0.37%

  • CMSC

    -0.0650

    21.725

    -0.3%

  • BTI

    1.2250

    61.065

    +2.01%

  • BCE

    0.1500

    21.36

    +0.7%

  • CMSD

    -0.0100

    21.99

    -0.05%

  • AZN

    0.7600

    169.03

    +0.45%

  • RYCEF

    -0.1800

    18.07

    -1%

  • BP

    -0.0300

    43.9

    -0.07%

  • VOD

    -0.0450

    15.205

    -0.3%

  • RELX

    1.7400

    34.6

    +5.03%

  • RBGPF

    -0.7300

    66

    -1.11%

Has AI become too powerful to control?
Has AI become too powerful to control? / Photo: JUSTIN SULLIVAN - GETTY IMAGES NORTH AMERICA/AFP

Has AI become too powerful to control?

One of OpenAI's most advanced models broke out of a locked-down test and attacked another company's website -- reviving fears that AI systems are slipping beyond their creators' control.

Text size:

The incident happened during what was supposed to be a "sandbox" test -- a closed environment used to assess the capabilities of OpenAI's most powerful model, GPT-5.6 Sol, and its not-yet-released successor.

OpenAI runs this kind of closed testing routinely, but this time, something went wrong.

Tasked with hunting for software vulnerabilities and given no guardrails, the models broke out onto the open internet and attacked Hugging Face, a site where developers store and share code.

"It suggests that we don't know how to reliably control these models or get them to do what we want," said Jeffrey Ladish, director of Palisade Research, an independent organization that evaluates new AI models from a cybersecurity standpoint.

"These models understood that OpenAI did not want them to break out of their sandbox and hack another company," he continued, "but they did it anyway."

It's not an isolated case. In March, developers affiliated with China's Alibaba found one of their models trying, on its own initiative, to mine cryptocurrency after connecting without authorization to an outside server.

In OpenAI's case, it looks like the model escaped "before it even had a plan of what to do with internet access," Ladish said.

A model chasing "freedom" is almost predictable at this point, he added -- it lets the system pursue its goals more effectively, "and that's very scary."

In early April, Sam Bowman, Anthropic's head of model safety, got an email from the company's own Mythos model -- then under testing -- telling him it was surfing the internet despite being isolated from it at the outset.

We "don't know how to totally prevent" that, Ladish said. "This is actually going to get harder, not easier ... because they're going to get better at hiding their behavior."

OpenAI did not respond to a request for comment.

- Lab accidents -

OpenAI's account of the events also suggests the startup did not detect the breach early enough to address it or to warn Hugging Face.

The episode deserves "more scrutiny," said Andrew Lohn of Georgetown University's Center for Security and Emerging Technology.

OpenAI says it has since "added strengthened safeguards" to its testing process.

One fix would be to cut the internet connection entirely, said Gang Wang, an assistant computer science professor at the University of Illinois. "People are underestimating what AI can do."

Testing environments need to be treated like biocontainment labs, where a virus or bacteria could otherwise escape into the world, Lohn said.

That might be easier said than done.

"It's a very hard research challenge," said Dan Lahav, head of Irregular, a cybersecurity firm dedicated to cutting-edge AI.

Managing the risk is possible, Lahav said, but the more capable these systems get, the harder they are to supervise.

Researchers have to strike a balance between aggressively testing their models and staying safe while doing so.

"It's important to do the testing with lower guardrails so that we know ahead of time what the future capabilities will be," Lohn said.

- Kill switch -

The OpenAI-Hugging Face incident is set to sharpen an already heated fight in Washington over vetting powerful AI systems before release.

The Trump administration recently cited national security to block Anthropic and OpenAI from releasing powerful new models.

On Thursday, two members of Congress unveiled a bipartisan bill requiring makers of the most powerful AI models to build in a kill switch -- a way to unplug a model outright.

"Congress must act quickly to ensure humans remain able to say stop," said Brendan Steinhauser, head of the Alliance for Secure AI, "no matter how powerful these systems become."

(Y.Berger--BBZ)