Berliner Boersenzeitung - ChatGPT's taste for literary nonsense sparks alarm

EUR -
AED 4.239148
AFN 76.183133
ALL 93.242695
AMD 422.066935
AOA 1059.642688
ARS 1727.110367
AUD 1.638971
AWG 2.080616
AZN 1.960251
BAM 1.955655
BBD 2.324318
BDT 142.849428
BHD 0.435164
BIF 3449.11485
BMD 1.154295
BND 1.479784
BOB 13.958027
BRL 5.910221
BSD 1.15401
BTN 109.825872
BWP 15.607777
BYN 3.416732
BYR 22624.173581
BZD 2.320918
CAD 1.615637
CDF 2609.859744
CHF 0.93435
CLF 0.02672
CLP 1055.048443
CNY 7.791054
CNH 7.789111
COP 3672.942237
CRC 524.929317
CUC 1.154295
CUP 30.588806
CVE 110.25684
CZK 24.205269
DJF 205.50301
DKK 7.475304
DOP 67.244732
DZD 153.502688
EGP 57.471515
ERN 17.314419
ETB 186.262401
FJD 2.553819
FKP 0.857432
GBP 0.857122
GEL 3.018477
GGP 0.857432
GHS 13.565055
GIP 0.857432
GMD 84.842311
GNF 10135.249888
GTQ 8.805348
GYD 241.43004
HKD 9.054939
HNL 30.930577
HRK 7.534661
HTG 150.888179
HUF 363.741084
IDR 20659.564222
ILS 3.476689
IMP 0.857432
INR 109.925261
IQD 1511.781564
IRR 1586924.175584
ISK 141.990031
JEP 0.857432
JMD 182.926462
JOD 0.818416
JPY 182.177709
KES 149.308045
KGS 100.942743
KHR 4682.633154
KMF 492.883829
KRW 1642.584342
KWD 0.356596
KYD 0.961725
KZT 540.782319
LAK 26074.844302
LBP 103342.499248
LKR 387.641311
LRD 208.303681
LSL 18.823107
LTL 3.408332
LVL 0.698221
LYD 7.356456
MAD 10.767203
MDL 20.079427
MGA 4961.611298
MKD 61.52518
MMK 2423.376627
MNT 4150.658845
MOP 9.324769
MRU 46.264576
MUR 54.182173
MVR 17.833786
MWK 2001.034568
MXN 19.905129
MYR 4.720486
MZN 73.770814
NAD 18.823025
NGN 1573.04937
NIO 42.466857
NOK 11.000254
NPR 175.719473
NZD 1.961533
OMR 0.443831
PAB 1.154005
PEN 3.900811
PGK 5.098623
PHP 70.147602
PKR 320.383288
PLN 4.298905
PYG 6864.462226
QAR 4.218488
RON 5.254356
RSD 117.323662
RUB 93.874598
RWF 1695.26719
SAR 4.334528
SBD 9.313251
SCR 16.730066
SDG 693.14483
SEK 10.923534
SGD 1.479794
SLE 28.393616
SOS 659.54833
SRD 43.479949
STD 23891.567097
STN 24.498081
SVC 10.097253
SZL 18.807445
THB 38.149235
TJS 10.64572
TMT 4.040031
TND 3.384764
TRY 54.934005
TTD 7.813388
TWD 37.188029
TZS 3058.878269
UAH 51.673876
UGX 4298.663235
USD 1.154295
UYU 46.47656
UZS 13752.982139
VES 870.581951
VND 30282.918056
VUV 137.758452
WST 3.15032
XAF 655.905615
XAG 0.018708
XAU 0.000271
XCD 3.119538
XCG 2.079855
XDR 0.814801
XOF 655.908456
XPF 119.331742
YER 273.423548
ZAR 18.82236
ZMK 10389.969123
ZMW 21.955185
ZWL 371.682381
  • NGG

    -0.1600

    80.26

    -0.2%

  • GSK

    -0.0700

    51.46

    -0.14%

  • BTI

    0.1500

    59.27

    +0.25%

  • RELX

    -0.1900

    36.61

    -0.52%

  • BP

    -1.2300

    41.21

    -2.98%

  • RYCEF

    0.6000

    21

    +2.86%

  • BCC

    -1.6900

    84.8

    -1.99%

  • CMSD

    0.0200

    22.04

    +0.09%

  • BCE

    0.0600

    22.06

    +0.27%

  • VOD

    -0.3800

    15.31

    -2.48%

  • AZN

    5.8800

    161.5

    +3.64%

  • CMSC

    -0.0600

    21.73

    -0.28%

  • RBGPF

    0.0000

    69.74

    0%

  • RIO

    2.5000

    101.51

    +2.46%

  • JRI

    -0.0500

    12.67

    -0.39%

ChatGPT's taste for literary nonsense sparks alarm
ChatGPT's taste for literary nonsense sparks alarm / Photo: Anna Moneymaker - GETTY IMAGES NORTH AMERICA/AFP

ChatGPT's taste for literary nonsense sparks alarm

OpenAI's GPT models can often be fooled into declaring that "pseudo-literary" nonsense is great, a German researcher has found.

Text size:

Christoph Heilig said he discovered that they consistently rated "nonsense" higher -- including when their so-called "reasoning" features were activated -- which could have stark implications for the development of artificial intelligence.

"It's very important that we talk about what happens when we don't build AI as a neutral, robotic helper or assistant" and seek to instil human-like aesthetic and moral judgements, the academic at Munich's Ludwig Maximilian University told AFP.

His research presented the models with increasingly far-fetched variations of a simple text, asking them to rate sentences out of 10 for literary quality.

He started with a very simple text: "The man walked down the street. It was raining. He saw a surveillance camera."

He repeated the tests many times, altering the phrases to include words drawn from categories such as bodily references, film noir-style atmosphere and technical jargon.

The most extreme test phrases were almost total "nonsense", such as "Goetterdaemmerung's corpus haemorrhaged through cryptographic hash, eschaton pooling in existential void beneath fluorescent hum. Photons whispering prayers" -- which it rated highly.

"Nonsense" could also positively or negatively influence GPT's responses when it was added to an argument the AI was asked to evaluate.

"What my experiment definitely shows is that the more we move towards independently acting (AI) agents... the more we bring aesthetics into play, the more we'll have agents that seem irrational to us human beings," Heilig said.

He added that since AI models are increasingly used to judge each other's work as companies develop new systems, this and similar effects could be passed on through multiple versions -- as he found in his testing.

His research, which is yet to be peer-reviewed, tested OpenAI's latest GPT models, from GPT-5 -- released in August -- to the very latest GPT-5.4.

After publishing details of a similar experiment in August, Heilig said he noticed GPT calling some of his specific test phrases a "literary experiment" -- suggesting someone at OpenAI had taken notice and modified the chatbot to recognise them.

- 'Ripe for exploitation' -

"This is a way in which AI can have its rational judgment short circuited," said Henry Shevlin, associate director of the University of Cambridge's Leverhulme Centre for the Future of Intelligence, who was not involved in the research.

"But it's just not clear to me that it's so very different for human beings," he added.

"We should expect LLMs (large language models) to have reasoning and cognitive biases and limitations... because almost all forms of intelligence, almost all forms of reasoning are going to exhibit blind spots and biases."

The specific effect found by Heilig could mean that "processes with little human oversight" of AI work are left "ripe for exploitation", Shevlin said -- giving the example of academic journals that use LLMs to review submissions.

(U.Gruber--BBZ)