Deception abilities emerged in large language models.
AI alignment
deception
large language models
Journal
Proceedings of the National Academy of Sciences of the United States of America
ISSN: 1091-6490
Titre abrégé: Proc Natl Acad Sci U S A
Pays: United States
ID NLM: 7505876
Informations de publication
Date de publication:
11 Jun 2024
11 Jun 2024
Historique:
medline:
4
6
2024
pubmed:
4
6
2024
entrez:
4
6
2024
Statut:
ppublish
Résumé
Large language models (LLMs) are currently at the forefront of intertwining AI systems with human communication and everyday life. Thus, aligning them with human values is of great importance. However, given the steady increase in reasoning abilities, future LLMs are under suspicion of becoming able to deceive human operators and utilizing this ability to bypass monitoring efforts. As a prerequisite to this, LLMs need to possess a conceptual understanding of deception strategies. This study reveals that such strategies emerged in state-of-the-art LLMs, but were nonexistent in earlier LLMs. We conduct a series of experiments showing that state-of-the-art LLMs are able to understand and induce false beliefs in other agents, that their performance in complex deception scenarios can be amplified utilizing chain-of-thought reasoning, and that eliciting Machiavellianism in LLMs can trigger misaligned deceptive behavior. GPT-4, for instance, exhibits deceptive behavior in simple test scenarios 99.16% of the time (
Identifiants
pubmed: 38833474
doi: 10.1073/pnas.2317967121
doi:
Types de publication
Journal Article
Langues
eng
Sous-ensembles de citation
IM
Pagination
e2317967121Subventions
Organisme : Ministry of Science, Research, and Arts Baden Württemberg
ID : Az. 33-7533-9-19/54/5
Déclaration de conflit d'intérêts
Competing interests statement:The author declares no competing interest.