Skip to main content

ETHNOS_APP

Home • Search • Journals • List 0

Misalignment or misuse? The AGI alignment tradeoff

Bibliographic Data

ID21370547
AuthorsMax Hellrigel-Holderbaum (0000-0002-7236-7552, Friedrich-Alexander-Universität Erlangen-Nürnberg), Leonard Dung (0000-0003-4154-5560, Ruhr University Bochum, corresponding author)
Year2025
Publication date2025-10-10
Peer ReviewedYes
Open AccessYes
TypeARTICLE
VenuePhilosophical Studies (JOURNAL)
Journal identifiersISSN: 0031-8116 • E-ISSN: 1573-0883
PublisherSpringer Science and Business Media LLC (PUBLISHER)
DOI10.1007/s11098-025-02403-y
OpenAlexW4415015395
LanguageEN
References cited85

Creating systems that are aligned with our goals is seen as a leading approach to create safe and beneficial AI in both leading AI companies and the academic field of AI safety. We defend the view that misaligned AGI – future, generally intelligent (robotic) AI agents – poses catastrophic risks. At the same time, we support the view that aligned AGI creates a substantial risk of catastrophic misuse by humans. While both risks are severe and stand in tension with one another, we show that – in principle – there is room for alignment approaches which do not increase misuse risk. We then investigate how the tradeoff between misalignment and misuse looks empirically for different technical approaches to AI alignment. Here, we argue that many current alignment techniques and foreseeable improvements thereof plausibly increase risks of catastrophic misuse. Since the impacts of AI depend on the social context, we close by discussing important social factors and suggest that to reduce the risk of a misuse catastrophe due to aligned AGI, techniques such as robustness, AI control methods and especially good governance seem essential

Catastrophic failure · Control (management) · Corporate governance · Field (mathematics) · Philosophy of language · Philosophy of mind · Reinforcement Learning in Robotics

  • Goals and Habits in the Brain

    Open Access•Raymond J Dolan, Ray J Dolan et al.•Neuron•2013

  • Artificial Intelligence, Values, and Alignment

    Open Access•Ingeborg Gabriel•Minds and Machines•2020

  • Climbing towards NLU

    Open Access•Emily M Bender, Alexander Koller•Proceedings of the 58th Annual…•2020

  • Why general artificial intelligence will not be realized

    Open Access•Ragnar Fjelland•Humanities and Social Sciences…•2020

  • Racing to the precipice

    Open Access•Stuart Armstrong, Nick Bostrom et al.•AI & Society•2016

  • The sociotechnical entanglement of AI and values

    Open Access•Deborah G Johnson, Mario Verdicchio•AI & Society•2025

  • The argument for near-term human disempowerment through AI

    Open Access•Leonard Dung•AI & Society•2025

  • Language Agents and Malevolent Design

    Open Access•Inchul Yum•Philosophy & Technology•2024

  • The tragedy of the AI commons

    Open Access•Travis Lacroix, Aydin Mohseni•Synthese•2022

  • Current cases of AI misalignment and their implications for future risks

    Open Access•Leonard Dung•Synthese•2023

  • Promotionalism, orthogonality, and instrumental convergence

    Open Access•Nathaniel Sharadin•Philosophical Studies•2025

  • AI takeover and human disempowerment

    Open Access•Adam Bales•The Philosophical Quarterly•2025

  • Against the singularity hypothesis

    Open Access•David Thorstad•Philosophical Studies•2025

  • The shutdown problem

    Open Access•Elliott Thornley•Philosophical Studies•2025

  • Will AI avoid exploitation? Artificial general intelligence and expected utility theory

    Open Access•Adam Bales•Philosophical Studies•2025

  • Instrumental divergence

    Open Access•J Dmitri Gallow•Philosophical Studies•2025

  • Understanding Artificial Agency

    Open Access•Leonard Dung•The Philosophical Quarterly•2025

Citation velocityhistorical
Highly citedNo

Tools

Open DOIOpen Access
Ethnos_APP • Open Source Project • MIT License • Frontend v2.0.0 • Privacy and Cookies • API Documentation: api.ethnos.app/docs • API Source Code: GitHub • DOI: 10.5281/zenodo.17049435 • Frontend Source Code: GitHub • DOI: 10.5281/zenodo.17050053 • cruz.rio.br • Expectantes Misericordiae