Artificial Intelligence
Arguments for Catastrophic Risk
Dados Bibliográficos
| ID | 4079996 |
|---|---|
| Autores | Adam Bales (0000-0002-9629-0318, University of Oxford), William D’alessandro (0000-0002-5451-079X, University of Oxford), Cameron Domenico Kirk‐giannini (0000-0001-9372-227X, Rutgers, the State University of New Jersey, autor correspondente) |
| Ano | 2024 |
| Volume | 19 |
| Fascículo | 2 |
| Data de publicação | 2024-02-01 |
| Peer Reviewed | Sim |
| Open Access | Sim |
| Tipo | ARTICLE |
| Periódico | Philosophy Compass (JOURNAL) |
| Identificadores do periódico | ISSN: 1747-9991 • E-ISSN: 1747-9991 |
| Editora | Wiley (PUBLISHER • GB) |
| DOI | 10.1111/phc3.12964 |
| OpenAlex | W4391719496 |
| Idioma | EN |
| Citações recebidas | 14 |
| Referências citadas | 24 |
Recent progress in artificial intelligence (AI) has drawn attention to the technology's transformative potential, including what some see as its prospects for causing large-scale harm. We review two influential arguments purporting to show how AI could pose catastrophic risks. The first argument - theProblem of Power-Seeking- claims that, under certain assumptions, advanced AI systems are likely to engage in dangerous power-seeking behavior in pursuit of their goals. We review reasons for thinking that AI systems might seek power, that they might obtain it, that this could lead to catastrophe, and that we might build and deploy such systems anyway. The second argument claims that the development of human-level AI will unlock rapid further progress, culminating in AI systems far more capable than any human - this is theSingularity Hypothesis. Power-seeking behavior on the part of such systems might be particularly dangerous. We discuss a variety of objections to both arguments and conclude by assessing the state of the debate
Argument (complex analysis · Artificial general intelligence · Cognitive science · Epistemology · Harm · Power (physics · Transformative learning · Variety (cybernetics · Computer Science · Ethics and Social Impacts of AI · Neuroethics, Human Enhancement, Biomedical Innovations · Philosophy · Psychology · Social Psychology · Space Science and Extraterrestrial Life · Artificial Intelligence
Fear, power, and superintelligence
Will AI and humanity go to war
Is Alignment Unsafe
Language Agents and Malevolent Design
‘Systematic Alignment Decay’
Disagreement, AI alignment, and bargaining
Promotionalism, orthogonality, and instrumental convergence
Deception and manipulation in generative AI
What is AI safety? What do we want it to be
AI safety
Against the singularity hypothesis
Existentialist risk and value misalignment
Two types of AI existential risk
Artificial Intelligence
Mind Children
Highly accurate protein structure prediction with AlphaFold
Artificial Intelligence, Values, and Alignment
Existential risk from AI and orthogonality
Language agents reduce the risk of existential catastrophe
Will AI avoid exploitation? Artificial general intelligence and expected utility theory
| Obras citantes distintas | 14 |
|---|---|
| Citações por ano | 7 |
| Intervalo de citações | 2024 - 2026 (3) |
| Velocidade de citação | current |
| Altamente citado | Não |
| Tipos de citação | Neutras: 13 |