Mention versus Action
Guiding safety policy for content related to harmful topics
Bibliographic Data
| ID | 22159232 |
|---|---|
| Authors | Hadas Kotek (0009-0006-0360-2105), Leon Gatys, Margit Bowler, Yu'an Yang, Shruti Palaskar (0000-0001-8637-1897), Ciro Sannino, Gunnar Lund, Joseph Cheng (0000-0002-1881-012X), Robert Daland (0000-0002-2773-2637), Charlie Maalouf, Jeffrey Bigham |
| Year | 2026 |
| Volume | 11 |
| Issue | 1 |
| Pages | 6080 |
| Publication date | 2026-06-19 |
| Peer Reviewed | Yes |
| Open Access | Yes |
| Type | ARTICLE |
| Venue | Proceedings of the Linguistic Society of America (JOURNAL) |
| Journal identifiers | ISSN: 2473-8689 • E-ISSN: 2473-8689 |
| Publisher | Linguistic Society of America (PUBLISHER • US) |
| DOI | 10.3765/plsa.v11i1.6080 |
| OpenAlex | W7165373804 |
| Language | EN |
As generative AI systems are integrated into high-stakes domains, designing safety policies that accurately distinguish harmful from harm-free content has become a central challenge. A natural starting point is the use/mention distinction from linguistics and philosophy of language: content that merely mentions a harmful topic is typically less harmful than content that uses it to express harm. However, we argue that this binary is insufficient as a basis for responsible AI policy. We propose a mention vs. action framework that extends the use/mention distinction along two additional dimensions: the level of gratuitous detail and the discourse level contribution of the content. Together, these dimensions ground safety assessments in narrowly scoped, operationalizable criteria rather than opaque intent judgments. We demonstrate the framework through case studies in instruction-following, jailbreak attempts, and image content moderation, showing its applicability across modalities and its practical value for safety policy development
Generative grammar · Modalities · Adversarial Robustness in Machine Learning · Ethics and Social Impacts of AI · Hate Speech and Cyberbullying Detection
| Citation velocity | historical |
|---|---|
| Highly cited | No |