Noah A Smith
Biographic Data
| ID | 4180323 |
|---|---|
| NAME | Noah A Smith |
| GIVEN NAMES | Noah A |
| FAMILY NAME | Smith |
| SIGNATURE | SMITH N A |
| AFFILIATIONS | Carnegie Mellon University |
| ORCID | 0000-0002-2310-6380 |
| VERIFIED | Yes |
| TOTAL WORKS | 17 |
| TOTAL CITATIONS | 159 |
| AUTHOR COUNT | 17 |
| EDITOR COUNT | 0 |
| FIRST PUBLICATION YEAR | 2003 |
| LATEST PUBLICATION YEAR | 2024 |
| H-INDEX | 5 |
Troubles in Text: Using Natural Language Processing to Recognize Government Rationalizations for Rights Abuses
Patterns of Bias: How Mainstream Media Operationalize Links between Mass Shootings and Terrorism
How do race and/or religion shape news media coverage of mass shooters and whether media associate mass shooters with terrorism? This article combines natural language processing (NLP), statistical analysis of U.S. mass shooting events (1990–2016) and an in-depth case-study comparison to evaluate whether media exhibit patterns in how they frame mass shooters from different racial and/or religious groups. First, we use NLP to target and model the …
Multilingual and Interlingual Semantic Representations for Natural Language Processing: A Brief Introduction
We introduce the Computational Linguistics special issue on Multilingual and Interlingual Semantic Representations for Natural Language Processing. We situate the special issue’s five articles in the context of our fast-changing field, explaining our motivation for this project. We offer a brief summary of the work in the issue, which includes developments on lexical and sentential semantic representations, from symbolic and neural perspectives
Etch-a-Sketching: Evaluating the Post-Primary Rhetorical Moderation Hypothesis
Candidates have incentives to present themselves as strong partisans in primary elections, and then move “toward the center” upon advancing to the general election. Yet, candidates also face incentives not to flip-flop on their policy positions. These competing incentives suggest that candidates might use rhetoric to seem more partisan in the primary and more moderate in the general, even if their policy positions remain fixed. We test this idea …
The Risk of Racial Bias in Hate Speech Detection
We investigate how annotators' insensitivity to differences in dialect can lead to racial bias in automatic hate speech detection models, potentially amplifying harm against minority populations. We first uncover unexpected correlations between surface markers of African American English (AAE) and ratings of toxicity in several widely-used hate speech datasets. Then, we show that models trained on these corpora acquire and propagate these biases,…
World Vaping Day: Contextualizing Vaping Culture in Online Social Media Using a Mixed Methods Approach
Few studies have demonstrated the use of mixed methods research to contextualize health topics using primary data from social media. To address this gap in the methodological literature, we present research about electronic nicotine delivery systems, using Twitter data from “World Vaping Day.” To engage with the quantitative breadth and qualitative depth of 5,149 collected tweets, we utilized a convergent parallel mixed methods framework, integra…
Greedy Transition-Based Dependency Parsing with Stack LSTMs
We introduce a greedy transition-based parser that learns to represent parser states using recurrent neural networks. Our primary innovation that enables us to do this efficiently is a new control structure for sequential neural networks—the stack long short-term memory unit (LSTM). Like the conventional stack data structures used in transition-based parsers, elements can be pushed to or popped from the top of the stack in constant time, but, in …
How party nationalization conditions economic voting
Linguistic Markers of Status in Food Culture: Bourdieu’s Distinction in a Menu Corpus
Beginnings are always hard to trace. They tend to belong more to the realm of myth, as Tristram Shandy well knew. At what point did it become necessary, in the sense of unavoidable, to use computation to study culture? Was it a certain polemic, new kinds of data (Google Books, Project Gutenberg), the rise of analytical techniques (natural language processing, machine learning), technologies such as the internet or social media, or simply that pow…
Narrative framing of consumer sentiment in online restaurant reviews
The vast increase in online expressions of consumer sentiment offers a powerful new tool for studying consumer attitudes. To explore the narratives that consumers use to frame positive and negative sentiment online, we computationally investigate linguistic structure in 900,000 online restaurant reviews. Negative reviews, especially in expensive restaurants, were more likely to use features previously associated with narratives of trauma: negativ…
Frame-Semantic Parsing
Frame semantics is a linguistic theory that has been instantiated for English in the FrameNet lexicon. We solve the problem of frame-semantic parsing using a two-stage statistical model that takes lexical targets (i.e., content words and phrases) in their sentential contexts and predicts frame-semantic structures. Given a target in context, the first stage disambiguates it to a semantic frame. This model uses latent variables and semi-supervised …
Phrase Dependency Machine Translation with Quasi-Synchronous Tree-to-Tree Features
Recent research has shown clear improvement in translation quality by exploiting linguistic syntax for either the source or target language. However, when using syntax for both languages (“tree-to-tree” translation), there is evidence that syntactic divergence can hamper the extraction of useful rules (Ding and Palmer 2005 ). Smith and Eisner ( 2006 ) introduced quasi-synchronous grammar, a formalism that treats non-isomorphic structure softly us…
Censorship and deletion practices in Chinese social media
With Twitter and Facebook blocked in China, the stream of information from Chinese domestic social media provides a case study of social media behavior under the influence of active censorship. While much work has looked at efforts to prevent access to information in China (including IP blocking of foreign Web sites or search engine filtering), we present here the first large–scale analysis of political content censorship in social media, i.e., t…
Empirical Risk Minimization for Probabilistic Grammars: Sample Complexity and Hardness of Learning
Probabilistic grammars are generative statistical models that are useful for compositional and sequential structures. They are used ubiquitously in computational linguistics. We present a framework, reminiscent of structural risk minimization, for empirical risk minimization of probabilistic grammars using the log-loss. We derive sample complexity bounds in this framework that apply both to the supervised setting and the unsupervised setting. By …
From Tweets to Polls: Linking Text Sentiment to Public Opinion Time Series
We connect measures of public opinion measured from polls with sentiment measured from text. We analyze several surveys on consumer confidence and political opinion over the 2008 to 2009 period, and find they correlate to sentiment word frequencies in contempora- neous Twitter messages. While our results vary across datasets, in several cases the correlations are as high as 80%, and capture important large-scale trends. The re- sults highlight th…
Weighted and Probabilistic Context-Free Grammars Are Equally Expressive
This article studies the relationship between weighted context-free grammars (WCFGs), where each production is associated with a positive real-valued weight, and probabilistic context-free grammars (PCFGs), where the weights of the productions associated with a nonterminal are constrained to sum to one. Because the class of WCFGs properly includes the PCFGs, one might expect that WCFGs can describe distributions that PCFGs cannot. However, Z. Chi…
The Web as a Parallel Corpus
Parallel corpora have become an essential resource for work in multilingual natural language processing. In this article, we report on our work using the STRAND system for mining parallel text on the World Wide Web, first reviewing the original algorithm and results and then presenting a set of significant enhancements. These enhancements include the use of supervised learning based on structural features of documents to improve classification pe…
From Tweets to Polls: Linking Text Sentiment to Public Opinion Time Series
We connect measures of public opinion measured from polls with sentiment measured from text. We analyze several surveys on consumer confidence and political opinion over the 2008 to 2009 period, and find they correlate to sentiment word frequencies in contempora- neous Twitter messages. While our results vary across datasets, in several cases the correlations are as high as 80%, and capture important large-scale trends. The re- sults highlight th…
Etch-a-Sketching: Evaluating the Post-Primary Rhetorical Moderation Hypothesis
Candidates have incentives to present themselves as strong partisans in primary elections, and then move “toward the center” upon advancing to the general election. Yet, candidates also face incentives not to flip-flop on their policy positions. These competing incentives suggest that candidates might use rhetoric to seem more partisan in the primary and more moderate in the general, even if their policy positions remain fixed. We test this idea …
Patterns of Bias: How Mainstream Media Operationalize Links between Mass Shootings and Terrorism
How do race and/or religion shape news media coverage of mass shooters and whether media associate mass shooters with terrorism? This article combines natural language processing (NLP), statistical analysis of U.S. mass shooting events (1990–2016) and an in-depth case-study comparison to evaluate whether media exhibit patterns in how they frame mass shooters from different racial and/or religious groups. First, we use NLP to target and model the …
How party nationalization conditions economic voting
The Web as a Parallel Corpus
Parallel corpora have become an essential resource for work in multilingual natural language processing. In this article, we report on our work using the STRAND system for mining parallel text on the World Wide Web, first reviewing the original algorithm and results and then presenting a set of significant enhancements. These enhancements include the use of supervised learning based on structural features of documents to improve classification pe…
Frame-Semantic Parsing
Frame semantics is a linguistic theory that has been instantiated for English in the FrameNet lexicon. We solve the problem of frame-semantic parsing using a two-stage statistical model that takes lexical targets (i.e., content words and phrases) in their sentential contexts and predicts frame-semantic structures. Given a target in context, the first stage disambiguates it to a semantic frame. This model uses latent variables and semi-supervised …
World Vaping Day: Contextualizing Vaping Culture in Online Social Media Using a Mixed Methods Approach
Few studies have demonstrated the use of mixed methods research to contextualize health topics using primary data from social media. To address this gap in the methodological literature, we present research about electronic nicotine delivery systems, using Twitter data from “World Vaping Day.” To engage with the quantitative breadth and qualitative depth of 5,149 collected tweets, we utilized a convergent parallel mixed methods framework, integra…
Weighted and Probabilistic Context-Free Grammars Are Equally Expressive
This article studies the relationship between weighted context-free grammars (WCFGs), where each production is associated with a positive real-valued weight, and probabilistic context-free grammars (PCFGs), where the weights of the productions associated with a nonterminal are constrained to sum to one. Because the class of WCFGs properly includes the PCFGs, one might expect that WCFGs can describe distributions that PCFGs cannot. However, Z. Chi…
The Web as a Parallel Corpus
Parallel corpora have become an essential resource for work in multilingual natural language processing. In this article, we report on our work using the STRAND system for mining parallel text on the World Wide Web, first reviewing the original algorithm and results and then presenting a set of significant enhancements. These enhancements include the use of supervised learning based on structural features of documents to improve classification pe…
Weighted and Probabilistic Context-Free Grammars Are Equally Expressive
This article studies the relationship between weighted context-free grammars (WCFGs), where each production is associated with a positive real-valued weight, and probabilistic context-free grammars (PCFGs), where the weights of the productions associated with a nonterminal are constrained to sum to one. Because the class of WCFGs properly includes the PCFGs, one might expect that WCFGs can describe distributions that PCFGs cannot. However, Z. Chi…
From Tweets to Polls: Linking Text Sentiment to Public Opinion Time Series
We connect measures of public opinion measured from polls with sentiment measured from text. We analyze several surveys on consumer confidence and political opinion over the 2008 to 2009 period, and find they correlate to sentiment word frequencies in contempora- neous Twitter messages. While our results vary across datasets, in several cases the correlations are as high as 80%, and capture important large-scale trends. The re- sults highlight th…
Empirical Risk Minimization for Probabilistic Grammars: Sample Complexity and Hardness of Learning
Probabilistic grammars are generative statistical models that are useful for compositional and sequential structures. They are used ubiquitously in computational linguistics. We present a framework, reminiscent of structural risk minimization, for empirical risk minimization of probabilistic grammars using the log-loss. We derive sample complexity bounds in this framework that apply both to the supervised setting and the unsupervised setting. By …
Censorship and deletion practices in Chinese social media
With Twitter and Facebook blocked in China, the stream of information from Chinese domestic social media provides a case study of social media behavior under the influence of active censorship. While much work has looked at efforts to prevent access to information in China (including IP blocking of foreign Web sites or search engine filtering), we present here the first large–scale analysis of political content censorship in social media, i.e., t…
Frame-Semantic Parsing
Frame semantics is a linguistic theory that has been instantiated for English in the FrameNet lexicon. We solve the problem of frame-semantic parsing using a two-stage statistical model that takes lexical targets (i.e., content words and phrases) in their sentential contexts and predicts frame-semantic structures. Given a target in context, the first stage disambiguates it to a semantic frame. This model uses latent variables and semi-supervised …
Phrase Dependency Machine Translation with Quasi-Synchronous Tree-to-Tree Features
Recent research has shown clear improvement in translation quality by exploiting linguistic syntax for either the source or target language. However, when using syntax for both languages (“tree-to-tree” translation), there is evidence that syntactic divergence can hamper the extraction of useful rules (Ding and Palmer 2005 ). Smith and Eisner ( 2006 ) introduced quasi-synchronous grammar, a formalism that treats non-isomorphic structure softly us…
Narrative framing of consumer sentiment in online restaurant reviews
The vast increase in online expressions of consumer sentiment offers a powerful new tool for studying consumer attitudes. To explore the narratives that consumers use to frame positive and negative sentiment online, we computationally investigate linguistic structure in 900,000 online restaurant reviews. Negative reviews, especially in expensive restaurants, were more likely to use features previously associated with narratives of trauma: negativ…
Linguistic Markers of Status in Food Culture: Bourdieu’s Distinction in a Menu Corpus
Beginnings are always hard to trace. They tend to belong more to the realm of myth, as Tristram Shandy well knew. At what point did it become necessary, in the sense of unavoidable, to use computation to study culture? Was it a certain polemic, new kinds of data (Google Books, Project Gutenberg), the rise of analytical techniques (natural language processing, machine learning), technologies such as the internet or social media, or simply that pow…
World Vaping Day: Contextualizing Vaping Culture in Online Social Media Using a Mixed Methods Approach
Few studies have demonstrated the use of mixed methods research to contextualize health topics using primary data from social media. To address this gap in the methodological literature, we present research about electronic nicotine delivery systems, using Twitter data from “World Vaping Day.” To engage with the quantitative breadth and qualitative depth of 5,149 collected tweets, we utilized a convergent parallel mixed methods framework, integra…
Greedy Transition-Based Dependency Parsing with Stack LSTMs
We introduce a greedy transition-based parser that learns to represent parser states using recurrent neural networks. Our primary innovation that enables us to do this efficiently is a new control structure for sequential neural networks—the stack long short-term memory unit (LSTM). Like the conventional stack data structures used in transition-based parsers, elements can be pushed to or popped from the top of the stack in constant time, but, in …
How party nationalization conditions economic voting
The Risk of Racial Bias in Hate Speech Detection
We investigate how annotators' insensitivity to differences in dialect can lead to racial bias in automatic hate speech detection models, potentially amplifying harm against minority populations. We first uncover unexpected correlations between surface markers of African American English (AAE) and ratings of toxicity in several widely-used hate speech datasets. Then, we show that models trained on these corpora acquire and propagate these biases,…
Multilingual and Interlingual Semantic Representations for Natural Language Processing: A Brief Introduction
We introduce the Computational Linguistics special issue on Multilingual and Interlingual Semantic Representations for Natural Language Processing. We situate the special issue’s five articles in the context of our fast-changing field, explaining our motivation for this project. We offer a brief summary of the work in the issue, which includes developments on lexical and sentential semantic representations, from symbolic and neural perspectives
Etch-a-Sketching: Evaluating the Post-Primary Rhetorical Moderation Hypothesis
Candidates have incentives to present themselves as strong partisans in primary elections, and then move “toward the center” upon advancing to the general election. Yet, candidates also face incentives not to flip-flop on their policy positions. These competing incentives suggest that candidates might use rhetoric to seem more partisan in the primary and more moderate in the general, even if their policy positions remain fixed. We test this idea …
Patterns of Bias: How Mainstream Media Operationalize Links between Mass Shootings and Terrorism
How do race and/or religion shape news media coverage of mass shooters and whether media associate mass shooters with terrorism? This article combines natural language processing (NLP), statistical analysis of U.S. mass shooting events (1990–2016) and an in-depth case-study comparison to evaluate whether media exhibit patterns in how they frame mass shooters from different racial and/or religious groups. First, we use NLP to target and model the …
Troubles in Text: Using Natural Language Processing to Recognize Government Rationalizations for Rights Abuses
Computer Science (15 works) · Artificial Intelligence (9 works) · Natural language processing (9 works) · Natural Language Processing Techniques (7 works) · Political science (7 works) · Topic Modeling (7 works) · Law (6 works) · Linguistics (6 works) · History (5 works) · Mathematics (5 works)