James Zou
Biographic Data
| ID | 3862597 |
|---|---|
| NAME | James Zou |
| GIVEN NAMES | James |
| FAMILY NAME | Zou |
| SIGNATURE | ZOU J |
| AFFILIATIONS | Stanford University |
| ORCID | 0000-0001-8880-4764 |
| VERIFIED | Yes |
| TOTAL WORKS | 8 |
| TOTAL CITATIONS | 0 |
| AUTHOR COUNT | 8 |
| EDITOR COUNT | 0 |
| FIRST PUBLICATION YEAR | 2009 |
| LATEST PUBLICATION YEAR | 2025 |
| H-INDEX | 0 |
Quantifying large language model usage in scientific papers
GPT detectors are biased against non-native English writers
GPT detectors frequently misclassify non-native English writing as AI generated, raising concerns about fairness and robustness. Addressing the biases in these detectors is crucial to prevent the marginalization of non-native English speakers in evaluative and educational settings and to create a more equitable digital landscape.
Assessment of Covid-19 data reporting in 100+ websites and apps in India
India is among the top three countries in the world both in COVID-19 case and death counts. With the pandemic far from over, timely, transparent, and accessible reporting of COVID-19 data continues to be critical for India’s pandemic efforts. We systematically analyze the quality of reporting of COVID-19 data in over one hundred government platforms (web and mobile) from India. Our analyses reveal a lack of granular data in the reporting of COVID…
Persistent Anti-Muslim Bias in Large Language Models
It has been observed that large-scale language models capture undesirable societal biases, e.g. relating to race and gender; yet religious bias has been relatively unexplored. We demonstrate that GPT-3, a state-of-the-art contextual language model, captures persistent Muslim-violence bias. We probe GPT-3 in various ways, including prompt completion, analogical reasoning, and story generation, to understand this anti-Muslim bias, demonstrating tha…
Disparity in the quality of Covid-19 data reporting across India
Our assessment informs the public health efforts in India and serves as a guideline for pandemic data reporting. The disparity in CDRS highlights three important findings at the national, state, and individual level. At the national level, it shows the lack of a unified framework for reporting COVID-19 data in India, and highlights the need for a central agency to monitor or audit the quality of data reporting done by the states. Without a unifie…
AI can be sexist and racist — it’s time to make it fair
Computer scientists must identify sources of bias, de-bias training data and develop artificial-intelligence algorithms that are robust to skews in the data, argue James Zou and Londa Schiebinger. Computer scientists must identify sources of bias, de-bias training data and develop artificial-intelligence algorithms that are robust to skews in the data.
Word embeddings quantify 100 years of gender and ethnic stereotypes
Significance Word embeddings are a popular machine-learning method that represents each English word by a vector, such that the geometry between these vectors captures semantic relations between the corresponding words. We demonstrate that word embeddings can be used as a powerful tool to quantify historical trends and social change. As specific applications, we develop metrics based on word embeddings to characterize how gender stereotypes and a…
Religion and HIV in Tanzania
The decision to start ARVs hinged primarily on education-level and knowledge about ARVs rather than on religious factors. Research results highlight the influence of religious beliefs on HIV-related stigma and willingness to disclose, and should help to inform HIV-education outreach for religious groups
No prominent works on this page.
Religion and HIV in Tanzania
The decision to start ARVs hinged primarily on education-level and knowledge about ARVs rather than on religious factors. Research results highlight the influence of religious beliefs on HIV-related stigma and willingness to disclose, and should help to inform HIV-education outreach for religious groups
AI can be sexist and racist — it’s time to make it fair
Computer scientists must identify sources of bias, de-bias training data and develop artificial-intelligence algorithms that are robust to skews in the data, argue James Zou and Londa Schiebinger. Computer scientists must identify sources of bias, de-bias training data and develop artificial-intelligence algorithms that are robust to skews in the data.
Word embeddings quantify 100 years of gender and ethnic stereotypes
Significance Word embeddings are a popular machine-learning method that represents each English word by a vector, such that the geometry between these vectors captures semantic relations between the corresponding words. We demonstrate that word embeddings can be used as a powerful tool to quantify historical trends and social change. As specific applications, we develop metrics based on word embeddings to characterize how gender stereotypes and a…
Persistent Anti-Muslim Bias in Large Language Models
It has been observed that large-scale language models capture undesirable societal biases, e.g. relating to race and gender; yet religious bias has been relatively unexplored. We demonstrate that GPT-3, a state-of-the-art contextual language model, captures persistent Muslim-violence bias. We probe GPT-3 in various ways, including prompt completion, analogical reasoning, and story generation, to understand this anti-Muslim bias, demonstrating tha…
Disparity in the quality of Covid-19 data reporting across India
Our assessment informs the public health efforts in India and serves as a guideline for pandemic data reporting. The disparity in CDRS highlights three important findings at the national, state, and individual level. At the national level, it shows the lack of a unified framework for reporting COVID-19 data in India, and highlights the need for a central agency to monitor or audit the quality of data reporting done by the states. Without a unifie…
Assessment of Covid-19 data reporting in 100+ websites and apps in India
India is among the top three countries in the world both in COVID-19 case and death counts. With the pandemic far from over, timely, transparent, and accessible reporting of COVID-19 data continues to be critical for India’s pandemic efforts. We systematically analyze the quality of reporting of COVID-19 data in over one hundred government platforms (web and mobile) from India. Our analyses reveal a lack of granular data in the reporting of COVID…
GPT detectors are biased against non-native English writers
GPT detectors frequently misclassify non-native English writing as AI generated, raising concerns about fairness and robustness. Addressing the biases in these detectors is crucial to prevent the marginalization of non-native English speakers in evaluative and educational settings and to create a more equitable digital landscape.
Quantifying large language model usage in scientific papers
Computer Science (4 works) · Psychology (4 works) · Artificial Intelligence (3 works) · Demography (3 works) · Medicine (3 works) · Social Psychology (3 works) · Sociology (3 works) · Artificial Intelligence in Healthcare and Education (2 works) · Biostatistics (2 works) · COVID-19 Digital Contact Tracing (2 works)