A machine learning approach to recognize bias and discrimination in job advertisements
Bibliographic Data
| ID | 20396151 |
|---|---|
| Authors | Richard Frissen (0000-0002-8302-3923, Maastricht School of Management, corresponding author), Kolawole John Adebayo (0000-0001-7126-7026, Dublin City University), Rohan Nanda (0000-0001-7124-4799, Maastricht University) |
| Year | 2023 |
| Volume | 38 |
| Issue | 2 |
| Pages | 1025-1038 |
| Publication date | 2023-04-01 |
| Peer Reviewed | Yes |
| Open Access | Yes |
| Type | ARTICLE |
| Venue | AI & Society (JOURNAL) |
| Journal identifiers | ISSN: 0951-5666 • E-ISSN: 1435-5655 |
| Publisher | Springer Science and Business Media LLC (PUBLISHER) |
| DOI | 10.1007/s00146-022-01574-0 |
| OpenAlex | W4306677180 |
| Language | EN |
| Citations received | 7 |
| References cited | 19 |
In recent years, the work of organizations in the area of digitization has intensified significantly. This trend is also evident in the field of recruitment where job application tracking systems (ATS) have been developed to allow job advertisements to be published online. However, recent studies have shown that recruiting in most organizations is not inclusive, being subject to human biases and prejudices. Most discrimination activities appear early but subtly in the hiring process, for instance, exclusive phrasing in job advertisement discourages qualified applicants from minority groups from applying. The existing works are limited to analyzing, categorizing and highlighting the occurrence of bias in the recruitment process. In this paper, we go beyond this and develop machine learning models for identifying and classifying biased and discriminatory language in job descriptions. We develop and evaluate a machine learning system for identifying five major categories of biased and discriminatory language in job advertisements, i.e., masculine-coded, feminine-coded, exclusive, LGBTQ-coded, demographic and racial language. We utilized the combination of linguistic features with recent state-of-the-art word embeddings representations as input features for various machine learning classifiers. Our results show that the machine learning classifiers were able to identify all the five categories of biased and discriminatory language with a decent accuracy. The Random Forest classifier with FastText word embeddings achieved the best performance with tenfolds cross-validation. Our system directly addresses the bias in the attraction phase of hiring by identifying and classifying biased and discriminatory language and thus encouraging recruiters to write more inclusive job advertisements
Digitization · Gender bias · Machine learning · Natural language processing · Random forest · Authorship Attribution and Profiling · Computer Science · Gender Studies in Language · Names, Identity, and Discrimination Research · Psychology · Social Psychology · Artificial Intelligence
Prediction of the gender inequality index based on data-driven interpretable ensemble learning methods
ChatGPT versus humans in judging discriminatory scenarios
Ideal on Paper, Excluded in Practice
Automating public policy
Productive power in social networks
Is artificial intelligence (AI) research biased and conceptually vague? A systematic review of research on bias and discrimination in the context of using AI in human resource management
Deepfakes and the crisis of digital authenticity
Enriching Word Vectors with Subword Information
Deep Contextualized Word Representations
A meta-analysis of gender stereotypes and bias in experimental simulations of employment decision making.
Are Emily and Greg More Employable Than Lakisha and Jamal? A Field Experiment on Labor Market Discrimination
Developing the Research Basis for Controlling Bias in Hiring
Evidence that gendered wording in job advertisements exists and sustains gender inequality
| Unique citing works | 7 |
|---|---|
| Citations per year | 7 |
| Citation span | 2025 - 2026 (2) |
| Citation velocity | current |
| Highly cited | No |
| Citation types | Neutral: 7 |