Racial bias and fairness in logistic regression models for AI-Based income prediction

Abstract

The use of artificial intelligence (AI) systems to aid socioeconomic decision-making is expanding; there are substantial ethical concerns because AI systems can replicate existing discrepancies among biased data and modelling procedures. This study uses logistic regression to assess model bias and fairness. An AI-based income categorisation model was built using the University of California, Irvine (UCI) Adult Census dataset, which has 32,561 entries and 15 demographic and socioeconomic parameters. The model’s objective was to predict whether an individual’s annual income would exceed $50,000. It established good prediction performance on the test data, with an AUC of 69.33, a recall of 93.64%, a precision of 83.70%, and an overall accuracy of 81.53%.
Despite these findings, fairness evaluations across ethnic groups displayed significant differences. Positive prediction rates were from 75.97% to 92.28%, while accuracy varied from 94.00% for those classified as “Other” to 80.52% for the Asian-Pacific-Islander category. In a similar pattern, recall scores ranged from 88.50% to 97.19%, indicating discrepancies in model outputs. These differences reveal an ingrained racial bias in the prediction procedure, as the model failed to meet fairness standards, including equal accuracy, group fairness, and equal opportunity.
The study findings establish how, when trained on datasets with imbalanced demographics, even interpretable methods, such as logistic regression, can aggravate disparities. To support more equitable AI-driven socioeconomic assessments, the study emphasises the significance of integrating fairness indicators in AI model evaluation and advocates for strategies containing data augmentation for underrepresented groups, fairness-aware regularisation, and ongoing bias audits.

Authors

Files

Link of Paper