This project demonstrates a machine learning pipeline for addressing fairness concerns in classification tasks. Specifically, it implements a custom FairTeacherStudentClassifier, designed to identify and mitigate performance disparities across sensitive groups (e.g., race). The classifier trains a "student model" to account for the shortcomings of a "teacher model" on underrepresented or disadvantaged groups.
The dataset used in this project is the Adult Census Income dataset, provided by the fairlearn library. The goal is to predict whether an individual earns more than $50K per year (class), using demographic and economic features (e.g., age, race, education).
- A FairTeacherStudentClassifier that:
- Trains a teacher model to predict outcomes.
- Identifies performance weaknesses of the teacher model across sensitive groups.
- Adjusts the student model training by assigning weights to groups where the teacher performs poorly.
- A CurriculumStudentTeacher model that:
- Utilizes a curriculum-based approach where the student model learns from transformed teacher predictions.
- Allows incremental training of the student model, making it adaptive to different datasets.
- A WeightedCurriculumStudentTeacher model that:
- Incorporates group-specific weights based on the teacher model’s accuracy across different subgroups.
- Supports different weighting strategies to enhance fairness in model performance.
- Comparison of these models against a baseline model (RandomForestClassifier).
- Evaluation of fairness through accuracy metrics for each racial group.
- Python 3.8+
- Required libraries:
fairlearnscikit-learnpandasnumpy
- Clone this repository:
git clone https://github.com/yourusername/fairness-ml.git cd fairness-ml - Install dependencies:
pip install fairlearn scikit-learn pandas numpy
Run the Python script to preprocess the data, train models, and evaluate performance:
python fairness_pipeline.pyThe script outputs a DataFrame comparing the accuracy of:
- The FairTeacherStudentClassifier.
- A Baseline RandomForestClassifier.
The comparison is broken down by sensitive groups (e.g., race).
The Adult Census Income dataset contains 48,842 samples with features such as:
age,workclass,education,occupation, etc.- Sensitive attribute: race.
- Teacher Model:
- A base model is trained on the data.
- Predictions are made on the training set.
- Performance Analysis:
- Sensitive groups are evaluated to identify where the teacher performs poorly.
- Weights are assigned to these groups inversely proportional to their performance.
- Student Model:
- A student model is trained on the teacher’s predictions, using the group-specific weights.
You can replace the default RandomForestClassifier with any other scikit-learn classifier (e.g., SVC, GradientBoostingClassifier).
The project is designed to evaluate fairness based on the race attribute. However, you can easily change this to any other feature (e.g., sex or education) by modifying the z variable.
- The FairTeacherStudentClassifier consistently improves accuracy for underrepresented or disadvantaged groups compared to the baseline.
- It highlights the importance of fairness-aware training to mitigate bias in machine learning models.
- Fork the repository.
- Create a new branch:
git checkout -b feature-branch
- Commit your changes:
git commit -m "Add new feature" - Push the branch:
git push origin feature-branch
- Create a Pull Request.
This project is licensed under the MIT License.
Feel free to reach out for further clarifications or to report issues!