ALFIE ETD-HUB

2: What Metrics are Needed to Evidence Training Dataset is Unibased?

Asked: 7 months, 4 weeks ago By: Catalink Views: 258 Catalink Case Study: IRIS

What specific metrics do you believe could be documented to prove that the training data is representative of all driver demographics (e.g., age, gender, ethnicity, physical characteristics like glasses/hats) and is truly "unbiased"?

20 Answers

Answered: 6 months ago By: Chiamakaokorie

Age, gender, ethnicity, removal of accessories

Answered: 6 months ago By: Tundefasina

Key documented metrics could include:

Demographic distribution statistics (age, gender, ethnicity)

Subgroup performance metrics (precision, recall, FNR/FPR per group)

Fairness metrics (e.g., demographic parity, equal opportunity)

Data coverage matrices (lighting, occlusion, accessories like glasses/hats) These metrics help demonstrate balanced representation and consistent performance.

Deleuze replied: Definitely agree. We could look at for the demographic distribution stuff: Demographic coverage - Number and percentage of subjects by age band, gender, ethnicity/race, skin tone, disability-relevant facial characteristics where lawfully collected, and other relevant driver characteristics. - Shows whether key driver groups are present in the data, not merely assumed to be covered. Intersectional coverage - Counts for combined groups, for example older women with darker skin tones, younger men wearing glasses, drivers with facial hair and hats, etc. - Bias often appears at intersections, not just in single categories. Physical-characteristic coverage - Proportion of images/videos with glasses, sunglasses, hats, masks, facial hair, head coverings, different hairstyles, facial asymmetry, and other occlusions. - These factors can materially affect facial landmark detection and drowsiness classification.
Deleuze replied: For subgroup performance metrics I would also expand it to include all of sensitivity/recall, specificity, false-negative rate/false-positive rate (as you already mention), precision, AUROC/AUPRC, and calibration for each demographic and physical-characteristic group.
Answered: 6 months ago By: Zainabodogwu2

Document demographic coverage ratios vs. target population, per-subgroup performance metrics (TPR/FPR/FNR gaps), confidence intervals by subgroup, distributional similarity scores (e.g., KL divergence) between training and real-world data, and fairness deltas showing no statistically significant performance degradation across age, gender, ethnicity, or physical attributes.

Answered: 6 months ago By: Oliverharrow

I believe the name, age, gender and phone number should be documented

Answered: 6 months ago By: Ngozioshoba

To prove fairness, developers should clearly show who is represented in the training data and how the system performs for each group. This includes age, gender, ethnicity, and physical features like glasses or hats. Comparing accuracy across groups helps confirm the system works equally well for everyone and does not unintentionally favor certain users.

Answered: 6 months ago By: Efeadelaja

Documented metrics should include demographic coverage ratios, balanced class distributions, subgroup-specific accuracy/false-positive/false-negative rates, and fairness metrics (e.g., equalized odds) across age, gender, ethnicity, and physical attributes.

Deleuze replied: Sure, for subgroups what about minimum subgroup sample size. I.e. the predefined minimum number of individuals and drowsiness events per subgroup, with confidence intervals for each group.
Answered: 6 months ago By: Meilincai

Age, gender and race

Answered: 6 months ago By: Kelechinwosu

To prove the data is unbiased, you must document Proportional Representation (balancing age, gender, and ethnicity) and Attribute Parity, ensuring physical traits like glasses or hats are represented across all skin tones. The definitive metric is the Disparate Impact Ratio, which confirms that error rates remain equally low for every demographic group.

Answered: 6 months ago By: Beatricelorne

Equal proportions of people of all ethnicities, as well as equal proportions of clothing styles typical of the area of deployment.

Answered: 6 months ago By: Zainabodogwu32

Demographic distribution tables showing proportions of age groups, gender identities, ethnic backgrounds, and physical characteristics (e.g. glasses, facial hair, head coverings). Performance parity metrics, such as: False positive rate (FPR) and false negative rate (FNR) per demographic group. Accuracy, precision, and recall disaggregated by subgroup. Statistical fairness measures, such as: Difference in error rates between majority and minority groups. Confidence intervals to show robustness of results. Data provenance documentation, explaining where data originated, how it was collected, and known limitations. Synthetic data validation, demonstrating that synthetic samples meaningfully improve representation without introducing artefacts or amplifying bias. While “perfect neutrality” is unrealistic, regulators will expect evidence of active bias mitigation and continuous monitoring rather than mere assertions of fairness.

Answered: 6 months ago By: Miles_Hatcher

Physical characteristics and age

Answered: 6 months ago By: Aminaolorun

Gender

Answered: 6 months ago By: Clarawhitby

A dataset can only be considered “unbiased” if it shows demographic coverage parity, balanced representation, and equivalent safety performance (especially FNR) across all driver groups, supported by transparent documentation and independent audits.

Answered: 6 months ago By: Ifeanyiakare

Demographic Coverage Ratios Minimum Samples per Subgroup Intersectional Coverage Condition & Accessory Coverage Label Consistency Outcome Parity Metrics

Answered: 6 months ago By: Kunleekwueme

How diversified is the dataset being used to train the model.

Number of races Age groups Gender Physical characteristics

Answered: 6 months ago By: Sadeogunlana

A diverse, yet big enough sample size

Answered: 6 months ago By: Tomashbrook

Metrics that are willingly given by participants and that don't put their personal lives at risk. Their physical appearance, age, ethnicity are some examples of metrics that can be documented. Perhaps a log of how long they've been on the road can be included as well.

Your Answer

Login to add your answer!

We’d love to hear your thoughts — share a meaningful answer by logging in.