Demographic parity gap
Do groups receive positive predictions at different rates?
Zero means equal selection rates in this sample.
Parity may be inappropriate when relevant base rates differ; context is essential.
A model can keep the same code and still behave differently after deployment. Change one data-generating mechanism, then inspect performance, group gaps, and uncertainty, calibration, and the full decision curve together—then compare the live run with a frozen 300-experiment report—then compare eight transparent decision policies without hiding their costs or data requirements. A separate external-evidence area then tests those ideas against a governed historical reference table without calling it the real world. A synthetic Robustness Lab then stress-tests the whole apparatus—label noise, measurement error, an unobserved subgroup, small samples, and misspecification—across two inspectable model families.
Shift microscope
Do groups receive positive predictions at different rates?
Zero means equal selection rates in this sample.
Parity may be inappropriate when relevant base rates differ; context is essential.
Intervals use 80 group-stratified percentile-bootstrap resamples in the browser. They describe sampling variability only—not label validity, model selection, or future drift.
Decision landscape
Chosen on an independent source calibration sample by minimum log loss.
Mean squared probability error after source-only calibration. Lower is better.
Weighted bin gap; before calibration: 0.054.
Source calibration is not a guarantee under shift. Inspect the target evidence.
| Series | Mean probability | Observed rate | Count |
|---|---|---|---|
| Raw | 0.067 | 0.000 | 22 |
| Raw | 0.152 | 0.138 | 29 |
| Raw | 0.253 | 0.255 | 47 |
| Raw | 0.352 | 0.182 | 55 |
| Raw | 0.449 | 0.493 | 67 |
| Raw | 0.558 | 0.583 | 60 |
| Raw | 0.646 | 0.738 | 84 |
| Raw | 0.751 | 0.723 | 83 |
| Raw | 0.850 | 0.901 | 101 |
| Raw | 0.938 | 0.962 | 52 |
| Calibrated | 0.057 | 0.000 | 5 |
| Calibrated | 0.151 | 0.037 | 27 |
| Calibrated | 0.254 | 0.132 | 38 |
| Calibrated | 0.351 | 0.250 | 64 |
| Calibrated | 0.450 | 0.430 | 86 |
| Calibrated | 0.558 | 0.629 | 89 |
| Calibrated | 0.651 | 0.733 | 101 |
| Calibrated | 0.751 | 0.847 | 98 |
| Calibrated | 0.841 | 0.907 | 75 |
| Calibrated | 0.932 | 1.000 | 17 |
| Threshold | Source | Target |
|---|---|---|
| 0.05 | 0.010 | 0.004 |
| 0.10 | 0.009 | 0.018 |
| 0.15 | 0.007 | 0.023 |
| 0.20 | 0.035 | 0.040 |
| 0.25 | 0.042 | 0.057 |
| 0.30 | 0.033 | 0.055 |
| 0.35 | 0.001 | 0.064 |
| 0.40 | 0.004 | 0.082 |
| 0.45 | 0.042 | 0.077 |
| 0.50 | 0.026 | 0.093 |
| 0.55 | 0.023 | 0.120 |
| 0.60 | 0.008 | 0.129 |
| 0.65 | 0.027 | 0.110 |
| 0.70 | 0.003 | 0.080 |
| 0.75 | 0.018 | 0.077 |
| 0.80 | 0.011 | 0.068 |
| 0.85 | 0.016 | 0.030 |
| 0.90 | 0.000 | 0.040 |
| 0.95 | 0.000 | 0.003 |
Among cases scored near 0.7, roughly 70% should be positive in a well-calibrated sample.
Moving the cutoff changes selections and error rates—even when ranking remains fixed.
Calibration and error-rate parity can conflict when groups have different outcome rates.
ECE depends on binning and is descriptive. Temperature scaling preserves score ranking, so AUROC remains unchanged; it does not repair a shifted label mechanism. Target reliability can only be evaluated once representative target labels are observed.
Verified report
Three interventions × five magnitudes × 20 independent seeds. The frozen registry contains 300 complete experiments and 1.2 million generated rows before resampling.
What the grid supports: different controlled mechanisms produce materially different combinations of predictive, calibration, and group-gap behavior inside this generator.
What it cannot support: claims that a real system is fair, lawful, causally understood, or ready to decide about people.
Policy Studio
Eight decision policies are evaluated on the same independently sampled target tests. Choose an intervention and error-cost declaration; the studio exposes consequences but never chooses a policy for you.
Data required: Source labels only
Source-calibrated score at threshold 0.50.
Efficient within this comparison. This status uses mean cost and mean equalized-odds gap only. It is not a recommendation.
| Policy | Data access | Cost mean (range) | Equalized-odds mean (range) | Accuracy mean | Status |
|---|---|---|---|---|---|
| Fixed baseline | Source labels only | 0.210 (0.195–0.226) | 0.257 (0.126–0.396) | 0.790 | Efficient here |
| Group thresholds · λ=0.1 | Source labels and protected attribute | 0.211 (0.196–0.227) | 0.262 (0.148–0.386) | 0.789 | Dominated here |
| Group thresholds · λ=0.3 | Source labels and protected attribute | 0.213 (0.196–0.228) | 0.261 (0.170–0.361) | 0.787 | Dominated here |
| Group thresholds · λ=1 | Source labels and protected attribute | 0.216 (0.195–0.237) | 0.251 (0.156–0.366) | 0.784 | Efficient here |
| Group thresholds · λ=3 | Source labels and protected attribute | 0.233 (0.195–0.339) | 0.258 (0.103–0.404) | 0.767 | Dominated here |
| Reweighed training | Source labels and protected attribute | 0.211 (0.189–0.226) | 0.276 (0.171–0.364) | 0.789 | Dominated here |
| Cost-sensitive threshold | Source labels only | 0.211 (0.195–0.226) | 0.260 (0.131–0.402) | 0.789 | Dominated here |
| Target recalibration | Requires recent labeled target data | 0.210 (0.194–0.223) | 0.273 (0.147–0.400) | 0.790 | Dominated here |
Target recalibration is an oracle-like adaptation benchmark. It uses a recent, representative, independently labeled target sample. Without that access, its result is unavailable—not a deployable promise.
Group thresholds use protected attributes at decision time. Their legality, legitimacy, construct validity, and impact depend on context that this synthetic laboratory cannot supply.
External evidence
This section uses a 1994 Census-derived UCI table. It is visually and analytically separate from the synthetic experiments above. The income label is not merit, qualification, need, or deservingness; the provider-coded binary sex field is not identity truth.
Read the admission gate ↗Mean · empirical 2.5th–97.5th split-seed range. Small differences can reverse under another declared cohort, missingness rule, cost, or split. No policy is selected.
| Policy | Accuracy | Normalized cost | DP gap | EO gap | EOdds gap | Brier | ECE |
|---|---|---|---|---|---|---|---|
| Fixed baseline | 0.800 · 0.799–0.800 | 0.200 · 0.200–0.201 | 0.080 · 0.079–0.081 | 0.037 · 0.036–0.038 | 0.037 · 0.036–0.038 | 0.139 · 0.139–0.139 | 0.022 · 0.022–0.023 |
| Group thresholds · λ=1 | 0.796 · 0.793–0.797 | 0.204 · 0.203–0.207 | 0.051 · 0.045–0.056 | 0.017 · 0.013–0.024 | 0.017 · 0.013–0.024 | 0.139 · 0.139–0.139 | 0.022 · 0.022–0.023 |
| Reweighed training | 0.801 · 0.800–0.802 | 0.199 · 0.198–0.200 | 0.074 · 0.071–0.082 | 0.021 · 0.014–0.033 | 0.021 · 0.014–0.033 | 0.139 · 0.139–0.139 | 0.023 · 0.023–0.024 |
| Cost-sensitive threshold | 0.800 · 0.799–0.800 | 0.200 · 0.200–0.201 | 0.082 · 0.079–0.090 | 0.039 · 0.036–0.046 | 0.039 · 0.036–0.046 | 0.139 · 0.139–0.139 | 0.022 · 0.022–0.023 |
Robustness Lab
Every value on this page comes from the preregistered v1.3 specification-stress protocol on entirely synthetic populations. It is visually and analytically separate from both the base synthetic experiments above and the governed UCI Adult external evidence: it must never share an unlabeled chart, table, or scale with either. No protected attribute, label, or subgroup here describes a real person or community.
Read the preregistered protocol ↗Group A=0 stays near a 5% flip rate; group A=1 rises to 45% at magnitude 1.00. At the selected magnitude and model family, 12 seeded replications defined 18 of 18 disaggregated rates in every replication, 0 were defined in only some replications, and 0 were undefined in every replication (an empty group, label, or subgroup cell — not a rate of zero). This tests: v1.0/v1.1: observed group-fairness gaps reflect the declared mechanism, not annotation artifacts.
| Magnitude | Model family | Accuracy | Normalized cost | DP gap | EO gap | EOdds gap | Brier | ECE | Stress diagnostic |
|---|---|---|---|---|---|---|---|---|---|
| 0.00 | Logistic regression | 0.712 | 0.288 | 0.030 | 0.035 | 0.043 | 0.188 | 0.034 | 0.050 |
| 0.00 | Shallow decision tree | 0.687 | 0.313 | 0.090 | 0.105 | 0.119 | 0.202 | 0.027 | 0.050 |
| 0.25 | Logistic regression | 0.712 | 0.288 | 0.024 | 0.026 | 0.038 | 0.190 | 0.050 | 0.101 |
| 0.25 | Shallow decision tree | 0.686 | 0.314 | 0.103 | 0.114 | 0.124 | 0.203 | 0.033 | 0.101 |
| 0.50 | Logistic regression | 0.713 | 0.287 | 0.037 | 0.033 | 0.053 | 0.193 | 0.070 | 0.151 |
| 0.50 | Shallow decision tree | 0.682 | 0.318 | 0.086 | 0.088 | 0.106 | 0.205 | 0.047 | 0.151 |
| 0.75 | Logistic regression | 0.712 | 0.288 | 0.047 | 0.046 | 0.063 | 0.198 | 0.088 | 0.203 |
| 0.75 | Shallow decision tree | 0.677 | 0.323 | 0.098 | 0.105 | 0.127 | 0.211 | 0.061 | 0.203 |
| 1.00 | Logistic regression | 0.705 | 0.295 | 0.077 | 0.083 | 0.100 | 0.204 | 0.107 | 0.254 |
| 1.00 | Shallow decision tree | 0.628 | 0.372 | 0.297 | 0.299 | 0.395 | 0.223 | 0.056 | 0.254 |
| Group | Selection rate | True-positive rate | False-positive rate | Defined replications |
|---|---|---|---|---|
| group one excluding subgroup | 0.678 | 0.829 | 0.485 | 12/12 · 12/12 · 12/12 |
| observed group 0 | 0.430 | 0.646 | 0.238 | 12/12 · 12/12 · 12/12 |
| observed group 1 | 0.506 | 0.727 | 0.316 | 12/12 · 12/12 · 12/12 |
| subgroup | 0.332 | 0.567 | 0.201 | 12/12 · 12/12 · 12/12 |
| true group 0 | 0.430 | 0.646 | 0.238 | 12/12 · 12/12 · 12/12 |
| true group 1 | 0.506 | 0.727 | 0.316 | 12/12 · 12/12 · 12/12 |
A null/"undefined" rate above means zero replications had any record in that group-label cell at this magnitude and model family — not that the rate is zero. This is this module's own explicit missing-value convention; it does not change how earlier releases report zero-denominator rates as 0.0.
Scientific method
Sample an explicit structural process.
Fit logistic regression on source data only.
Fit one temperature on a separate source holdout.
Change exactly one target mechanism.
Compare reliability, curves, gaps, and intervals.
P(Y=1) = σ(−0.15 + 1.1X₁ − 0.7X₂ − 0.45A)pᶜ = σ(logit(p) / T)This equation is a transparent teaching mechanism—not a representation of any real demographic group. The protected attribute is synthetic and binary; intersectionality, institutions, measurement error, and lived context are outside this release.
A numerical statement about this generated sample.
A reproducible statement inside the declared structural process.
Not established by synthetic data or a single statistical definition.
Research trail
Guo et al. evaluate post-hoc calibration and fit temperature scaling on held-out validation data.
↗NeurIPS · 2019Can You Trust Your Model’s Uncertainty?Ovadia et al. show why calibration must be re-evaluated under dataset shift.
↗ITCS · 2017Inherent Trade-Offs in Fair Risk ScoresKleinberg, Mullainathan & Raghavan formalize incompatibilities among fairness conditions.
↗NeurIPS · 2016Equality of Opportunity in Supervised LearningHardt, Price & Srebro formalize equal opportunity and equalized odds.
↗KAIS · 2012Data preprocessing without discriminationKamiran & Calders develop reweighing and related preprocessing approaches for discrimination-aware classification.
↗JMLR · 2023The Measure and Mismeasure of FairnessCorbett-Davies et al. connect decision policy, utility, fairness criteria, and Pareto-dominated rules.
↗Research · 2021Pareto Efficient FairnessKamani et al. frame model loss and fairness criteria as a multi-objective frontier rather than one scalar answer.
↗FAccT · 2022Fairness Transferability Subject to Bounded Distribution ShiftChen et al. study when statistical fairness can transfer across shifted distributions.
↗AISTATS · 2024Uncertainty MattersBarrainkua et al. show why uncertainty-aware fairness comparisons matter.
↗ICML · 2025Optimal Fair Learning Robust to Adversarial Distribution ShiftAgarwal et al. analyze fairness-constrained learning under malicious distribution noise.
↗NIST AI RMF · 1.0AI Risks and TrustworthinessFairness belongs inside a broader socio-technical risk-management process.
↗