Verified interactive research releasev1.3.0

Move the population.
Trace every decision.

A model can keep the same code and still behave differently after deployment. Change one data-generating mechanism, then inspect performance, group gaps, and uncertainty, calibration, and the full decision curve together—then compare the live run with a frozen 300-experiment report—then compare eight transparent decision policies without hiding their costs or data requirements. A separate external-evidence area then tests those ideas against a governed historical reference table without calling it the real world. A synthetic Robustness Lab then stress-tests the whole apparatus—label noise, measurement error, an unobserved subgroup, small samples, and misspecification—across two inspectable model families.

One controlled interventionSource-only calibrationWhole threshold curve20-seed replication gridClaims boundedNo automatic best policy
01

Shift microscope

Change the mechanism,
not the story.

Distribution shift mechanism
Observed feature X₁Source → target
live simulation
Source populationTarget population
Target Group B share53%
Threshold0.50
Independent target seed45
Metric lens

Demographic parity gap

Question

Do groups receive positive predictions at different rates?

What this run supports

Zero means equal selection rates in this sample.

Do not overclaim

Parity may be inappropriate when relevant base rates differ; context is essential.

Intervals use 80 group-stratified percentile-bootstrap resamples in the browser. They describe sampling variability only—not label validity, model selection, or future drift.

02

Decision landscape

A score is not
yet a decision.

Source-fitted temperature1.32

Chosen on an independent source calibration sample by minimum log loss.

Target Brier score0.168

Mean squared probability error after source-only calibration. Lower is better.

Target ECE0.077

Weighted bin gap; before calibration: 0.054.

Critical readingWorsened here

Source calibration is not a guarantee under shift. Inspect the target evidence.

Target reliabilityPredicted probability → observed rate
RawCalibrated
Target reliability diagramRaw and source-calibrated predicted probabilities compared with observed outcome rates. The complete values follow in a table.0101
Closer to the dotted diagonal means predicted probabilities align more closely with observed frequencies in this sample.
Read reliability data
Target reliability bins
SeriesMean probabilityObserved rateCount
Raw0.0670.00022
Raw0.1520.13829
Raw0.2530.25547
Raw0.3520.18255
Raw0.4490.49367
Raw0.5580.58360
Raw0.6460.73884
Raw0.7510.72383
Raw0.8500.901101
Raw0.9380.96252
Calibrated0.0570.0005
Calibrated0.1510.03727
Calibrated0.2540.13238
Calibrated0.3510.25064
Calibrated0.4500.43086
Calibrated0.5580.62989
Calibrated0.6510.733101
Calibrated0.7510.84798
Calibrated0.8410.90775
Calibrated0.9321.00017
Threshold sensitivityDemographic parity gap across the decision curve
SourceTarget
Demographic parity gap threshold sensitivitySource and target measurements across 19 decision thresholds. The active threshold is 0.50. Complete values follow in a table.0101
The vertical marker is your current threshold. Select any metric card above to change the curve.
Read threshold data
Demographic parity gap by decision threshold
ThresholdSourceTarget
0.050.0100.004
0.100.0090.018
0.150.0070.023
0.200.0350.040
0.250.0420.057
0.300.0330.055
0.350.0010.064
0.400.0040.082
0.450.0420.077
0.500.0260.093
0.550.0230.120
0.600.0080.129
0.650.0270.110
0.700.0030.080
0.750.0180.077
0.800.0110.068
0.850.0160.030
0.900.0000.040
0.950.0000.003
Calibration asks

Do scores mean what they say?

Among cases scored near 0.7, roughly 70% should be positive in a well-calibrated sample.

Thresholding asks

Where does a probability become action?

Moving the cutoff changes selections and error rates—even when ranking remains fixed.

Fairness asks

Which differences matter here?

Calibration and error-rate parity can conflict when groups have different outcome rates.

ECE depends on binning and is descriptive. Temperature scaling preserves score ranking, so AUROC remains unchanged; it does not repair a shifted label mechanism. Target reliability can only be evaluated once representative target labels are observed.

03

Verified report

One grid.
No cherry-picking.

Three interventions × five magnitudes × 20 independent seeds. The frozen registry contains 300 complete experiments and 1.2 million generated rows before resampling.

20replications per cell
0.50fixed decision threshold
Magnitude 1.00 · descriptive result

The changed label mechanism produced the largest predictive and calibration failure.

Accuracy0.466-0.246 from source
AUROC0.451-0.333 from source
DP gap0.046-0.007 from source
Equalized-odds gap0.062-0.010 from source
ECE0.247+0.214 from source

What the grid supports: different controlled mechanisms produce materially different combinations of predictive, calibration, and group-gap behavior inside this generator.

What it cannot support: claims that a real system is fair, lawful, causally understood, or ready to decide about people.

04

Policy Studio

Declare the stakes.
See the trade-offs.

Eight decision policies are evaluated on the same independently sampled target tests. Choose an intervention and error-cost declaration; the studio exposes consequences but never chooses a policy for you.

Target intervention · magnitude 1.00
Cost of one false negative · false positive = 1Both errors cost the same
Benchmark-relative frontierDecision cost ↔ equalized-odds gap
Policy cost and equalized-odds comparisonEight policies compared by mean normalized expected error cost on the horizontal axis and mean equalized-odds gap on the vertical axis. Lower is better on both axes. Whiskers show empirical 2.5th to 97.5th percentile ranges across 20 seeded replications. Diamonds mark policies that are Pareto-efficient only within this benchmark and scenario.Normalized expected error cost · lower is betterEqualized-odds gap · lower is better
Lower-left is preferable only for these two declared measurements. It does not settle validity, rights, downstream harm, or which errors a real institution should value.
Selected policy

Fixed baseline

Data required: Source labels only
Source-calibrated score at threshold 0.50.

Normalized cost0.2100.1950.226
Equalized-odds gap0.2570.1260.396
Accuracy0.7900.7740.805
Mean thresholds · A=0 / A=10.500 / 0.500selected before target testing

Efficient within this comparison. This status uses mean cost and mean equalized-odds gap only. It is not a recommendation.

Read the complete policy comparison
Covariate shift at magnitude 1.00 with false-negative cost 1
PolicyData accessCost mean (range)Equalized-odds mean (range)Accuracy meanStatus
Fixed baselineSource labels only0.210 (0.1950.226)0.257 (0.1260.396)0.790Efficient here
Group thresholds · λ=0.1Source labels and protected attribute0.211 (0.1960.227)0.262 (0.1480.386)0.789Dominated here
Group thresholds · λ=0.3Source labels and protected attribute0.213 (0.1960.228)0.261 (0.1700.361)0.787Dominated here
Group thresholds · λ=1Source labels and protected attribute0.216 (0.1950.237)0.251 (0.1560.366)0.784Efficient here
Group thresholds · λ=3Source labels and protected attribute0.233 (0.1950.339)0.258 (0.1030.404)0.767Dominated here
Reweighed trainingSource labels and protected attribute0.211 (0.1890.226)0.276 (0.1710.364)0.789Dominated here
Cost-sensitive thresholdSource labels only0.211 (0.1950.226)0.260 (0.1310.402)0.789Dominated here
Target recalibrationRequires recent labeled target data0.210 (0.1940.223)0.273 (0.1470.400)0.790Dominated here

Target recalibration is an oracle-like adaptation benchmark. It uses a recent, representative, independently labeled target sample. Without that access, its result is unavailable—not a deployable promise.

Group thresholds use protected attributes at decision time. Their legality, legitimacy, construct validity, and impact depend on context that this synthetic laboratory cannot supply.

05

External evidence

History is evidence.
Not ground truth.

Observational boundary

This section uses a 1994 Census-derived UCI table. It is visually and analytically separate from the synthetic experiments above. The income label is not merit, qualification, need, or deservingness; the provider-coded binary sex field is not identity truth.

Read the admission gate ↗
Held-out records15,060Female-coded 4,913 · Male-coded 10,147
Baseline accuracy0.800Range 0.7990.800
Baseline equalized-odds gap0.037Descriptive split-seed range, not a justice score
Replications5Only the development/tuning split changes

Uncertainty before ranking

Mean · empirical 2.5th–97.5th split-seed range. Small differences can reverse under another declared cohort, missingness rule, cost, or split. No policy is selected.

All policies for the selected historical cohort and decision-cost declaration
PolicyAccuracyNormalized costDP gapEO gapEOdds gapBrierECE
Fixed baseline0.800 · 0.799–0.8000.200 · 0.200–0.2010.080 · 0.079–0.0810.037 · 0.036–0.0380.037 · 0.036–0.0380.139 · 0.139–0.1390.022 · 0.022–0.023
Group thresholds · λ=10.796 · 0.793–0.7970.204 · 0.203–0.2070.051 · 0.045–0.0560.017 · 0.013–0.0240.017 · 0.013–0.0240.139 · 0.139–0.1390.022 · 0.022–0.023
Reweighed training0.801 · 0.800–0.8020.199 · 0.198–0.2000.074 · 0.071–0.0820.021 · 0.014–0.0330.021 · 0.014–0.0330.139 · 0.139–0.1390.023 · 0.023–0.024
Cost-sensitive threshold0.800 · 0.799–0.8000.200 · 0.200–0.2010.082 · 0.079–0.0900.039 · 0.036–0.0460.039 · 0.036–0.0460.139 · 0.139–0.1390.022 · 0.022–0.023
06

Robustness Lab

Break the assumptions.
See what still holds.

Synthetic stress-test boundary

Every value on this page comes from the preregistered v1.3 specification-stress protocol on entirely synthetic populations. It is visually and analytically separate from both the base synthetic experiments above and the governed UCI Adult external evidence: it must never share an unlabeled chart, table, or scale with either. No protected attribute, label, or subgroup here describes a real person or community.

Read the preregistered protocol ↗
Stressor
Metric
Magnitude for the table and undefined-rate summary below
Model family for the undefined-rate summary below

Uncertainty and limitations come before the ranking

Group A=0 stays near a 5% flip rate; group A=1 rises to 45% at magnitude 1.00. At the selected magnitude and model family, 12 seeded replications defined 18 of 18 disaggregated rates in every replication, 0 were defined in only some replications, and 0 were undefined in every replication (an empty group, label, or subgroup cell — not a rate of zero). This tests: v1.0/v1.1: observed group-fairness gaps reflect the declared mechanism, not annotation artifacts.

Both model families · same stressEqualized-odds gap (observed grouping) across magnitude
Equalized-odds gap (observed grouping) under Group-conditional label noiseTwo inspectable model families, fit and thresholded identically, compared across five preregistered stress magnitudes for one stressor. Whiskers show the empirical 2.5th to 97.5th percentile replication range across seeded replications. The complete values follow in a table.Stressor magnitude · 0.00 is the no-stress controlEqualized-odds gap (observed grouping)
Neither shape marks a winner. A model family that degrades less on one metric can degrade more on another; both are reported at every magnitude.
Train/tuning · test samples2,000 · 2,000Disjoint seeds; test is evaluated once per replication.
Realized stress diagnostic0.254Measured on the adaptation split, never used for fitting or thresholding.
Selected global threshold0.492Range 0.4500.536
Observed vs. true grouping0.100 vs. 0.100Equalized-odds gap by recorded vs. ground-truth protected attribute.
Group-conditional label noise: every magnitude and model family
MagnitudeModel familyAccuracyNormalized costDP gapEO gapEOdds gapBrierECEStress diagnostic
0.00Logistic regression0.7120.2880.0300.0350.0430.1880.0340.050
0.00Shallow decision tree0.6870.3130.0900.1050.1190.2020.0270.050
0.25Logistic regression0.7120.2880.0240.0260.0380.1900.0500.101
0.25Shallow decision tree0.6860.3140.1030.1140.1240.2030.0330.101
0.50Logistic regression0.7130.2870.0370.0330.0530.1930.0700.151
0.50Shallow decision tree0.6820.3180.0860.0880.1060.2050.0470.151
0.75Logistic regression0.7120.2880.0470.0460.0630.1980.0880.203
0.75Shallow decision tree0.6770.3230.0980.1050.1270.2110.0610.203
1.00Logistic regression0.7050.2950.0770.0830.1000.2040.1070.254
1.00Shallow decision tree0.6280.3720.2970.2990.3950.2230.0560.254
Read disaggregated group and subgroup rates for the selected cell
Group-conditional label noise at magnitude 1.00, Logistic regression
GroupSelection rateTrue-positive rateFalse-positive rateDefined replications
group one excluding subgroup0.6780.8290.48512/12 · 12/12 · 12/12
observed group 00.4300.6460.23812/12 · 12/12 · 12/12
observed group 10.5060.7270.31612/12 · 12/12 · 12/12
subgroup0.3320.5670.20112/12 · 12/12 · 12/12
true group 00.4300.6460.23812/12 · 12/12 · 12/12
true group 10.5060.7270.31612/12 · 12/12 · 12/12

A null/"undefined" rate above means zero replications had any record in that group-label cell at this magnitude and model family — not that the rate is zero. This is this module's own explicit missing-value convention; it does not change how earlier releases report zero-denominator rates as 0.0.

07

Scientific method

See exactly what
the experiment knows.

1Generate

Sample an explicit structural process.

2Train once

Fit logistic regression on source data only.

3Calibrate

Fit one temperature on a separate source holdout.

4Intervene

Change exactly one target mechanism.

5Audit

Compare reliability, curves, gaps, and intervals.

Source structural equationP(Y=1) = σ(−0.15 + 1.1X₁ − 0.7X₂ − 0.45A)pᶜ = σ(logit(p) / T)

This equation is a transparent teaching mechanism—not a representation of any real demographic group. The protected attribute is synthetic and binary; intersectionality, institutions, measurement error, and lived context are outside this release.

Observation

The target gap changed.

A numerical statement about this generated sample.

Supported inference

The intervention altered measured behavior.

A reproducible statement inside the declared structural process.

Unsupported leap

“The model is fair in the real world.”

Not established by synthetic data or a single statistical definition.

08

Research trail

Built on arguments
you can inspect.

ICML · 2017On Calibration of Modern Neural Networks

Guo et al. evaluate post-hoc calibration and fit temperature scaling on held-out validation data.

NeurIPS · 2019Can You Trust Your Model’s Uncertainty?

Ovadia et al. show why calibration must be re-evaluated under dataset shift.

ITCS · 2017Inherent Trade-Offs in Fair Risk Scores

Kleinberg, Mullainathan & Raghavan formalize incompatibilities among fairness conditions.

NeurIPS · 2016Equality of Opportunity in Supervised Learning

Hardt, Price & Srebro formalize equal opportunity and equalized odds.

KAIS · 2012Data preprocessing without discrimination

Kamiran & Calders develop reweighing and related preprocessing approaches for discrimination-aware classification.

JMLR · 2023The Measure and Mismeasure of Fairness

Corbett-Davies et al. connect decision policy, utility, fairness criteria, and Pareto-dominated rules.

Research · 2021Pareto Efficient Fairness

Kamani et al. frame model loss and fairness criteria as a multi-objective frontier rather than one scalar answer.

FAccT · 2022Fairness Transferability Subject to Bounded Distribution Shift

Chen et al. study when statistical fairness can transfer across shifted distributions.

AISTATS · 2024Uncertainty Matters

Barrainkua et al. show why uncertainty-aware fairness comparisons matter.

ICML · 2025Optimal Fair Learning Robust to Adversarial Distribution Shift

Agarwal et al. analyze fairness-constrained learning under malicious distribution noise.

NIST AI RMF · 1.0AI Risks and Trustworthiness

Fairness belongs inside a broader socio-technical risk-management process.