00
Medical AI / Fairness

Medical AI
Bias Lab

Can a model look accurate overall while performing very differently across groups?

A single average can hide uneven performance. This experiment isolates one idea: how differences in representation can coincide with differences in simulated subgroup performance. Real medical AI systems are considerably more complicated.

Build the dataset.

0 experiments run
01 / Choose a scenario
02 / Adjust representation
Group A 34
Group B 33
Group C 33
Normalized dataset 100%

The sliders are treated as relative representation. Labs automatically normalizes them to 100%.

your results go here →

Change the dataset or choose a preset, then run the simulation.

Running simplified model simulation...

Overall performance — %
Fairness gap — pts

Look group by group.

Group A
—
Group B
—
Group C
—

Compare the experiments.

Balanced reference
—% / — pt gap
A simulated dataset with approximately equal representation.
Your dataset
—% / — pt gap

What changed?

Can you shrink the gap?

Adjust the simulated dataset and try to produce a fairness gap below eight percentage points while keeping overall performance above 80%.

Notice that optimizing an average and reducing differences between groups are related — but they aren't exactly the same objective.

Target: gap < 8 pts · overall > 80%

Recent experiments

Run A / B / C Overall Gap
No experiments yet.

A tiny glossary.

Representation

How much of a dataset comes from different groups or types of examples.

Subgroup performance

Evaluating a model separately on a particular subset instead of looking only at an average.

Overall performance

A combined metric across the full simulated dataset. It can conceal differences within it.

Fairness gap

In this Lab, the difference between the highest and lowest simulated subgroup performance.

Imbalance

A situation where some groups or examples appear much more frequently than others.

Simulation

A simplified model created to investigate an idea. Its numbers are not measurements from patients.

This Lab is an educational simulation, not a trained clinical model and not evidence of real-world performance. Its values are generated from a simplified mathematical relationship created for teaching. In real biomedical AI, disparities can arise from many interacting factors, including sampling, labels, measurement quality, disease prevalence, acquisition methods, model design, evaluation, and deployment conditions.