Medical AI
Bias Lab
Can a model look accurate overall while performing very differently across groups?
A single average can hide uneven performance. This experiment isolates one idea: how differences in representation can coincide with differences in simulated subgroup performance. Real medical AI systems are considerably more complicated.
Build the dataset.
The sliders are treated as relative representation. Labs automatically normalizes them to 100%.
Change the dataset or choose a preset, then run the simulation.
Running simplified model simulation...
Look group by group.
Compare the experiments.
What changed?
Can you shrink the gap?
Adjust the simulated dataset and try to produce a fairness gap below eight percentage points while keeping overall performance above 80%.
Notice that optimizing an average and reducing differences between groups are related — but they aren't exactly the same objective.
Recent experiments
| Run | A / B / C | Overall | Gap |
|---|---|---|---|
| No experiments yet. | |||
A tiny glossary.
Representation
How much of a dataset comes from different groups or types of examples.
Subgroup performance
Evaluating a model separately on a particular subset instead of looking only at an average.
Overall performance
A combined metric across the full simulated dataset. It can conceal differences within it.
Fairness gap
In this Lab, the difference between the highest and lowest simulated subgroup performance.
Imbalance
A situation where some groups or examples appear much more frequently than others.
Simulation
A simplified model created to investigate an idea. Its numbers are not measurements from patients.
This Lab is an educational simulation, not a trained clinical model and not evidence of real-world performance. Its values are generated from a simplified mathematical relationship created for teaching. In real biomedical AI, disparities can arise from many interacting factors, including sampling, labels, measurement quality, disease prevalence, acquisition methods, model design, evaluation, and deployment conditions.