Ethics

Algorithmic Bias

A label may record who was noticed rather than what actually happened.

Suppose a system predicts complaints. It has learned from records of who was reported, by whom, and under what conditions. Calling its output a prediction of wrongdoing changes the claim without changing the model.

Begin there: what does the target actually record?

Then ask who is represented, where errors occur, and how the deployment differs from the setting in which the system was evaluated. Removing a sensitive field does not necessarily remove information about it; other features may act as proxies. A good aggregate result can conceal a severe problem for a smaller group.

A metric cannot choose the good

For imperfect predictions, some common fairness criteria conflict when relevant outcome rates differ between groups. Chouldechova’s linked analysis examines such a conflict. It does not relieve an institution of the work of choosing and defending its standard.

Equal error rates, equal predictive value, and a fair procedure answer different questions. Say which question a measurement answers before offering it as proof that a system is fair.

Compare with the real alternative, not a perfect imaginary human. Human judgment can be biased and difficult to audit too. An improvement is relevant; it is not an exemption from asking who still loses and what recourse they have.

The person receiving a decision should not have to perform a population-level study before a relevant error about their own case is heard. Evaluation and appeal are different protections. A responsible deployment needs both.