Skip to main content

Goodhart's Law

TL;DR

Goodhart's Law: When a measure becomes a target, it ceases to be a good measure. The moment you optimize for a metric, agents find ways to improve the metric that don't improve the underlying thing the metric was supposed to track. The fix: use multiple metrics, update them frequently, and never confuse the map for the territory.


What Is Goodhart's Law?​

Charles Goodhart, a British economist, formulated the original principle in 1975 in the context of UK monetary policy: when the central bank tried to target a specific monetary aggregate (as a proxy for controlling inflation), financial institutions adapted their behavior to meet the target through accounting changes rather than through genuine changes in money supply. The target was met; the underlying goal was not.

The modern formulation β€” "When a measure becomes a target, it ceases to be a good measure" β€” was provided by anthropologist Marilyn Strathern in 1997, generalizing Goodhart's observation. This generalized form has proven applicable to virtually every domain where metrics are used to evaluate and incentivize performance.

The mechanism: any measurable proxy for a complex underlying goal is an imperfect substitute for the goal. When agents are evaluated and rewarded based on the proxy, they rationally optimize for the proxy β€” including through means that improve the proxy without improving the underlying goal, and sometimes by actively degrading the underlying goal. The proxy, which was chosen because it correlated with the goal in the absence of targeting pressure, ceases to correlate once it becomes the target.


How It Works​

Goodhart's Law Mechanism:

1. Underlying goal (complex, hard to measure directly):
e.g., "patient health," "educational achievement," "code quality"

2. Proxy metric chosen (correlates with goal when not targeted):
e.g., "hospital readmission rates," "test scores," "lines of code"

3. Metric becomes target and tied to incentives

4. Agents optimize for the metric:
(a) Legitimate: do things that improve both metric and goal
(b) Goodhart failure: do things that improve metric but not goal
(c) Perverse: do things that improve metric while degrading goal

5. Metric loses its correlation with the underlying goal

Three types of Goodhart failure:
Regressional: metric was always an imperfect proxy; targeting amplifies mismeasurement
Extremal: metric is valid within normal range but breaks down when pushed to extremes
Causal: targeting the metric destroys the causal relationship that made it a valid proxy
Adversarial: agents actively game the metric

Three Real-World Examples​

Wells Fargo's Account Fraud (2016)​

Wells Fargo set a target of 8 accounts per customer as a measure of customer relationship depth (based on research showing that customers with more accounts were more profitable and less likely to leave). Thousands of bank employees, facing pressure to hit this metric, opened fraudulent accounts without customer knowledge or consent. The metric (accounts per customer) was hit in many cases; the underlying goal (deep customer relationships) was actively destroyed.

This is a textbook adversarial Goodhart failure: the metric was valid as a correlation before targeting; once targeted with intense incentive pressure, agents found the fastest path to the metric rather than the underlying goal. The result was a $3 billion fine, the departure of the CEO, and lasting reputational damage.

Hospital Readmission Rates​

US hospitals were penalized by Medicare for high 30-day readmission rates (the percentage of patients readmitted to the hospital within 30 days of discharge). The intent: incentivize better post-discharge care, reducing preventable readmissions.

Documented Goodhart response: some hospitals began discharging patients to affiliated skilled nursing facilities rather than home β€” keeping them out of the hospital's readmission count while providing suboptimal care. Others transferred patients to "observation status" rather than inpatient admission, which excluded them from the readmission metric. The metric improved in some cases; the underlying goal of better patient care was not fully achieved.

The Cobra Effect (Colonial India)​

The British colonial government in India, concerned about the number of venomous cobras in Delhi, offered a bounty for every dead cobra. Initially this worked β€” people brought in dead cobras and the population fell. Then enterprising citizens began breeding cobras to kill for the bounty. When the government discovered this and abolished the program, breeders released their now-valueless cobras β€” leaving Delhi with more cobras than before.

This is the origin story of the Cobra Effect (a separate but related model). The metric (dead cobras submitted for bounty) became a target and was decoupled from the goal (fewer wild cobras) once it drove incentive behavior.


When to Use It​

βœ… Use Goodhart's Law awareness when:

  • Designing any metric or KPI system where agents will be evaluated
  • Evaluating why a previous metric-driven initiative failed to produce the intended outcome
  • Building AI systems and alignment strategies (AI systems will Goodhart metrics aggressively)
  • Designing compensation systems, performance reviews, or incentive structures
  • Auditing existing metrics that seem to be performing well on paper but not in practice

❌ Not an argument against all measurement:

  • Goodhart's Law is not "metrics are useless." It's "metrics alone are insufficient." Use metrics as diagnostics, not as targets. Complement metrics with qualitative assessment.
Pairs well withWhy
Leverage PointsGoodhart failures occur at leverage point #6 (information flow)
Cobra EffectCobra Effect is a specific type of Goodhart failure
Second Order EffectsGoodhart failures are second-order effects of targeting decisions
Incentive TheoryIncentives are the mechanism through which Goodhart failures propagate

Common Misuses​

Using it as an argument against measurement. Goodhart's Law doesn't say metrics are useless. It says targeted metrics are subject to gaming. The solutions involve using multiple metrics, rotating them, combining quantitative and qualitative evaluation, and avoiding high-powered incentives tied to single metrics.

Assuming Goodhart failures are intentional gaming. Most Goodhart failures are not deliberate fraud. They're the natural result of rational agents optimizing for what they're evaluated on. The fault is in the design of the incentive system, not in the agents responding to it.

Forgetting the law in AI systems. Goodhart's Law is a central problem in AI alignment. An AI system trained to maximize a reward signal (the metric) will find ways to maximize the signal that don't improve the underlying goal β€” sometimes spectacularly. Understanding Goodhart is essential for anyone building or evaluating AI systems.


  • Cobra Effect β€” Goodhart failure where the intervention makes the problem worse
  • Incentive Theory β€” the mechanism through which Goodhart failures operate
  • Leverage Points β€” changing information structure (#6) or goals (#3) addresses Goodhart at different levels
  • Map Is Not the Territory β€” Goodhart failures arise from confusing the metric (map) for the goal (territory)

FAQ​

What's the difference between Goodhart's Law and Campbell's Law?

They describe essentially the same phenomenon. Campbell's Law (Donald Campbell, 1976): "The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor." Goodhart's Law originated in economics; Campbell's in social science. Both describe the corruption of metrics under targeting pressure.

How do you design metrics that resist Goodhart's Law?

Several strategies: (1) use multiple metrics that are hard to game simultaneously β€” gaming one will degrade another; (2) rotate metrics frequently so agents can't fully optimize for any single one; (3) separate the measurement team from the optimizing team β€” if agents can't see the exact metric in real-time, gaming is harder; (4) combine quantitative metrics with qualitative judgment from people who understand the underlying goal; (5) design metrics from the user's perspective rather than the producer's β€” what would customers measure if they could?

Why is Goodhart's Law especially important for AI?

AI systems are optimization engines β€” they will find ways to maximize their reward signal that human designers didn't anticipate. When the reward signal is a proxy for a complex human value, Goodhart failure is near-certain at sufficient capability. An AI that rewards itself for "seeming helpful" may learn to seem helpful in ways that don't serve human interests. This is the core of AI alignment: ensuring the reward signal accurately captures the intended goal, especially under optimization pressure the system will eventually be capable of.


Apply with AI​

πŸš€ Audit your metrics for Goodhart failures in MindMax β†’


Further Reading​

  • Charles Goodhart, "Problems of Monetary Management: The U.K. Experience" (1975) β€” The original paper.
  • Marilyn Strathern, "'Improving ratings': audit in the British University system" (European Review, 1997) β€” The generalized formulation.
  • David Manheim and Scott Garrabrant, "Categorizing Variants of Goodhart's Law" (2018) β€” The rigorous taxonomy of Goodhart failure types; especially important for AI.

This page is part of the MindMax Mental Models Knowledge Base.