Power Laws
Power Laws: In many systems, a small number of causes produce a disproportionate fraction of effects. Unlike normal distributions (where averages are meaningful), power law distributions are dominated by extreme values β the top 1% matters more than the other 99% combined. This changes how you should invest, prioritize, and build strategies.
What Is a Power Law?β
A power law is a mathematical relationship between two quantities where one varies as a power of the other: y = kx^Ξ±. In distribution terms, a power law means that the frequency of an event decreases as a power function of its size β so very large events are rare, but they exist and they dominate the aggregate.
The signature of a power law distribution: the top item is much larger than the second, which is much larger than the third, and so on. On a log-log plot, a power law appears as a straight line β whereas a normal distribution appears as a bell curve on a linear plot.
Power laws are pervasive across natural and human systems:
Wealth: The Pareto principle originated from Vilfredo Pareto's observation that 80% of Italy's land was owned by 20% of the population. More extreme: the Forbes 400 richest Americans hold more wealth than the bottom 60% combined.
Cities: The largest city in a country is typically twice the size of the second-largest, which is twice the third-largest (Zipf's Law). New York, Los Angeles, Chicago.
Earthquakes: The Richter scale is logarithmic. A magnitude 8 earthquake releases 31.6x the energy of a 7, which releases 31.6x the energy of a 6.
Website traffic: A small number of pages receive the overwhelming majority of visits. The top 1% of YouTube videos account for the majority of views.
Startup outcomes: A small number of investments in a VC portfolio return more than the entire rest of the fund combined. The power law distribution of startup outcomes is the central fact of venture capital.
How It Worksβ
Normal Distribution (Bell Curve):
Most outcomes near the average
Extreme values exist but are very rare and bounded
Standard deviation is a meaningful measure
Example: human height, IQ, most biological measurements
Power Law Distribution:
Most outcomes are small; a tiny number are enormous
No characteristic scale β any "average" is misleading
Standard deviation may be undefined (infinite)
Examples: wealth, city size, earthquake magnitude, startup returns
Key implication: In power law domains, average β typical
VC example:
Fund A invests in 100 companies at $1M each.
90 return $0. 8 return $1M. 1 returns $10M. 1 returns $90M.
Total invested: $100M. Total returned: $0+8+10+90 = $108M.
Average return: 1.08x (barely positive)
The single best outcome ($90M) is 83% of all returns.
Without that outlier, the fund loses money.
In a normal distribution world: diversify to reduce variance.
In a power law world: concentrate on finding the outlier.
Three Real-World Examplesβ
Venture Capital and Power Lawsβ
Peter Thiel, in Zero to One, makes the power law argument explicitly: a venture fund's returns are dominated by its single best investment, which typically outperforms the rest of the portfolio combined. This has profound implications for how VCs should operate:
- Invest in companies that can potentially return the entire fund (not ones that can return 5β10x)
- Be willing to hold concentrated positions in winners rather than diversifying
- The selection of any one company matters more than the entire portfolio construction process
Andreessen Horowitz's first fund had WhatsApp as its single dominant performer. Sequoia's famous early Facebook investment returned more than the rest of many portfolios combined. In power law systems, the tails are the strategy.
The Long Tail (Content and Commerce)β
Chris Anderson's 2004 essay "The Long Tail" described how digital distribution changes the economics of power law distributions in content. In a physical store, only top-selling products are profitable to stock. Online, even products with tiny demand can be profitably served β the aggregation of low-demand items (the "long tail") can exceed the revenue from the "head" (blockbuster items).
Netflix's recommendation engine is designed to drive traffic to the long tail β not because the long tail is more profitable per title, but because the aggregate of the long tail represents significant revenue, and the ability to serve niche preferences is a differentiator unavailable to brick-and-mortar competitors.
Pareto Principle in Businessβ
The 80/20 rule (Pareto Principle) is a specific form of power law. In most businesses: 20% of customers generate 80% of revenue; 20% of products generate 80% of profit; 20% of salespeople close 80% of deals. These ratios are not always exactly 80/20, but some version of disproportionate concentration reliably applies.
The strategic implication: identify the 20% that drives 80% and design your organization around serving, retaining, and expanding it β rather than spreading resources uniformly. This is particularly powerful for customer segmentation (identify high-LTV customers and allocate customer success resources accordingly) and product prioritization (identify the 20% of features used by 80% of customers and make those best-in-class before adding new ones).
When to Use Itβ
β Use Power Law thinking when:
- Making investment decisions where a few outcomes will dominate returns
- Prioritizing resources (customer segments, product features, marketing channels)
- Building distribution strategies for content or products
- Understanding competitive dynamics in winner-take-most markets
- Interpreting data: if the mean is much larger than the median, suspect a power law
β Be cautious when:
- The domain is genuinely normally distributed (physical measurements, most biological data)
- Using power law thinking to justify ignoring average performance β not all decisions are power law domains
| Pairs well with | Why |
|---|---|
| Pareto Principle | The Pareto Principle is the business application of power laws |
| Expected Value | In power law domains, EV calculations are dominated by rare high-value outcomes |
| Matthew Effect | Matthew Effect is the mechanism that generates power law distributions over time |
Common Misusesβ
Applying power law thinking to normal distribution domains. Not all systems are power law distributed. Applying "find the outlier" strategy to employee management (hiring normal curves) or inventory management (normally distributed demand) produces bad decisions.
Confusing skewed distributions with power laws. Many distributions are skewed (the right tail is longer than the left) without being power law distributions. The specific mathematical form matters for the strongest claims about power law behavior.
Common Misuses and Limitationsβ
Assuming all distributions are normal. The core error that power laws correct: most statistical training uses normal (Gaussian) distributions, which have thin tails. Power law distributions have fat tails β extreme events are far more common than Gaussian models predict. Applying normal distribution thinking to power law domains (wealth, earthquake magnitude, word frequency, city size) systematically underestimates extreme outcomes.
Fitting power laws to data that isn't power-law distributed. Power law distributions look roughly linear on a log-log plot, but so do several other fat-tailed distributions. Proper statistical testing (not just visual inspection of log-log plots) is required to confirm a power law. Clauset et al. (2009) found that many claimed power laws in empirical data don't survive rigorous testing.
Using power laws to justify winner-take-all thinking prematurely. Not all competitive markets converge to power law distributions. Local businesses, professional services, and many B2B markets sustain multiple viable competitors. Power law dynamics are strongest where network effects, information asymmetries, or scale advantages are large. Apply the model where the underlying mechanism (preferential attachment, returns to scale) is actually present.
Ignoring the exponent. Power laws differ in how extreme their tails are. A power law with exponent Ξ±=2 has much fatter tails than one with Ξ±=3. The practical consequences (how extreme are the most extreme events?) depend critically on the exponent. Treating all power laws as equivalent ignores this crucial variation.
Related Modelsβ
| Model | Relationship |
|---|---|
| Matthew Effect | Matthew Effect is the generative mechanism that produces power law distributions |
| Network Effects | Network effects create power law competitive dynamics through preferential attachment |
| Black Swan | Black Swans are the extreme tail events that power law distributions make much more probable than Gaussian models |
| Pareto Principle | Pareto's 80/20 rule is a special case of power law distribution |
Frequently Asked Questionsβ
What is "preferential attachment" and how does it generate power laws?
Preferential attachment (also called the "rich-get-richer" or "cumulative advantage" mechanism) is a process where new connections are proportionally more likely to attach to already-highly-connected nodes. In citation networks, heavily-cited papers attract more citations. In social networks, popular accounts attract more followers. In cities, larger cities attract more migrants. This mechanism β formalised by BarabΓ‘si and Albert in 1999 as a model of network growth β naturally produces power law degree distributions. Once established, power law distributions are self-perpetuating: the mechanism that produced them continues to operate.
How should investors think about power laws?
Venture capital is explicitly a power law business. Andreessen Horowitz has noted that in VC, the best investment in a fund typically returns more than the entire rest of the fund combined β a textbook power law distribution. This has two implications: (1) portfolio construction β make enough bets to have a reasonable chance of capturing the power law winner; (2) selection β prioritise investing in companies with the potential to be the power law outlier (typically through network effects, platform dynamics, or exponential market growth). The firms that try to optimise for average returns across their portfolio consistently underperform those that accept high variance in pursuit of the tail.
What is Zipf's Law and how does it relate to power laws?
Zipf's Law is a specific power law observed in word frequency: the most common word in any natural language corpus appears approximately twice as often as the second most common word, three times as often as the third, and so on. The frequency of any word is inversely proportional to its rank (rank 1, rank 2, rank 3...). Remarkably, the same Zipf distribution appears in city sizes, income distributions, protein interaction networks, and many other seemingly unrelated phenomena. Zipf's Law is one of the most universally documented power laws and its recurrence across domains suggests a very general underlying mechanism of self-organisation.
Further Readingβ
- BarabΓ‘si, A.L. (2002). Linked: The New Science of Networks β accessible introduction to power laws and network science
- Taleb, N.N. (2007). The Black Swan β the consequences of fat-tailed (power law) distributions
- Newman, M.E.J. (2005). "Power Laws, Pareto Distributions and Zipf's Law." Contemporary Physics
Apply with AIβ
π Apply power law thinking to your strategy in MindMax β
Further Readingβ
- Peter Thiel, Zero to One (2014) β Chapter 7 on power laws in venture capital.
- Chris Anderson, "The Long Tail" (Wired, 2004) β The original article.
- Nassim Taleb, The Black Swan (2007) β The distinction between "Mediocristan" (normal distributions) and "Extremistan" (power law distributions).
This page is part of the MindMax Mental Models Knowledge Base.