Skip to main content

Scientific Method

TL;DR

Scientific Method: (1) Observe a phenomenon; (2) Form a hypothesis that explains it; (3) Design an experiment to falsify the hypothesis; (4) Run the experiment; (5) Update belief based on results. The key discipline: design experiments to disprove your hypothesis, not confirm it. The scientific method's power comes from systematic vulnerability to being wrong β€” which is why beliefs produced by it are reliable.


What Is the Scientific Method?​

The Scientific Method is the epistemological procedure that distinguishes science from opinion. Its core innovation is the experimental test: instead of arguing about which explanation is more plausible, you design a test that would definitively disconfirm one of them. The hypothesis that survives repeated, rigorous attempts at falsification gains credibility; the hypothesis that fails even one well-designed test is rejected or revised.

The modern formulation owes most to Karl Popper's 20th-century philosophy of science, which identified falsifiability as the criterion distinguishing scientific from non-scientific claims. A claim is scientific if it can in principle be proven wrong. "All swans are white" is a scientific claim β€” one black swan falsifies it. "Everything happens for a reason" is not a scientific claim β€” no observation could falsify it.

Applied outside pure science, the Scientific Method provides a template for any situation where you want to reliably distinguish true beliefs from false ones: product development (A/B testing, usability testing), business strategy (market experiments, pilot programmes), medicine (clinical trials), and personal decision-making (small experiments before large commitments).

The scientific method doesn't require a laboratory. It requires: a specific, falsifiable prediction; a method of testing that prediction that minimises confounding variables; honest reporting of results including results that disconfirm your hypothesis; and willingness to update beliefs based on evidence.


How It Works​

Step 1: Observe and define the phenomenon
β€” What specific pattern or problem are you investigating?
β€” Be precise: "conversion rate is 2.3% for mobile users" not "mobile converts poorly"

Step 2: Research existing knowledge
β€” What is already known? What explanations exist?
β€” Avoid reinventing the wheel or contradicting established findings

Step 3: Form a specific, falsifiable hypothesis
β€” "If X is true, then we should observe Y under condition Z"
β€” Must be specific enough to be disproved

Step 4: Design the experiment
β€” What evidence would falsify the hypothesis?
β€” Control confounding variables
β€” Pre-register your success criteria (before seeing results)

Step 5: Collect data
β€” Execute the experiment rigorously
β€” Measure what you said you'd measure

Step 6: Analyse results
β€” Do results support or contradict the hypothesis?
β€” Apply appropriate statistical analysis

Step 7: Update beliefs and report
β€” Revise hypothesis if results disconfirm
β€” Report methodology and results honestly, including negative results

Three Real-World Examples​

A/B Testing in Product Development​

Duolingo hypothesised that streaks (consecutive days of learning) would increase retention. Scientific method applied:

  • Hypothesis: Users who see a streak-maintenance prompt on day 7 will have 15% higher 30-day retention than users who don't
  • Experiment: Randomly assign 50,000 users to streak prompt vs. no prompt
  • Pre-registered success criterion: Lift β‰₯15% at 95% statistical confidence
  • Result: 18% retention lift; hypothesis confirmed
  • Update: Deploy streak prompts to all users and explore further streak-related mechanics

Without the scientific method (just intuiting "streaks seem good"), Duolingo would have no reliable way to distinguish what actually works from what feels intuitively good. A/B testing has allowed Duolingo to run thousands of experiments per year, with each result building a reliable, evidence-based product strategy.

Clinical Trial: COVID-19 Vaccines​

Pfizer-BioNTech's mRNA vaccine development used the scientific method at scale:

  • Hypothesis: Two doses of BNT162b2 will reduce symptomatic COVID-19 infection by β‰₯50% in adults
  • Experiment: 43,661 participants randomised to vaccine or placebo; double-blind (neither participant nor researcher knew assignment)
  • Primary endpoint: Symptomatic COVID-19 confirmed by PCR, β‰₯7 days after second dose
  • Result: 95% efficacy (170 confirmed cases: 162 in placebo group, 8 in vaccine group)
  • Update: Regulatory emergency authorisation, deployment to billions

The double-blind randomised controlled trial is the gold standard precisely because it maximally controls for confounding variables (placebo effect, selection bias, measurement bias).

Business Hypothesis Testing​

A startup hypothesised that adding a free trial would increase paid conversion from landing page.

  • Hypothesis: A "Start Free Trial" CTA will produce higher 90-day paid conversion than "Buy Now" CTA
  • Experiment: 50/50 split test on new visitors, tracked for 90 days
  • Confounders controlled: Same ad spend, same landing page content, only CTA changed
  • Result: Free trial CTA showed 40% higher signup but 25% lower paid conversion β€” net effect negative
  • Update: Free trial attracts tire-kickers; refine qualifying mechanisms before retest

The result contradicted the hypothesis. Without a controlled experiment, the team might have attributed higher signups to the free trial and declared it a success, missing the monetisation degradation.


When to Use It​

βœ… Scientific Method disciplines thinking for:

  • Any experiment where you're testing a causal claim
  • Product decisions based on A/B testing
  • Business strategy validation through pilot programmes
  • Personal beliefs that should be updated based on evidence, not defended regardless of evidence

❌ Less directly applicable for:

  • One-time decisions with no ability to repeat the experiment
  • Ethical questions (science informs but doesn't resolve normative questions)
  • Time-critical decisions where experimental cycles are too slow
Pairs well withWhy
FalsificationFalsification is the philosophical core of the scientific method
Bayesian ThinkingBayesian updating formalises how to revise beliefs from experimental results
Minimum Viable TestMVT applies scientific method to business experiments efficiently
5 Whys5 Whys generates hypotheses that the scientific method then tests

Common Misuses and Limitations​

HARKing β€” Hypothesising After Results are Known. Forming the hypothesis after seeing the data, then presenting it as if it were pre-specified. This produces seemingly significant results that are actually pattern-matching to noise. Pre-registration (publishing your hypothesis and methods before seeing results) prevents this.

Underpowered studies. Running an experiment too small to detect real effects. An A/B test with 200 users can't reliably detect a 5% lift; you need thousands. Underpowered studies frequently produce false negatives (fail to detect real effects) and, paradoxically, inflated false positives.

Confusing correlation with causation. Scientific experiments are designed to isolate causal relationships. Observational studies (without random assignment) can show correlation but not causation. This confusion produces enormous amounts of bad health advice, business mythology, and policy mistakes.

Ignoring negative results. Publication bias (preferentially publishing positive results) corrupts the scientific record. In business contexts, the equivalent is "cherry-picking" experiments that confirmed your prior beliefs. Honest reporting of negative results is essential to learning.


ModelRelationship
FalsificationThe philosophical criterion that makes hypothesis testing meaningful
Bayesian ThinkingThe formal framework for updating beliefs from experimental evidence
Minimum Viable TestApplies scientific method efficiently in resource-constrained contexts
Black Box ThinkingThe cultural mindset that creates appetite for scientific experimentation

Frequently Asked Questions​

What makes a hypothesis "scientific" vs "unscientific"?

Following Karl Popper: a hypothesis is scientific if it is in principle falsifiable β€” if there exists some possible observation that would prove it wrong. "Exercise improves cardiovascular health" is scientific (we can design studies that could disprove it). "Everything happens for a reason" is unfalsifiable β€” no observation could disprove it, because any counter-evidence can be reinterpreted as part of the "reason." Unfalsifiable claims aren't necessarily false; they're just not scientifically testable.

How do you apply the scientific method without a formal lab or statistics team?

The key disciplines are achievable without formal infrastructure: (1) state your hypothesis before collecting data; (2) design a test that would actually prove your hypothesis wrong; (3) collect and report results honestly, including failures; (4) update your belief based on what you found. In business contexts, even simple before/after comparisons, controlled pilots in a subset of markets, or basic A/B tests with sufficient sample sizes provide scientific discipline without requiring a statistics PhD.

Why is replication important and what is the "replication crisis"?

Science progresses by independent replication β€” other researchers running the same experiment and checking if they get the same results. A finding that can't be replicated may have been a fluke or the result of bias. The "replication crisis" (2011–present) refers to the discovery that a surprisingly large fraction of published findings in psychology, medicine, and social science fail to replicate. Contributing factors include small sample sizes, HARKing, publication bias, and p-hacking. The crisis has prompted reforms including pre-registration, open data, and larger replication studies.


Further Reading​

  • Popper, K. (1959). The Logic of Scientific Discovery β€” the philosophical foundation of falsifiability
  • Kuhn, T. (1962). The Structure of Scientific Revolutions β€” how scientific paradigms shift
  • Kohavi, R. et al. (2020). Trustworthy Online Controlled Experiments β€” A/B testing at scale

Apply with AI​

πŸš€ Design an experiment for your hypothesis with MindMax β†’


This page is part of the MindMax Mental Models Knowledge Base.