LazyTools

🔒 Every tool runs in your browser, the files and values you enter are never uploaded to any server. How it works

explainer

Which Statistical Test Should I Use? A Decision Guide

By Uttam Regmi · Published 2026-08-01 · Updated 2026-08-23 · 8 min read · Fact-checked, sources cited

A decision tree for choosing a statistical test based on data type, number of groups and normality

The right statistical test comes down to three questions: is your data numeric or categorical, how many groups are you comparing, and can you assume it’s roughly normal? Answer those three and the choice, t-test, z-test, ANOVA, Mann-Whitney or chi-square, is almost automatic. This guide walks the decision path, adds worked examples, and links a browser-based calculator for each, so nothing you enter leaves your device.

The three questions

  1. What type of data? Measurements/numbers (heights, times, revenue) vs categories/counts (yes/no, which-of-three, pass/fail).
  2. How many groups? One group against a target, two groups, or three-plus.
  3. Is it roughly normal? Symmetric and outlier-free → parametric tests are fine. Skewed, ordinal, tiny samples, or heavy outliers → reach for a non-parametric test.
Choosing a test What data? Categories / counts chi-square 2 proportions → z-test Numbers → how many groups? 2 groups normal → t-test σ known → z-test not normal → Mann-Whitney 3+ groups normal → ANOVA not normal → Kruskal-Wallis

Parametric or non-parametric? The normality question

The word parametric just means a test assumes your data follows a known distribution shape, usually the normal (bell) curve, and works from summary parameters like the mean and standard deviation. Parametric tests (t-test, z-test, ANOVA) are more powerful when that assumption holds: they can detect a real difference with a smaller sample. When the assumption is shaky, they can mislead.

Non-parametric tests (Mann-Whitney, Wilcoxon, Kruskal-Wallis) make no bell-curve assumption. They convert values to ranks and compare those, so a single wild outlier or a skewed tail can’t drag the result around. Reach for them when any of these is true:

  • Your sample is small (say, fewer than ~15-20 per group) and you can’t verify normality.
  • The data is visibly skewed or has heavy outliers.
  • The data is ordinal, ranks or ratings like a 1-5 satisfaction scale, where the gap between 4 and 5 isn’t guaranteed to equal the gap between 1 and 2.

A practical rule: if you’re unsure and the sample is small, the non-parametric test is the safer default. You lose a little power when the data really was normal, but you avoid a false conclusion when it wasn’t.

Categorical data → chi-square

If you’re counting things in categories, how many users chose A vs B vs C, or whether an outcome relates to a group. You want a chi-square test. Goodness-of-fit compares observed counts to expected proportions; the test of independence checks whether two categorical variables are related.

Worked example. You survey 200 visitors and record which of three plans they picked: 90 Basic, 70 Pro, 40 Team. Are those preferences different from an even 1/3-each split? Goodness-of-fit compares the observed counts (90, 70, 40) against the expected 66.7 each and returns a p-value. A small p-value says the uneven split is unlikely to be chance. If instead you had a two-way table, say plan choice (rows) by device type (columns), the test of independence tells you whether choice and device are linked.

The one special case: comparing exactly two proportions (an A/B conversion test) is cleaner with a two-proportion z-test.

Two numeric groups → t-test (or z-test, or Mann-Whitney)

Comparing two groups of measurements:

  • Roughly normal, σ unknown → the t-test. This is the default for most real data, because you almost never know the true population standard deviation.
  • σ known or very large sample → the z-test.
  • Not normal, skewed, ordinal, small, or outlier-heavy → the Mann-Whitney U test, which compares groups by rank and needs no normality assumption.

Worked example. Version A of a checkout takes, in seconds, 12, 14, 11, 13, 15 for five users; version B takes 9, 10, 8, 11, 10. Both samples look roughly symmetric with no wild outliers, and you don’t know the true spread of load times, so a two-sample t-test is the right call. It compares the two means (13.0 vs 9.6) relative to the variation within each group and reports whether that 3.4-second gap is bigger than you’d expect from noise. Had one user in A taken 60 seconds (a stuck request), the mean would lurch and normality would be doubtful, that’s when you’d switch to Mann-Whitney.

Paired data → paired t-test or Wilcoxon signed-rank

Before you settle on a two-group test, ask whether the two sets of numbers are independent or paired. Paired means each value in one group is tied to a specific value in the other, the same subjects measured twice (before vs after a change), or matched pairs. Blood pressure for 20 patients before and after a drug is paired; heights of 20 men and 20 unrelated women are not.

Pairing is worth exploiting because it cancels out person-to-person differences and gives you more power. For paired numeric data that’s roughly normal, use a paired t-test; if it’s non-normal or ordinal, use the Wilcoxon signed-rank test, the paired counterpart of Mann-Whitney. Using an unpaired test on paired data throws away that advantage and can hide a real effect.

Three or more groups → ANOVA (not many t-tests)

With 3+ groups, don’t run a t-test on every pair, each comparison carries its own false-positive risk, and doing many inflates the overall error rate. One-way ANOVA (calculator) tests all groups at once with a single F test. A significant result says the groups aren’t all equal; a follow-up post-hoc test (e.g. Tukey’s HSD) tells you which ones differ. If the groups are clearly non-normal, the rank-based Kruskal-Wallis test is the non-parametric counterpart.

Worked example. You test three landing-page headlines and measure time-on-page for visitors in each. That’s three numeric groups, so ANOVA is the tool, not three separate t-tests. If ANOVA returns a significant F, a Tukey post-hoc test then pins down whether headline C beat both A and B, or only A. If the time-on-page values were badly skewed (a common shape for durations), Kruskal-Wallis would replace ANOVA.

Quick reference

Your situationTestParametric?
Category counts, 1 variable vs expectedChi-square goodness-of-fit,
Two categorical variables related?Chi-square independence,
Two proportions (A/B)Two-proportion z-test,
Two independent numeric groups, normalt-testYes
Two independent numeric groups, σ knownz-testYes
Two independent numeric groups, non-normalMann-Whitney UNo
Paired numeric (before/after), normalPaired t-testYes
Paired numeric (before/after), non-normalWilcoxon signed-rankNo
3+ independent numeric groups, normalOne-way ANOVAYes
3+ independent numeric groups, non-normalKruskal-WallisNo

Common mistakes to avoid

A few traps catch people far more often than picking the “wrong” family of test:

  • Running many t-tests instead of ANOVA. Every extra pairwise test adds another chance of a false positive. Use ANOVA for 3+ groups, then a post-hoc test.
  • Ignoring pairing. Treating before/after measurements as independent groups discards the biggest source of power you have. Match the test to the design.
  • Assuming normality on tiny samples. With a handful of points you usually can’t confirm the bell-curve shape, so a rank-based test is the honest choice.
  • Reading only the p-value. Statistical significance is not the same as practical importance, which is the next section.

Don’t forget effect size

Effect size is the most common gap in reported results. A p-value answers “could this be chance?” but says nothing about magnitude. With a large enough sample, a trivially small difference becomes “significant.” Pair every test with an effect size, Cohen’s d for mean differences, or the appropriate measure for your test, so a reader can see whether a detectable difference is also a difference worth acting on.

The bottom line

Pick your test from data type → group count → normality: categories go to chi-square, two normal groups to a t-test, three-plus to ANOVA, and anything non-normal to a rank-based test like Mann-Whitney. Then read the p-value next to an effect size. Run any of them locally with the LazyTools statistics calculators, your data never leaves the browser.

Frequently asked questions

How do I choose the right statistical test?

Start with three questions: what kind of data do you have (numbers or categories), how many groups are you comparing, and can you assume the data is roughly normal? Categorical counts point to a chi-square test; comparing two normal groups points to a t-test; three or more groups to ANOVA; and non-normal or ordinal data to a non-parametric test like Mann-Whitney. Each of these has a LazyTools calculator that runs in your browser.

What's the difference between a t-test and ANOVA?

A t-test compares the means of two groups; ANOVA compares three or more at once. You could run many t-tests instead, but each carries its own false-positive risk and running lots of them inflates the overall error rate, ANOVA tests all groups together at one significance level, which is why it's the correct tool for 3+ groups.

When should I use a non-parametric test like Mann-Whitney?

When the assumption that your data is roughly normal is doubtful, small samples, skewed distributions, ordinal ratings (like 1-5 scales), or data with strong outliers. The Mann-Whitney U test compares two groups using ranks instead of raw values, so it doesn't need normality. Its 3-plus-group counterpart is the Kruskal-Wallis test.

What's the difference between a z-test and a t-test?

Both compare means, but a z-test assumes the population standard deviation is known (or the sample is large enough that it's effectively known) and uses the normal distribution, while a t-test estimates the standard deviation from the sample and uses the t-distribution to account for that extra uncertainty. In practice, use a t-test unless you truly know the population SD or have proportion data (where a z-test applies).

Which test do I use for categorical data like yes/no or counts?

A chi-square test. Use goodness-of-fit to compare observed counts against expected proportions, and the test of independence to check whether two categorical variables (rows and columns of a contingency table) are related. For two proportions specifically, like an A/B conversion test, a two-proportion z-test is the standard choice.

Are these calculators private?

Yes, every LazyTools statistics calculator runs entirely in your browser using JavaScript. Your data is never uploaded, and the tools work offline.