Where Statistical Thresholds Come From

A statistical threshold can be a formal boundary, an author’s convention, a discipline-specific practice, a software default, a rule of thumb, a disputed citation, or a value with no universal cutoff at all. ResultAtlas keeps those categories separate so a convenient label does not acquire more authority than its source supports.

A number does not explain its own authority

Statements such as “.70 is acceptable,” “.80 is large,” or “above .90 is excellent” sound precise. But precision in the number is not the same as strength of evidence. Before using a cutoff, ask:

  1. Who proposed or implemented it?
  2. What did the original source actually say?
  3. In which discipline, design, and use case?
  4. Was the value a definition, a convention, a default, or an informal heuristic?
  5. Do credible sources disagree?
  6. Is there evidence that no universal threshold exists?

The answer determines how the threshold should be presented and cited.

The ResultAtlas provenance labels

LabelWhat it meansHow to write it
FORMALA mathematical definition or model-implied boundaryState the definition and its assumptions.
ORIGINAL CONVENTIONA named source proposed an interpretive value or bandAttribute the author, year, and context.
DISCIPLINE CONVENTIONA field or use case adopted a practiceName the discipline and intended decision.
SOFTWARE DEFAULTA program chooses a value or behavior unless changedName the software, version, and option.
RULE OF THUMBA heuristic aids screening or communicationCall it a heuristic, not a law.
DISPUTEDThe citation, interpretation, or generality is contestedShow the source trail and disagreement.
NO UNIVERSAL THRESHOLDContext prevents one defensible cross-setting cutoffExplain what should be evaluated instead.

These labels describe provenance, not a ranking in which every formal value is important and every heuristic is useless. A software default can be practical. A convention can support communication. The point is to prevent category confusion.

Case 1: Cronbach’s alpha at .70

The claim that α = .70 is a universal acceptable minimum is often linked to Nunnally. A source audit by Lance and colleagues found that this retelling overstates and decontextualizes the cited guidance.[Lance (2006)] The underlying discussion varied expectations with the purpose of measurement rather than establishing one eternal gate.

ResultAtlas therefore labels .70 as DISPUTED, shows the source trail, and asks what the score will be used for. Cronbach’s alpha interpretation includes the full case and the distinction between internal consistency, dimensionality, and validity.

Case 2: Cohen’s small, medium, and large effects

Cohen proposed conventional effect-size values for settings where stronger substantive knowledge was unavailable, while explicitly describing conventional definitions as arbitrary.[Cohen (1988)] The values are useful reference points only when their authorship and scope stay attached.

ResultAtlas labels them ORIGINAL CONVENTION. A report can say “medium by Cohen’s 1988 convention,” then show the estimate, interval, original units, and domain consequences. Cohen’s d interpretation demonstrates this approach.

Case 3: competing ICC scales

Koo and Li label ICC values below .50 poor, .50–.75 moderate, .75–.90 good, and above .90 excellent.[Koo (2016)] Cicchetti uses different boundaries in a psychological-assessment guideline, including ≥ .75 for excellent.[Cicchetti (1994)]

Both are citable. They do not assign identical labels to the same number. ResultAtlas presents them side by side, marks the contextual difference, and refuses to manufacture a source-free compromise. Kappa and ICC interpretation shows how the coefficient form and interval matter alongside the selected scale.

Case 4: R-squared has no universal “good” value

R² is formally defined for a specified model and comparison, but a “good” value depends on the outcome, field, design, and purpose. The honest threshold entry is therefore NO UNIVERSAL THRESHOLD. What is a good R-squared value? replaces a floating cutoff with checks of model definition, residuals, baseline performance, and intended use.

This is not indecision. It is a positive conclusion about what the statistic can support.

How a threshold becomes a methodological urban legend

A common path looks like this:

  1. An author gives context-dependent guidance.
  2. A later source shortens it to a memorable number.
  3. The shortened claim is cited to the earlier author.
  4. Teaching materials repeat the number without its conditions.
  5. Software, peer review, or publication incentives turn it into a gate.

Lance and colleagues documented this kind of source inflation across four common cutoff criteria.[Lance (2006)] Preventing it requires checking the original locator and preserving the language category in structured source data.

How ResultAtlas evaluates a threshold claim

Every threshold-ready article should record:

  • the exact value or band;
  • the label attached to it;
  • the source and page, table, section, or software locator;
  • the discipline and intended use;
  • the provenance category;
  • credible competing scales;
  • whether the source supports a universal claim;
  • the independent provenance-QA status.

The prose draft does not certify its own evidence. A separate reviewer must verify the source, locator, claim strength, and category before publication. If that gate remains pending, the article remains a draft even when its writing and technical checks are complete.

A practical reading rule

When a result is placed beside a cutoff, use this sentence structure:

The observed value is X. Source Y (year) labels that range Z for context C. This is a provenance category, not a universal fact. The decision also depends on relevant uncertainty and domain evidence.

For assumption checks, assumptions and sensitivity analysis applies the same principle: a diagnostic value is evidence to interpret, not a button that declares a model valid.

Sources and notes

  1. Charles E. Lance, Marcus M. Butts, and Lawrence C. Michels (2006). The Sources of Four Commonly Reported Cutoff Criteria: What Did They Really Say? Full article; four cutoff case studies
  2. Jacob Cohen (1988). Statistical Power Analysis for the Behavioral Sciences, Second Edition §2.2 pp. 25–27; caveats §1.4 pp. 12–13 and §11.2 p. 532
  3. Terry K. Koo and Mae Y. Li (2016). A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research Interpretation guideline
  4. Domenic V. Cicchetti (1994). Guidelines, Criteria, and Rules of Thumb for Evaluating Normed and Standardized Assessment Instruments in Psychology Reliability guideline table

Optional analytics help us understand site use. They remain off unless you accept.

Your preferences