A number does not explain its own authority
Statements such as “.70 is acceptable,” “.80 is large,” or “above .90 is excellent” sound precise. But precision in the number is not the same as strength of evidence. Before using a cutoff, ask:
- Who proposed or implemented it?
- What did the original source actually say?
- In which discipline, design, and use case?
- Was the value a definition, a convention, a default, or an informal heuristic?
- Do credible sources disagree?
- Is there evidence that no universal threshold exists?
The answer determines how the threshold should be presented and cited.
The ResultAtlas provenance labels
| Label | What it means | How to write it |
|---|---|---|
| FORMAL | A mathematical definition or model-implied boundary | State the definition and its assumptions. |
| ORIGINAL CONVENTION | A named source proposed an interpretive value or band | Attribute the author, year, and context. |
| DISCIPLINE CONVENTION | A field or use case adopted a practice | Name the discipline and intended decision. |
| SOFTWARE DEFAULT | A program chooses a value or behavior unless changed | Name the software, version, and option. |
| RULE OF THUMB | A heuristic aids screening or communication | Call it a heuristic, not a law. |
| DISPUTED | The citation, interpretation, or generality is contested | Show the source trail and disagreement. |
| NO UNIVERSAL THRESHOLD | Context prevents one defensible cross-setting cutoff | Explain what should be evaluated instead. |
These labels describe provenance, not a ranking in which every formal value is important and every heuristic is useless. A software default can be practical. A convention can support communication. The point is to prevent category confusion.
Case 1: Cronbach’s alpha at .70
The claim that α = .70 is a universal acceptable minimum is often linked to Nunnally. A source audit by Lance and colleagues found that this retelling overstates and decontextualizes the cited guidance.1 The underlying discussion varied expectations with the purpose of measurement rather than establishing one eternal gate.
ResultAtlas therefore labels .70 as DISPUTED, shows the source trail, and asks what the score will be used for. Cronbach’s alpha interpretation includes the full case and the distinction between internal consistency, dimensionality, and validity.
Case 2: Cohen’s small, medium, and large effects
Cohen proposed conventional effect-size values for settings where stronger substantive knowledge was unavailable, while explicitly describing conventional definitions as arbitrary.2 The values are useful reference points only when their authorship and scope stay attached.
ResultAtlas labels them ORIGINAL CONVENTION. A report can say “medium by Cohen’s 1988 convention,” then show the estimate, interval, original units, and domain consequences. Cohen’s d interpretation demonstrates this approach.
Case 3: competing ICC scales
Koo and Li label ICC values below .50 poor, .50–.75 moderate, .75–.90 good, and above .90 excellent.3 Cicchetti uses different boundaries in a psychological-assessment guideline, including ≥ .75 for excellent.4
Both are citable. They do not assign identical labels to the same number. ResultAtlas presents them side by side, marks the contextual difference, and refuses to manufacture a source-free compromise. Kappa and ICC interpretation shows how the coefficient form and interval matter alongside the selected scale.
Case 4: R-squared has no universal “good” value
R² is formally defined for a specified model and comparison, but a “good” value depends on the outcome, field, design, and purpose. The honest threshold entry is therefore NO UNIVERSAL THRESHOLD. What is a good R-squared value? replaces a floating cutoff with checks of model definition, residuals, baseline performance, and intended use.
This is not indecision. It is a positive conclusion about what the statistic can support.
How a threshold becomes a methodological urban legend
A common path looks like this:
- An author gives context-dependent guidance.
- A later source shortens it to a memorable number.
- The shortened claim is cited to the earlier author.
- Teaching materials repeat the number without its conditions.
- Software, peer review, or publication incentives turn it into a gate.
Lance and colleagues documented this kind of source inflation across four common cutoff criteria.1 Preventing it requires checking the original locator and preserving the language category in structured source data.
How ResultAtlas evaluates a threshold claim
Every threshold-ready article should record:
- the exact value or band;
- the label attached to it;
- the source and page, table, section, or software locator;
- the discipline and intended use;
- the provenance category;
- credible competing scales;
- whether the source supports a universal claim;
- the independent provenance-QA status.
The prose draft does not certify its own evidence. A separate reviewer must verify the source, locator, claim strength, and category before publication. If that gate remains pending, the article remains a draft even when its writing and technical checks are complete.
A practical reading rule
When a result is placed beside a cutoff, use this sentence structure:
The observed value is X. Source Y (year) labels that range Z for context C. This is a provenance category, not a universal fact. The decision also depends on relevant uncertainty and domain evidence.
For assumption checks, assumptions and sensitivity analysis applies the same principle: a diagnostic value is evidence to interpret, not a button that declares a model valid.