Assessment Reliability & Methodology

How the Enneagram.guide assessment performs statistically, what those numbers mean, and what they don't.

7 min read
Updated 8/22/2026

Last Updated: August 22, 2026 • Sample Size: n = 4,417 completed assessments • Question Set Analysed: Version 1.4


Summary

We analyse how the assessment performs and publish the results. This page reports what we measured, on how many people, and what it does and does not establish.

Completed assessments analysed4,417
Types meeting the standard reliability threshold (α ≥ 0.70)7 of 9
Reliability range across the nine types0.63 – 0.80
Questions contributing meaningfully to their type92%

Two of the nine type scales fell below the conventional 0.70 threshold in this analysis. Both are named below, and both were revised as a direct result.


Reliability by Type

Reliability here means internal consistency — whether the questions making up a single type's scale actually behave as though they measure the same thing. It is reported per type rather than as one figure for the whole assessment, because the nine types are intentionally distinct: a single instrument-wide number would blur them together and overstate how much any one type has been verified.

The conventional threshold in personality research is α ≥ 0.70.

TypeReliability (α)
1 — The Reformer0.76
2 — The Helper0.79
3 — The Achiever0.76
4 — The Individualist0.79
5 — The Investigator0.73
6 — The Loyalist0.80
7 — The Enthusiast0.63
8 — The Challenger0.69
9 — The Peacemaker0.78

Seven types sit in the 0.73–0.80 band. The Enthusiast (0.63) and The Challenger (0.69) did not meet the threshold.

Both were investigated in detail, and three questions across them were rewritten in the August 2026 revision specifically to address what the analysis found. Whether that worked is an open question until enough responses accumulate on the new version; the next analysis reports the answer.


Question Quality

Every question is checked against the rest of its own type's scale — that is, whether a person's answer to it lines up with how they answered the other questions measuring the same type. A question that doesn't track with its own scale isn't earning its place, regardless of how well-written it looks.

MeasureResult
Questions clearing the standard contribution threshold92%
Questions performing strongly75%
Questions requiring replacement0
Questions flagged for revision5 of 63

No question in the assessment fell into the range where replacement is indicated. Five were flagged as worth improving, and three of those were rewritten in August 2026.


What These Numbers Do Not Establish

Reliability and validity are different things, and it matters which one you have.

Reliability — the figures above — means the assessment measures something consistently. Validity means it measures what it claims to. The two are often conflated in assessment marketing. They shouldn't be.

We have solid reliability evidence. We do not yet have published evidence for:

Test-retest stability. Whether the same person taking the assessment months apart gets the same result. This is arguably the question users care most about, and we have not yet measured it. It is next on our list.

Comparison against other instruments. Whether our results agree with other established Enneagram measures or with expert-assessed type.

Predictive validity. Whether results forecast anything measurable outside the assessment itself.

Any assessment claiming a number "proves validity" is describing something a reliability statistic cannot do. We would rather say plainly what we have and haven't established.


How We Analyse

Sample. All complete responses to the analysed question set, excluding internal test accounts. Responses are separated by question-set version so that figures never mix data from different versions of the assessment.

Reliability. Cronbach's alpha, computed per type — the standard measure of internal consistency in personality research.

Question-level analysis. Each question is evaluated against the other questions measuring the same type, using the corrected form of the statistic — meaning a question is never scored against a total that includes itself, which would flatter it.

Aggregate only. The analysis runs on aggregate statistics computed inside our database. No individual response data is extracted.

Cadence. We re-analyse periodically, and whenever the question set changes. When a revision is made, the next analysis reports whether it actually helped.


Limitations

The sample is self-selected. People who seek out an Enneagram assessment are not a random sample of the population, and the distribution of types here should not be read as a population distribution.

Self-report has limits. Results reflect how you currently see yourself, which is influenced by self-awareness, mood, and how you interpret the wording. That is true of every self-report instrument.

Primarily English-speaking respondents, taking the assessment online.

Reported figures describe question set 1.4. The assessment was revised to version 1.5 on 22 August 2026, incorporating three rewritten questions. Figures for the current version will be published once enough responses have accumulated to make them meaningful; until then, we report the version the data actually came from rather than implying the numbers describe the current set.

No assessment is accurate for every individual. These statistics describe average performance across thousands of people. Your result is best treated as a well-informed starting point for reflection, not a verdict.


Common Questions

Why report reliability per type instead of one overall number?

A single figure across all 63 questions would mostly reflect a general tendency to agree or disagree with statements, rather than anything about the Enneagram. It also tends to look flattering for reasons that have nothing to do with quality. Per-type figures are the meaningful ones, so those are what we publish.

What happens when a type or question underperforms?

It gets flagged by the analysis, investigated, and revised. The August 2026 revision rewrote three questions across The Enthusiast and The Challenger scales for that reason. The following analysis measures whether the revision worked, and that figure is published here in turn.

This is the loop the page exists to document: measure, act on what the measurement shows, measure again.

Is this assessment scientifically validated?

It has been tested for reliability, on a large sample, with the results published above. It has not been through independent peer-reviewed validation studies, which is what that phrase usually implies. We would rather show our figures and let them speak.

How does this compare to other Enneagram assessments?

Most do not publish reliability figures or sample sizes at all, which makes direct comparison impossible. Among those that do, ours are in a comparable range. We think publishing the underlying data is more useful than a ranking we would be grading ourselves on.

What if my result doesn't feel right?

The Enneagram is about underlying motivation, not behaviour, and the type that fits is often not the one that feels most flattering. Read the descriptions for your top two or three results. If another fits better, it probably is better. The assessment is a starting point for that investigation, not a substitute for it.


References

  1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. American Educational Research Association.

  2. Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16(3), 297–334.

  3. Nunnally, J. C., & Bernstein, I. H. (1994). Psychometric theory (3rd ed.). McGraw-Hill.

  4. Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Lawrence Erlbaum Associates.