In 2001, a research team led by Drs. Kurt Kroenke, Robert Spitzer, and Janet Williams published a validation study in the Journal of General Internal Medicine that would quietly reshape how depression gets measured. Drawing on roughly 6,000 patients across primary care and obstetrics-gynecology clinics, they tested a nine-question self-report scale against independent, structured interviews conducted by mental health professionals.
That scale was the PHQ-9. Twenty-plus years later, it's the default depression measure in primary care, research, and electronic health records across much of the world — and the question people keep asking about it is the right one: can nine questions answered in two minutes really be trusted?
The evidence-based answer: for screening — the job it was designed for — the PHQ-9 is among the most thoroughly validated instruments in mental health. But "accurate" has a precise meaning for a screening tool, and knowing that meaning is what separates using your score wisely from over-reading it. If you'd like your own number to refer to as you read, our free PHQ-9 test takes about two minutes.
Built From the Diagnostic Criteria Themselves
One design decision explains much of the PHQ-9's staying power: its nine items aren't a loose collection of "things depressed people feel." They're a direct translation of the nine DSM criteria for a major depressive episode into self-report form — the same symptoms a psychiatrist would ask about, over the same two-week window. (The 9 PHQ-9 questions explained walks through each item.)
That one-to-one mapping means the scale isn't measuring something loosely adjacent to clinical depression. It's measuring the criteria themselves, from the patient's side of the desk.
The Two Numbers That Define a Screener
Every screening tool lives or dies by two statistics:
- Sensitivity — how well it catches people who do have the condition. High sensitivity means few missed cases.
- Specificity — how well it clears people who don't. High specificity means few false alarms.
At the standard cut-point of 10 or higher, the original validation research found the PHQ-9 achieves:
| Measure | At cut-point ≥10 | Plain meaning |
|---|---|---|
| Sensitivity | ~88% | Catches roughly 9 in 10 true cases |
| Specificity | ~88% | Correctly clears roughly 9 in 10 non-cases |
For a two-minute questionnaire, those numbers are genuinely strong — and they've survived two decades of independent scrutiny. A 2012 meta-analysis by Manea, Gilbody, and McMillan in CMAJ pooled studies across settings and confirmed that cut-points around 10 perform well for detecting major depression. In 2019, Levis, Benedetti, and Thombs went further in BMJ with an individual-participant-data meta-analysis — recombining raw data from thousands of participants across dozens of studies — and again landed on sensitivity around 88% with specificity in the mid-80s at the ≥10 cutoff.
Why 10, specifically? Because it balances the two kinds of error. Drop the threshold and you catch more cases but trigger more false alarms; raise it and false alarms fall but more people slip through. Ten is where the data keeps pointing.
Consistency: The Other Half of Accuracy
A test that gave different answers every time you took it would be useless no matter how well-designed. The PHQ-9 holds up here too.
Its internal consistency — whether the nine items measure one coherent thing rather than pulling in different directions — is high, with a Cronbach's alpha in the high 0.80s in the original validation work. Its test-retest reliability is also strong: scores stay stable across short, unchanged periods, which is precisely what makes the PHQ-9 useful for tracking treatment. A score that falls from 16 to 8 over two months is far more likely to reflect real improvement than statistical noise. (How PHQ-9 scoring works covers the bands your score moves through.)
What 88% Accurate Means You Should — and Shouldn't — Conclude
This is the section that matters most, because both misreadings of a PHQ-9 score come from forgetting that a screen is not a diagnosis.
Scoring high does not mean you're diagnosed. With ~88% specificity, roughly 12% of people above the cutoff won't actually meet criteria for major depressive disorder. A high score is a strong, evidence-based signal to look further — not a conclusion. Diagnosis weighs your history, duration, functional impact, and whether something else explains the symptoms better.
Scoring low does not mean you're fine. Sensitivity of 88% still misses about 1 in 10 true cases. People who minimize symptoms, or who've normalized a chronic low-grade heaviness through high-functioning depression, routinely score below what their experience warrants. When the number says "fine" and your life says otherwise, trust the life data. (What is a normal PHQ-9 score digs into why a low number isn't always the full story.)
Four Honest Limitations
Self-report is only as good as self-awareness. The PHQ-9 measures how you perceive and report the last two weeks. Denial, minimization, and simple inattention all move the number.
Two weeks is a snapshot, not a baseline. One unusually brutal — or unusually good — fortnight can pull a single score away from your true trend. Repeated administrations tell you more than any single result.
It can't distinguish depression from its look-alikes. Hypothyroidism, anemia, medication side effects, grief, burnout, and bipolar depression can all elevate a PHQ-9 score. If chronic work stress is a plausible driver, it's worth understanding how burnout differs from depression, because the two call for different responses. Bipolar disorder is the case that matters most: its depressive episodes look identical on this scale, yet the treatment differs significantly. That's exactly why a clinician, not a questionnaire, makes the call.
Depression rarely travels alone. High PHQ-9 scores often come with anxiety symptoms the scale doesn't measure at all. Depression and anxiety comorbidity explains why screening for both is usually worth the extra two minutes.
Where It Sits Among Depression Measures
The Hamilton Depression Rating Scale (HAM-D) is a respected classic — but clinician-administered, which rules out self-screening. The Beck Depression Inventory (BDI-II) is self-report and well validated — but longer, and a copyrighted commercial instrument. The PHQ-9's edge is fit-for-purpose design: two minutes, mapped directly onto DSM criteria, validated against clinical interviews, free to use, and sensitive to change over time.
One more thing distinguishes it from most casual online depression quizzes: it keeps Item 9, the question about thoughts of self-harm. A quiz that skips the most clinically important question is easier to publish — and less honest about what depression can involve.
The Verdict
Trust it — with the right frame. The PHQ-9 screens for major depression with roughly 88% sensitivity and 88% specificity at the standard cutoff, confirmed repeatedly by independent meta-analyses across two decades. Used as intended, it's about as good as a brief self-report instrument gets.
What it cannot do is replace a clinician. The most accurate use of your score is as the start of a conversation — a concrete number you bring to a doctor or therapist, not a verdict you reach alone.
And regardless of your total: if you're having thoughts of self-harm, or you answered anything other than "not at all" on the ninth question, please reach out now. In the US, the Suicide & Crisis Lifeline is available 24/7 by call or text at 988, and you can text HOME to 741741 for the Crisis Text Line.
Ready for your own evidence-based number? Take our free PHQ-9 test — two minutes, no signup, with a full interpretation of what your score means.