If you've spent any time reading about depression screening, you may have run into two nearly identical tools with nearly identical names: the PHQ-9 and the PHQ-8. The difference between them is exactly one question โ but it happens to be the most sensitive question on the entire scale. The PHQ-8 is the PHQ-9 with Item 9, the question about thoughts of self-harm, removed.
That single deletion raises reasonable questions. Why would anyone remove the one item that asks about safety? Does taking it out change the score? Which version should you actually take? This article walks through the answers โ where the PHQ-8 came from, what the research says about how the two versions compare, and why this site uses the complete PHQ-9. (If you want to skip ahead, our free online PHQ-9 test uses all nine questions, scored and interpreted instantly.)
The One-Question Difference
The PHQ-9 asks how often nine problems have bothered you over the last two weeks โ low mood, loss of interest, sleep changes, fatigue, appetite changes, feelings of worthlessness, trouble concentrating, psychomotor changes, and finally Item 9: thoughts that you would be better off dead, or of hurting yourself in some way.
The PHQ-8 asks the first eight of those questions and stops. Everything else is unchanged: the same two-week window, the same four answer options worth 0 to 3 points each, the same instructions. The PHQ-8's maximum score is 24 instead of 27, simply because there's one fewer item to sum. Our line-by-line guide to the PHQ-9 questions covers what each of the nine items is actually asking; the PHQ-8 is that list minus the last entry.
Why Was Item 9 Removed?
The PHQ-8 wasn't created because Item 9 is a bad question. It was created because of a practical, ethical problem with where screening sometimes happens.
The PHQ-8 was described by Kurt Kroenke โ one of the original PHQ-9 authors โ and colleagues in a 2009 paper in the Journal of Affective Disorders, in the context of large population-based health surveys. Think of a telephone survey that dials tens of thousands of households to estimate how common depression is across a state, or a mailed questionnaire in an epidemiological study. In those settings, there is a real dilemma: if a survey asks about thoughts of self-harm and someone answers "nearly every day," there is often no clinician on the line, no established care relationship, and no reliable way to follow up with that person immediately.
Asking a safety question you cannot act on is an ethical problem. Rather than ask it and leave a potentially at-risk person hanging, researchers running anonymous or remote surveys chose to drop the item. The PHQ-8 exists so that population research can still measure depression responsibly when immediate clinical follow-up is impossible.
That origin story matters, because it defines what the PHQ-8 is for. It is a research and surveillance instrument. It was never meant to replace the PHQ-9 in settings where a human being โ a doctor, a nurse, a therapist, or a well-designed screening tool with crisis resources built in โ can respond to what Item 9 reveals.
Do the Scores Actually Differ?
Here's the part that surprises people: for the total score, barely.
In the 2009 Kroenke et al. paper, the PHQ-8 was evaluated as a measure of current depression in the general population, and its scores and operating characteristics track the PHQ-9 closely. The reason is straightforward. Item 9 is the least frequently endorsed item on the scale โ most people taking a depression screener answer "not at all" to it, contributing zero points. And among people who do endorse it, they almost always endorse several other items too, which means their total is already elevated with or without Item 9's contribution.
The practical consequences:
- The same cut-point works for both. A score of 10 or greater signals possible major depression on the PHQ-8 just as it does on the PHQ-9. You don't need to adjust the threshold when one item is removed.
- The severity bands transfer. The familiar bands โ minimal, mild, moderate, moderately severe, severe โ apply to both instruments with essentially the same boundaries. (Our full scoring guide walks through all five bands on the 0โ27 scale.)
- Population estimates match. For the purpose the PHQ-8 was built for โ estimating how common depression is in a large group โ the eight-item version does the job about as well as the nine-item version.
So if the totals are nearly interchangeable, why does the missing question matter at all?
Why Item 9 Still Matters: The Score Isn't the Point
Because Item 9 was never really about the total score.
On the PHQ-9, Item 9 does double duty. It contributes 0โ3 points to the sum like every other question โ but it also functions as a standalone safety signal that overrides the total. The scale's authors recommend that any answer above "not at all" on Item 9 deserves follow-up on its own, regardless of what the overall score says. Someone could answer "several days" to Item 9 and "not at all" to everything else, total a 1, and land in the "minimal" band โ and that person still deserves a follow-up conversation far more urgently than the band label suggests.
The PHQ-8 cannot capture that person at all. By design, it doesn't ask.
That's an acceptable trade-off in an anonymous telephone survey where nobody could respond anyway. It's a meaningful loss in almost every other context:
- In a clinic, Item 9 routinely surfaces thoughts a patient hadn't planned to mention. A checkbox can feel easier than saying the words out loud, and clinicians describe the item as an opening for a conversation that might not otherwise happen.
- In ongoing treatment, tracking Item 9 across repeat administrations tells a clinician something the total score can't โ whether the most serious symptom is appearing, persisting, or fading.
- In self-screening online, a well-built PHQ-9 tool can do what the survey researcher can't: respond immediately. When Item 9 is elevated, crisis resources can appear on the spot, not buried at the bottom of a results page.
If you're reading this and thoughts of being better off dead or of hurting yourself are part of your experience right now โ whatever any score says โ please reach out. In the US, call or text 988 (Suicide & Crisis Lifeline, free and available 24/7), or text HOME to 741741 for the Crisis Text Line. If you're in immediate danger, call 911. These thoughts are a symptom that responds to treatment, and they're exactly the thing screening exists to catch early.
Which Version Will You Encounter Where?
A quick orientation to where each instrument shows up in the wild:
| Setting | Typical version | Why |
|---|---|---|
| Doctor's office / primary care | PHQ-9 | Clinical follow-up is available; Item 9 is wanted |
| Therapy intake and progress tracking | PHQ-9 | Safety monitoring is part of care |
| Large phone or mail health surveys | PHQ-8 | No way to respond to an elevated Item 9 |
| Some workplace wellness surveys | PHQ-8 | Anonymous, no follow-up pathway |
| Research studies | Either | Depends on whether follow-up is possible |
There's also a much shorter sibling, the two-question PHQ-2, which is used as a quick first-pass filter rather than a severity measure โ we compare it to the full instrument in PHQ-9 vs PHQ-2.
Why This Site Uses the Full PHQ-9
Some online depression tests quietly drop Item 9. The site avoids a sensitive topic, the score still looks like a PHQ score, and most visitors never notice. We made the opposite choice, deliberately.
First, the validated instrument is the nine-item version. The severity bands, the cut-point of 10, and the accuracy evidence that makes the PHQ-9 trustworthy were all established for the complete scale. The PHQ-8 has its own validation for its own purpose โ population surveillance โ but a self-screener that someone takes because they're worried about themselves is much closer to the clinical use case than to a telephone survey.
Second, and more importantly, the ethical logic that justifies the PHQ-8 doesn't apply here. The PHQ-8 removes Item 9 because remote surveys can't respond to it. An online screening tool can. Our quiz includes Item 9, and any non-zero answer surfaces crisis resources immediately โ within the experience, before the score, exactly as the instrument's authors recommend. Removing the question wouldn't protect anyone; it would just mean the one symptom most worth catching goes unasked.
The Bottom Line
The PHQ-9 and PHQ-8 are the same instrument with one difference: the PHQ-8 omits Item 9, the question about thoughts of self-harm. The omission exists for a legitimate reason โ large research surveys can't follow up on a concerning answer โ and for total-score purposes the two versions behave almost identically, with the cut-point of 10 working for both. But Item 9 was never mainly about the total. It's a safety question, and in any setting where someone can act on the answer, it belongs on the page.
Neither version diagnoses depression โ both are screeners, and only a qualified professional can make a diagnosis. But if you're going to screen yourself, use the complete tool. Take the free, full nine-question PHQ-9 โ it takes about two minutes, it's scored instantly, and it asks everything the clinical version asks. And if Item 9 is the question you've been carrying: 988 is there around the clock.