Core Function IV: Assessment

Criterion 11: Select assessment tools

Choose the instruments that fit this client and the decision they are being used to make.

Working draft
This chapter is part of an unfinished manual.Worksheet Under construction

This criterion opens Core Function IV, Assessment. Screening decided who was eligible, intake documented the decision, and orientation prepared the client. Assessment is the structured evaluation of the person the program has admitted, and it begins with a prior question: which instruments are the right ones to use. Criterion 11 is about the selection of assessment tools, the reasoning that matches an instrument to a client, a purpose, and a context. The instruments named here are stated with their psychometric properties, and the two lethal medical hazards that some assessments screen for are cross-referenced to the safety spine of Criteria 1 and 2 rather than re-taught.

Learning Objectives

By the end of this module, the trainee will be able to:

  • Explain the purpose of structured assessment, and why the selection of tools is a distinct competency that precedes their administration.
  • Distinguish the domains an assessment battery should cover: psychological, medical and physiological, social and environmental, and cultural and spiritual.
  • State the psychometric properties of the core standardized psychological instruments, including item count, score range, and validated cutoffs, and cite each to its primary source.
  • Distinguish a screening instrument from a diagnostic one, and explain why a self-report score provides a provisional signal rather than a diagnosis.
  • Match an instrument to a client and context, accounting for literacy, language, culture, and the specific population.
  • Identify which assessments carry safety-critical stakes, connecting the relevant medical screening to the serotonin syndrome and cardiac hazards of Criteria 1 and 2.
  • Recognize the limits of an instrument and of the facilitator's scope, and route findings that exceed either to the appropriate professional.
  • Assemble a defensible assessment battery for a given program and explain the rationale for each tool selected.

Key Terms

Assessment. The structured procedures by which a program evaluates a client's psychological, medical, social, and other relevant conditions, in order to understand strengths, risks, and needs and to inform care and integration planning.

Assessment tool or instrument. A standardized or structured method for gathering information, ranging from a validated self-report questionnaire to a structured interview to a medical screening. Each tool has a defined purpose, a defined output, and defined limits.

Standardized instrument. A questionnaire or measure with established, published psychometric properties, administered and scored the same way each time, which allows a score to be interpreted against validated norms and cutoffs.

Psychometric properties. The measured performance characteristics of an instrument, including reliability (consistency) and validity (whether it measures what it claims). Sensitivity and specificity describe how well a cutoff detects and rules out a condition.

Screening versus diagnosis. Screening flags the likelihood that a condition is present and warrants further evaluation; diagnosis is a determination made by a qualified professional, typically via structured clinical interview. A self-report screening score is a provisional signal, not a diagnosis.

Sensitivity and specificity. Sensitivity is the proportion of true cases a cutoff correctly identifies; specificity is the proportion of true non-cases it correctly clears. A cutoff trades one against the other, which is why the chosen cutoff matters.

Cutoff score. The threshold on an instrument at or above which further action is indicated. Cutoffs are validated empirically and differ by instrument and purpose; using the wrong cutoff misclassifies clients.

Battery. The set of assessment tools a program selects to cover the domains relevant to its clients and its medicine, assembled so that the tools together give a complete enough picture without redundant burden.

Corroborative information. Information obtained, with the client's authorized release, from secondary sources such as a treating clinician, to confirm and supplement self-report. This is the subject of Criterion 12 and is noted here as part of a complete assessment picture.

Core Teaching

Why the selection of tools is its own competency

The criterion that opens Assessment names a specific task, identifies appropriate assessment tools, and makes a deliberate choice of that task over administering the assessments. Choosing the right instrument for a given client, purpose, and context is a competency that precedes and governs the act of administering it, and choosing badly cannot be repaired by careful administration. A tool that does not fit the client's language or literacy produces unreliable data no matter how well it is administered. A tool aimed at the wrong construct answers a question no one asked. A tool used beyond its validated purpose, such as treating a screening questionnaire as a diagnosis, produces false confidence. The value of structured assessment is that it moves a program past subjective impression toward reliable, comparable information, and that value is only realized when the instruments are chosen well. This criterion is therefore about judgment applied before data collection begins.

Structured assessment serves three functions that unstructured impression cannot. It provides measurable baselines against which change can later be tracked. It surfaces risks a client may not volunteer, including the medication and medical factors that carry lethal stakes in this field. And it gives a multidisciplinary team a shared, common language, so that a facilitator, a prescriber, and an integration provider work from the same data rather than from separate and possibly conflicting impressions. Each of these depends on selecting instruments whose outputs are reliable and interpretable, which is the work of this criterion.

The domains an assessment battery should cover

A complete assessment picture spans several domains, and a program assembles its battery by choosing at least one appropriate tool for each domain relevant to its clients and its medicine. Four domains recur. The psychological domain covers mood, anxiety, trauma, and risk, and is where the standardized self-report instruments are most developed. The medical and physiological domain covers the cardiovascular, hepatic, and pharmacological factors that determine physical safety, and it is where the safety-critical screening lives. The social and environmental domain covers the support, stability, and circumstances into which a client will integrate the experience, since integration happens in a context rather than a vacuum. The cultural and spiritual domain covers the frameworks and expectations a client brings, which shape both the experience and what will count as a good outcome for them. A battery that covers only the psychological domain, because that is where the convenient questionnaires are, leaves the medical, social, and cultural picture unexamined, and each of those gaps has produced real harm in practice.

The standardized psychological instruments, and their properties

Several standardized self-report instruments are well validated, freely available, and widely used, and a facilitator should know their actual properties rather than a vague sense of them. The properties below are stated from the primary validation sources, which are cited in full in the references.

The Patient Health Questionnaire-9, the PHQ-9, is a nine-item self-report measure of depression severity. Each item is scored from 0 (not at all) to 3 (nearly every day), giving a total range of 0 to 27, and the nine items correspond to the DSM depression criteria on which the instrument was built. Its validation study reported that a cutoff of 10 or greater yielded sensitivity and specificity of about 88 percent each, and the conventional severity bands are 0 to 4 minimal, 5 to 9 mild, 10 to 14 moderate, 15 to 19 moderately severe, and 20 to 27 severe (Kroenke, Spitzer, & Williams, 2001). A reduction of about 5 points is generally taken to mark clinically meaningful improvement, which makes the instrument useful for tracking change and not only for a single reading.

The Generalized Anxiety Disorder-7, the GAD-7, is a seven-item self-report measure of anxiety, each item scored 0 to 3 for a total range of 0 to 21. Its validation reported that a cutoff of 10 or greater yielded sensitivity of about 89 percent and specificity of about 82 percent for generalized anxiety disorder, with severity bands commonly given as 0 to 4 minimal, 5 to 9 mild, 10 to 14 moderate, and 15 to 21 severe (Spitzer, Kroenke, Williams, & Löwe, 2006). Like the PHQ-9, it functions both as a screen and as a severity tracker.

The PTSD Checklist for DSM-5, the PCL-5, is a twenty-item self-report measure of post-traumatic stress symptoms, each item scored 0, not at all, to 4, extremely, for a total range of 0 to 80, with the twenty items mapping to the twenty DSM-5 PTSD symptoms across four clusters. The measure itself was developed by Weathers and colleagues at the National Center for PTSD in 2013, and its peer-reviewed development and initial psychometric evaluation was published by Blevins and colleagues, who reported strong internal consistency and test-retest reliability (Blevins, Weathers, Davis, Witte, & Domino, 2015). A crucial point of scope: the PCL-5 yields a provisional determination and a severity score, and it does not itself diagnose PTSD; the reference standard for diagnosis is a structured clinical interview such as the Clinician-Administered PTSD Scale, conducted by a qualified professional.

Screening is not diagnosis, and a score is not a person

The single most important interpretive discipline in this criterion is distinguishing between screening and diagnosis. The self-report instruments above are screening and severity tools. A PHQ-9 score of 18 indicates that a client is reporting symptoms in the moderately severe range and warrants further evaluation; it does not, by itself, diagnose major depressive disorder, which is a clinical determination made by a qualified professional through interview and judgment. Treating a self-report score as a diagnosis is an error in both directions: it can pathologize a transiently distressed person and can miss a person who underreports. The instruments are valuable precisely because they are structured signals, and they are safe to use only when the facilitator understands that a signal is what they are. This connects directly to scope of practice: a facilitator selecting and administering these tools is gathering structured information, and when a score indicates a clinical question beyond their competence, the finding is routed to a qualified professional, exactly as the coexisting-conditions review of Criterion 2 requires. An instrument does not license a facilitator to practice beyond their scope; it structures the information they hand to those who can.

One risk flag warrants specific mention because it recurs across these instruments. Item 9 of the PHQ-9 asks about thoughts of being better off dead or of self-harm, and any endorsement of it is a signal that requires immediate attention and an appropriate safety response, regardless of the total score. A facilitator using the PHQ-9 must know this in advance and have a plan, since a suicidality signal is not something to note and move past. The handling of that signal is a clinical safety matter, and the program must have an established response and referral pathway involving appropriately qualified professionals and local crisis services.

Medical and physiological screening, and the safety spine

The medical domain is where assessment selection meets the safety spine of the workbook, and it is where a poorly chosen or omitted tool can be fatal rather than merely inaccurate. Appropriate medical assessment includes a structured medical history that captures allergies, chronic conditions, and, most importantly, a complete and exact list of all medications and supplements. It also includes physiological measures appropriate to the medicine, for example, blood pressure given the cardiovascular load of these compounds, and, for specific substances, targeted screening: cardiac evaluation and attention to the QT interval where ibogaine is involved, and hepatic function where relevant. The reason these are safety-critical rather than routine is taught in depth in Criteria 1 and 2 and is only cross-referenced here: an undisclosed serotonergic medication combined with an MAO-inhibiting brew such as ayahuasca can precipitate serotonin syndrome, and a QT-prolonging drug combined with ibogaine can precipitate a fatal arrhythmia. The instrument that captures the medication list is therefore not a psychological nicety; it is the data-collection step on which two lethal hazards turn. A facilitator selecting medical assessment tools must know which screening each medicine requires and, crucially, must know the boundary of their own competence, referring the interpretation of medical findings to a licensed medical professional rather than adjudicating cardiac or pharmacological risk themselves.

Social, environmental, cultural, and spiritual assessment

The remaining domains are assessed with more flexible tools, often structured interviews rather than scored questionnaires, and they matter because they determine whether a client can metabolize and integrate the experience. Social and environmental assessment explores support, housing, and vocational stability, as well as the relationships into which a client will return, since a supportive context is part of what makes integration possible and its absence is a risk that may need to be addressed before or alongside the work. Cultural and spiritual assessment explores the frameworks a client brings, including the beliefs, practices, and expectations that shape both the experience and what a good outcome means to them. A structured spiritual history or an open-ended interview can surface these dimensions, and attending to them is a matter of both respect and accuracy, since a purely clinical battery administered to a person whose framework is spiritual will miss much of what is relevant to their care. The biopsychosocial assessment and structured clinical interview traditions supply the tools for these domains, and the workbook's own assessment and interviewing sources model them.

Matching the tool to the client and the context

Selecting appropriate tools means tailoring them to a specific client and context, not adopting a generic battery and applying it uniformly. Several matching considerations govern the choice. Literacy and language determine whether a self-report instrument will produce valid data at all; an instrument in a language the client reads poorly measures reading, not mood. Cultural fit determines whether items are interpretable and non-alienating. The specific population matters: a combat veteran may be better served by a trauma assessment developed and validated in military populations than by a generic one, and an adolescent, an older adult, or a person from a particular cultural background may each warrant instruments validated in their group. Purpose matters: an instrument chosen to establish a baseline for tracking change is selected differently from one used as a one-time screen. The discipline is to ask, for each tool, whether it is valid for this client and this purpose, rather than whether it is familiar. A well-chosen smaller battery outperforms a large one administered without regard to fit.

Assembling and defending the battery

The output of this criterion is a defensible assessment battery: a set of tools, one or more per relevant domain, each chosen for a stated reason, that together give the program the picture it needs without imposing a redundant burden on the client. Assembling it well means covering the psychological, medical, social, and cultural domains that apply, selecting instruments whose properties fit the client and purpose, giving the safety-critical medical screening the weight its stakes warrant, and being able to state why each tool was included. A facilitator who can explain the rationale for every instrument in their battery and who knows the limits of each and the boundary of their own scope has met this criterion. The battery then feeds the criteria that follow in Core Function IV: obtaining corroborative information under release (Criterion 12), gathering the history with these tools (Criterion 13), explaining their rationale to the client (Criterion 14), and integrating the results into an evaluation (Criterion 15).

Clinical and Decision Tools

Tool 1. Standardized psychological instrument reference matrix

Properties of the core self-report instruments, each cited to its primary source in the references. All three are freely available and in the public domain or free for clinical use. Scores are provisional signals, not diagnoses.

Instrument

Measures

Items / range

Validated cutoff

Source

PHQ-9

Depression severity

9 items; 0-27

≥ 10 (sens ~88%, spec ~88%)

Kroenke et al., 2001

GAD-7

Anxiety severity

7 items; 0-21

≥ 10 (sens ~89%, spec ~82%)

Spitzer et al., 2006

PCL-5

PTSD symptoms (DSM-5)

20 items; 0-80

Provisional only; CAPS-5 for dx

Blevins et al., 2015

Tool 2. Assessment domains and tool types

Cover each domain relevant to the program. The shaded row is safety-critical: its tools screen for the lethal hazards of Criteria 1 and 2.

Domain

What it evaluates

Typical tools

Psychological

Mood, anxiety, trauma, risk

PHQ-9, GAD-7, PCL-5; structured interview

Medical / physiological

Cardiovascular, hepatic, pharmacological safety

Medical history; medication list; BP; substance-specific screening (C1, C2)

Social / environmental

Support, stability, integration context

Structured psychosocial interview; biopsychosocial assessment

Cultural / spiritual

Frameworks, beliefs, expectations

Spiritual history; open-ended interview

Tool 3. Screening versus diagnosis

Hold this distinction whenever a score is interpreted. A self-report score is a structured signal that routes to a professional, not a determination the facilitator makes.

Screening / severity tool

Diagnosis

What it does

Flags likelihood and severity; tracks change

Determines a clinical condition

Who

Can be administered by trained facilitator

Made by a qualified professional

Output

A provisional signal (e.g. PHQ-9 = 18)

A clinical determination via interview and judgment

On a positive result

Route to a professional for evaluation

Informs a treatment or eligibility decision

Tool 4. Tool-selection checklist

For each instrument considered, confirm fit before adopting it. Shaded items are the checks whose failure most compromises safety or validity.

Selection check

Confirmed?

The tool covers a domain relevant to this client and program

Yes / No

The tool is validated for this client's population, language, and literacy

Yes / No

The tool is used within its validated purpose (screen vs. diagnose)

Yes / No

Safety-critical medical screening for the specific medicine is included (C1, C2)

Yes / No

A plan exists for risk flags (e.g. PHQ-9 item 9 suicidality)

Yes / No

Findings beyond the facilitator's scope have a referral route

Yes / No

The battery avoids redundant burden while covering the needed domains

Yes / No

A rationale can be stated for including each tool

Yes / No

Worked Example: Assembling and Defending a Battery

The following models the selection reasoning this criterion should produce, for a specific client and program. Details are fictional. The point is that each tool is chosen for a stated reason and used within its validated purpose.

Client and program: A 46-year-old military veteran seeking psilocybin work for treatment-resistant depression with a trauma history, at a clinically supervised program with prescriber support.

Psychological tools, with rationale: PHQ-9 to establish a depression-severity baseline and track change, chosen because the presenting concern is depression and the instrument doubles as a progress measure. PCL-5 given the trauma history, chosen because it maps the DSM-5 PTSD symptoms and yields a severity baseline; a trauma measure with military validation is preferred given the population. GAD-7 to capture co-occurring anxiety. All three understood as screening and severity tools, not diagnoses.

Risk handling: Because the PHQ-9 includes item 9, a suicidality-response plan is in place before administration, and any endorsement triggers the program's safety protocol regardless of total score.

Medical tools, safety-critical: A structured medical history with a complete, exact medication and supplement list, and blood pressure, chosen because cardiovascular load and medication interactions are the frontline safety questions (C1, C2). Interpretation of any medication or cardiac finding is routed to the prescriber, not adjudicated by the facilitator.

Social and cultural tools: A structured psychosocial interview to assess support and stability for integration, and an open-ended conversation about the client's frameworks and what a good outcome would mean to them. Chosen because integration happens in context and the client's meaning-frame shapes the work.

Defense of the battery: Each domain relevant to this client is covered; each instrument is validated for the purpose it serves; the safety-critical medical screening is included and its interpretation referred out; and the battery is complete without redundant burden. The facilitator can state why each tool is present.

Case Vignettes

Work each vignette by identifying the selection or interpretation error, the principle at stake, and the correct handling. Fillable response sheets are in the companion worksheet PDF.

Vignette A

A facilitator administers the PHQ-9 to a client who scores 19 and, on that basis alone, tells the client they have moderately severe major depression and adjusts the plan accordingly, without any clinical interview or referral.

Guided questions: What interpretive error has the facilitator made? What does a PHQ-9 of 19 actually indicate, and what does it not? What was the correct use of this score, and who should make a diagnostic determination?

Vignette B

A program uses a single generic depression questionnaire as its entire assessment battery. A client with a significant cardiac history and several medications is admitted, and the medication interaction is never surfaced because no medical assessment tool was selected.

Guided questions: Which domains were omitted from this battery, and why is the medical omission the most dangerous? How does this connect to the hazards of Criteria 1 and 2? What would a complete battery have included?

Vignette C

A facilitator administers a set of English-language self-report questionnaires to a client whose first language is not English and who reads English with difficulty. The scores come back oddly inconsistent, and the facilitator treats them as valid data anyway.

Guided questions: What selection consideration was ignored, and what is the instrument actually measuring in this case? Why are the inconsistent scores unsurprising? How should the tools have been matched to this client?

Vignette D

During assessment, a client endorses item 9 of the PHQ-9, indicating thoughts of being better off dead. The total score is only 8, in the mild range, so the facilitator files the questionnaire without acting on the item.

Guided questions: Why does the item-9 endorsement matter independently of the total score? What should have happened the moment that item was endorsed? What does this reveal about needing a risk-response plan before administering the instrument?

Vignette E

A facilitator selects the PCL-5 and, seeing a high score, records in the file that the client has PTSD and proceeds as though a diagnosis has been established, without any structured clinical interview.

Guided questions: What is the scope error in treating the PCL-5 score as a diagnosis? What is the reference standard for a PTSD diagnosis? How should the high PCL-5 score have been used and communicated instead?

Role-Play and Practice Scripts

Practice in pairs, then switch. The aim is to select and introduce assessment tools in a way that gathers valid data, respects the client, and stays honest about what a score means.

Introducing a questionnaire honestly

“I am going to ask you to fill out a few short questionnaires. These are structured tools that help me understand where you are starting from, and later they help us see what has changed. They are not tests you can pass or fail, and a score on one of them is a signal we look at together, not a label. Answer as honestly as you can.”

Responding to a risk flag

“I noticed on this questionnaire you marked that you have had some thoughts of being better off dead. I take that seriously and I am glad you were honest about it. I want to pause the paperwork and talk with you about that directly right now.” Practice stopping immediately and shifting to a direct safety conversation, regardless of the total score.

Explaining screening versus diagnosis

“This questionnaire suggests your symptoms are in a higher range, and I want to be precise about what that means. It tells us something worth taking seriously, and it does not by itself diagnose anything. A diagnosis is made by a qualified clinician through a full conversation. What this does is tell us where to look and whom to involve.”

Matching a tool to the client

“Before I hand you a standard form, I want to make sure it actually fits you. Is English the language you are most comfortable reading? And is there anything about your background I should know so I choose tools that make sense for you rather than ones that would feel foreign or miss the point?” Practice selecting for fit rather than defaulting to a familiar form.

Self-Assessment and Reflection

Knowledge check

  1. Explain why the selection of assessment tools is a distinct competency that precedes their administration.
  2. Name the four assessment domains and give an appropriate tool type for each.
  3. State the item count, score range, and validated cutoff for the PHQ-9 and the GAD-7, and cite each to its source.
  4. Explain why the PCL-5 score is provisional and what the reference standard for a PTSD diagnosis is.
  5. Explain the screening-versus-diagnosis distinction and why treating a score as a diagnosis is an error in both directions.
  6. Explain why the medication-list tool is safety-critical, connecting it to Criteria 1 and 2.
  7. Give three considerations that govern matching a tool to a specific client.

Reflection

  1. Review your program's current assessment battery. Which domains does it cover, and which does it omit? Is the medical domain given the weight its stakes warrant?
  2. Do you have a risk-response plan in place before administering any instrument that includes a suicidality item? If not, that is a gap to close before the next assessment.
  3. Where might you be tempted to read a self-report score as more than it is? What would keep you disciplined about the line between a signal and a diagnosis?

Summary

Criterion 11 opens Assessment with a competency that precedes administration: identifying the appropriate assessment tools for a given client, purpose, and context. Structured assessment moves a program beyond subjective impression toward reliable, comparable information, provides baselines for tracking change, surfaces risks a client may not volunteer, and gives a multidisciplinary team a shared language, but only when the instruments are chosen well. A complete battery covers the psychological, medical and physiological, social and environmental, and cultural and spiritual domains. The core standardized psychological instruments have properties a facilitator should know precisely: the PHQ-9, nine items scored 0 to 27 with a validated cutoff of 10 (Kroenke et al., 2001); the GAD-7, seven items scored 0 to 21 with a cutoff of 10 (Spitzer et al., 2006); and the PCL-5, twenty items scored 0 to 80 mapping the DSM-5 PTSD symptoms, which yields a provisional signal rather than a diagnosis (Blevins et al., 2015). The governing interpretive discipline is that screening is not diagnosis: a self-report score is a structured signal that routes to a qualified professional, and a risk flag such as PHQ-9 item 9 requires an immediate safety response regardless of total score. The medical domain is where selection meets the safety spine, since the medication list and substance-specific screening tools guard against the serotonin syndrome and cardiac hazards of Criteria 1 and 2, and their interpretation is referred to a medical professional. Tools are matched to the client by literacy, language, culture, population, and purpose, and the output is a defensible battery in which every instrument has a stated rationale. The battery feeds the criteria that follow: corroboration, history-gathering, explaining rationale, and integrated evaluation.

References

Blevins, C. A., Weathers, F. W., Davis, M. T., Witte, T. K., & Domino, J. L. (2015). The Posttraumatic Stress Disorder Checklist for DSM-5 (PCL-5): Development and initial psychometric evaluation. Journal of Traumatic Stress, 28(6), 489–498. https://doi.org/10.1002/jts.22059

Kroenke, K., Spitzer, R. L., & Williams, J. B. W. (2001). The PHQ-9: Validity of a brief depression severity measure. Journal of General Internal Medicine, 16(9), 606–613. https://doi.org/10.1046/j.1525-1497.2001.016009606.x

Spitzer, R. L., Kroenke, K., Williams, J. B. W., & Löwe, B. (2006). A brief measure for assessing generalized anxiety disorder: The GAD-7. Archives of Internal Medicine, 166(10), 1092–1097. https://doi.org/10.1001/archinte.166.10.1092

Weathers, F. W., Litz, B. T., Keane, T. M., Palmieri, P. A., Marx, B. P., & Schnurr, P. P. (2013). The PTSD Checklist for DSM-5 (PCL-5). National Center for PTSD. https://www.ptsd.va.gov/professional/assessment/adult-sr/ptsd-checklist.asp

First, M. B. (2015). Structured Clinical Interview for the DSM (SCID). In R. L. Cautin & S. O. Lilienfeld (Eds.), The Encyclopedia of Clinical Psychology (pp. 1–6). Wiley. https://doi.org/10.1002/9781118625392.wbecp351

Sheehan, D. V., Lecrubier, Y., Sheehan, K. H., Amorim, P., Janavs, J., Weiller, E., Hergueta, T., Baker, R., & Dunbar, G. C. (1998). The Mini-International Neuropsychiatric Interview (M.I.N.I.): The development and validation of a structured diagnostic psychiatric interview for DSM-IV and ICD-10. Journal of Clinical Psychiatry, 59(Suppl 20), 22–33. https://www.psychiatrist.com/jcp/mini-international-neuropsychiatric-interview-mini/

Help develop this chapter

What would you bring to Criterion 11?

Relevant research, practice experience, and thoughtful review can help strengthen this material.

Request to contribute