Delirium tools: making sense of the landscape
Commentary. These are the author’s interpretations of the cited evidence. Disclosures.
A practical guide to the different purposes, the evidence and routine clinical use.
There are so many delirium tools that it can be difficult to work out what they all mean. A clinician looking for a practical assessment can encounter dozens of names, acronyms, adaptations and scoring systems. A researcher faces much the same problem. Which tools do similar things? Which have been studied beyond their first publication? And which have actually been used successfully in everyday care?
There are more than 60 published instruments and named variants within the scope of this site. The number depends partly on what we count. A new instrument, a shorter version, a translation and a severity score derived from an existing tool are different developments. Putting them all in one list can obscure the smaller number of underlying approaches.
Many have limited follow-up evidence. In the current directory, 34 of the 82 core profiles have one identified primary diagnostic-accuracy publication and eight have two. These are provisional counts of distinct publications, not proof that no further papers or clinical use exist. Some tools remain part of the history of the field; others have been revised, as with the DRS and DRS-R-98, or have become more prominent in particular settings. The picture continues to evolve.
This commentary builds on the classification and distinction between episodic assessment and monitoring that I described on Delirium Words in 2020 (MacLullich, 2020a; 2020b). The aim is to help clinicians and researchers find their way through the field. For pragmatic reasons, I have cited selected papers rather than the original publication for every instrument. Original references and further evidence can be found in the individual profiles on DeliriumTools.
How we acquired so many tools
The history helps explain the variety. The Delirium Rating Scale, published in 1988, provided a structured way to describe symptoms. The Confusion Assessment Method, or CAM, followed in 1990, with an algorithm intended to help non-psychiatric clinicians identify delirium (Trzepacz, Baker and Greenhouse, 1988; Inouye et al., 1990).
Subsequent development followed several clinical and research needs. More detailed severity measurement led to instruments such as the Memorial Delirium Assessment Scale, or MDAS, and the revised Delirium Rating Scale, DRS-R-98. Intensive care required approaches suitable for patients with limited verbal communication, including people receiving mechanical ventilation. CAM-ICU and the Intensive Care Delirium Screening Checklist, ICDSC, were published in 2001 (Breitbart et al., 1997; Trzepacz et al., 2001; Ely et al., 2001; Bergeron et al., 2001).
Nursing observation provided another route. The Delirium Observation Screening Scale, DOSS, and Nursing Delirium Screening Scale, Nu-DESC, organise observations made during care. Brief bedside assessments also developed further, including the 4AT and 3D-CAM (Schuurmans, Shortridge-Baggett and Duursma, 2003; Gaudreau et al., 2005; Bellelli et al., 2014; Marcantonio et al., 2014).
There are now tools for children, informant questionnaires, measures of arousal and attention, severity scales, and instruments addressing distress and the experience of delirium. Research has also produced methods for identifying delirium retrospectively from records and predicting who may develop it. This variety reflects the range of work that needs to be done. It also makes the term “delirium tool” rather imprecise.
Begin with the purpose
For detection in clinical care, I find it useful to distinguish two processes.
Episodic assessment means assessing for delirium at a particular point: when a patient presents with acute illness, at a relevant transition in care, during a period of high risk, or when delirium is suspected. “Episodic” needs that explanation. The practical question is whether this person may have delirium now.
Monitoring for new delirium means repeated observation for changes during continuing care. A patient who has no delirium on admission may develop it later. Staff need a way to recognise that change and arrange assessment. An observation tool or a brief question can form part of this process.
These are patterns of use, and they can overlap. An assessment tool can be repeated within a monitoring programme. An observational measure can contribute to an assessment at a particular time. A family member’s account of a recent change may be useful in either process. The tool’s instructions, the clinical setting and the local pathway determine how it should be used.
I do not think we have a settled solution for monitoring across hospital settings. The evidence is mixed, and there is no single agreed approach covering which tool to use, how often to use it, and whether a positive observation should trigger another assessment or constitutes the service’s only structured assessment. General recommendations to observe patients regularly leave much of this work to local services.
Shift-based scoring also creates a practical risk. If staff first record and act on changes at the end of a shift, delirium may have been present for hours. A further handover can add delay. That is a concern about how the process is organised, rather than an inevitable property of every observational tool. A new change should prompt attention when it is noticed; the scheduled score should not determine when staff begin to respond.
The way a tool is performed needs equal attention. In my CAM-Lite commentary, I use that term for an unofficial observation-based use of CAM features in which the preceding cognitive assessment is omitted. It is my descriptive label, not the name of a validated instrument from the CAM developers. An altered assessment cannot simply inherit the validation results of the original method (MacLullich, 2025). I would not use observation-only CAM ratings as a replacement for a structured assessment with cognitive testing. The concern is missed delirium despite the time spent recording scores. The separate commentary on CAM and monitoring explains the evidence, including what the large implementation studies can and cannot establish.
Following the course of established delirium is another task. Severity measurement can describe symptoms over time and support research into prognosis and treatment. It should have its own place in the classification. For example, CAM-S was developed to quantify severity, with validation addressing measurement properties and associations with outcomes (Inouye et al., 2014).
The distinction also affects how we read studies. Agreement with a delirium diagnosis is relevant to a detection tool. A severity scale needs evidence about what its scores measure, how consistently they are rated and whether they reflect change. A distress measure must represent the experience of the person answering it. Patient and family accounts deserve separate attention; the development of the Delirium Burden instruments illustrates this (Racine et al., 2019).
A practical classification
The table groups tools by the work they help people do. Some rows describe purpose; others identify a useful characteristic, information source or population. A tool can therefore appear in more than one row.
The three-paper column is a provisional shortlist. It uses the distinct primary diagnostic-accuracy publications identified in the website’s current working audit, checked on 4 September 2026. These counts have not all been independently adjudicated for study eligibility, overlapping cohorts or tool versions. They are not counts of independent external validations. The website does not yet have an equivalent, consistently populated count of validation studies for every other type of measurement.
| Group | What it helps with | Tools with three or more primary diagnostic-accuracy papers identified |
|---|---|---|
| Episodic bedside assessment in adults | Assessing for delirium at a particular encounter | 3D-CAM, 4AT, bCAM, CAM, DDT-Pro, DSI |
| Observation and monitoring for new delirium | Recognising changes during continuing care | DOSS, ICDSC, NEECHAM, Nu-DESC, RADAR, S-PTD |
| Delirium severity | Describing the degree and course of symptoms | DDS, DRS, DRS-R-98, MDAS. Inclusion here records identified diagnostic-accuracy studies, not a judgement of severity measurement quality. |
| Ultra-brief tools | A short initial assessment or prompt for further assessment | RADAR, SQiD, UB-2. These use different information and have different instructions. |
| Informant tools | Using observations from relatives, carers or others who know the patient | FAM-CAM, I-AGeD, Sour Seven, SQiD |
| Cognitive tests and specific domains | Measuring cognition or a component of delirium | DelApp |
| Adult intensive care assessment | Assessing delirium in critical care | CAM-ICU, ICDSC |
| Paediatric assessment | Assessing delirium in defined child age groups and settings | CAPD, pCAM-ICU, psCAM-ICU, SOS-PD |
| Arousal measures | Describing level of alertness, agitation or sedation | A dependable three-paper shortlist requires a separate review of RASS, mRASS and OSLA, with attention to versions and the outcome being validated. |
| Distress, experience and burden | Describing how delirium affects patients, families and staff | The three-paper threshold has not been established here. Examples to explore include DEQ, DEL-B and TDDS. |
| Motor features, retrospective identification, risk prediction and causes | Other clinical and research questions | Each requires its own evidence criteria. Explore these through the classification and risk-prediction pages. |
The Cognitive Test for Delirium (CTD) is another example of a cognitive battery. It has two identified primary DTA papers in this audit, so it does not meet the table’s three-paper threshold.
The table contains 27 distinct profiles meeting the current diagnostic-accuracy threshold. Its gaps are also informative. CAM-S belongs in a discussion of severity despite falling outside this particular count. Arousal measures have relevant published evaluations that require reconciliation with the directory. An empty cell must not be interpreted as an absence of validation evidence.
“Ultra-brief” describes length, while “informant” describes where information comes from. Neither is a separate clinical purpose. They are useful ways to narrow a search once the task is clear. The same applies to setting, language, training requirements and whether the patient needs to communicate verbally.
Put the tools along the patient’s course
An admission assessment is one point in a longer period of care. Consider a patient arriving in the emergency department, moving to a medical admissions unit, and then spending several days or weeks on a ward. Assessment may be needed at presentation and at later points, while observation for new changes continues.
The perioperative course needs its own view: before surgery, recovery or the post-anaesthesia care unit, and then the postoperative ward. The relevant population, effects of anaesthesia, communication and level of arousal can change between these stages. UK NICE guidance makes a specific distinction: it recommends CAM-ICU or ICDSC in critical care or the recovery room after surgery, with the 4AT recommended for assessment when indicators of delirium are identified in other covered settings (NICE, 2023).
Our review of routine care described four reported patterns of assessment: admission only; admission followed by repeated inpatient assessment; repeated inpatient assessment; and postoperative assessment. Supplementary Figure 1 shows these patterns (Penfold et al., 2024). They describe what the included services did. The timing of an individual service’s assessments still needs a clinical rationale.
The patient journey diagram shows assessment points and continuing observation. A new concern prompts assessment and clinical review. That response matters as much as the choice of tool.
Which tools have become used in practice?
This question requires evidence about use. Publication counts alone cannot answer it.
A UK survey obtained responses from 154 of 169 NHS organisations. Among the 146 reporting formal delirium assessment processes, 80% reported using the 4AT, 45% CAM and 36% SQiD. Organisations could report more than one tool. These were reports of adoption, rather than observations of every assessment performed at the bedside (Tieges, Lowrey and MacLullich, 2021).
The evidence on recorded routine care is narrower than the full directory. Penfold and colleagues’ systematic review included 22 research studies and four audit reports involving at least 1,000 patients in general acute hospitals. Six tools appeared: CAM, 4AT, DOSS, bCAM, Nu-DESC and ICDSC. The search ended in December 2022, so this is a defined evidence snapshot rather than a complete current inventory of clinical use (Penfold et al., 2024).
There is also routine ICU evidence for CAM-ICU and for RASS as an arousal and sedation measure. SQiD has published improvement-project evidence, including a small project integrating it into nursing care rounds (Vasilevskis et al., 2011; McCleary and Cumming, 2015). These reports differ greatly in scale and in what they establish. MDAS has clinical validation and research applications, but those alone would not justify assigning it a badge for substantial routine-care implementation.
Validation and everyday performance
A validation study can establish how an instrument performs under the study’s conditions. Research staff may have extensive training, time to seek informant information and access to expert supervision. Eligibility criteria can exclude patients who are difficult to assess, and the final analysed sample may be more selected still. These features vary between studies and need to be examined in the methods.
Routine care can be quite a different context. The person completing an assessment may have little training in the tool or in delirium, while also caring for several other unwell patients. Information may be missing. Assessments may be delayed, shortened or completed retrospectively. The relevant question is how the specified tool performs in that setting, with the support actually available.
The variation can be considerable. In a study across ten ICUs, bedside CAM-ICU assessments had sensitivity of 47% and specificity of 98% against expert assessment. Another study found sensitivity and specificity of 81%, with sustained agreement between bedside and research nurses over three years. The populations and reference assessments differed, so these figures should not be treated as a controlled comparison. Together, they show why performance needs to be examined in the setting where the tool is used (van Eijk et al., 2011; Vasilevskis et al., 2011).
Across the studies in the Penfold review, overall completion rates ranged from 19% to 100%, with substantial variation in positive-score rates. Completion, positive scores and diagnostic accuracy answer different questions. A high completion rate says that assessments were recorded. A low positive-score rate can reflect case mix, timing, selective assessment or missed delirium. Estimating missed cases requires an appropriate reference assessment; the positive-score rate alone cannot do that (Penfold et al., 2024).
The 4AT also has implementation limitations to understand. In a two-centre study of 82,770 emergency admissions, completion was 77% in Lothian and 49% in Salford. Scores were associated with subsequent outcomes, but the study did not establish diagnostic accuracy against an independent reference assessment (Anand et al., 2022). For CAM, Rohatgi and colleagues reported a programme covering 105,455 encounters, with 98.8% screening completion after CAM was introduced within a sustained hospital programme. That is evidence of implementation at scale; it still leaves questions about the assessment process and detection performance (Rohatgi et al., 2019).
For a clinician, useful evidence therefore includes whether usual staff can complete the assessment appropriately and whether its findings lead to clinical review. For a researcher, it includes the assessment schedule, rater training, reference standard, missing assessments and the distinction between patients and repeated observations. These details determine how confidently results can be interpreted or compared across studies.
Making a choice
I would begin with the patient group, the clinical or research task, and the point in care. Then examine the relevant evidence, the practical requirements and the response to a concerning result. A short assessment may still require information from someone who knows the patient. A tool with substantial research evidence may require training or time that a service has not provided.
The directory’s two- and three-paper filters help narrow the field. They count identified primary diagnostic-accuracy publications, with the limitations described above. The routine-care evidence offers a separate view. It includes programmes reporting difficulties as well as successful delivery, and distinguishes large-scale implementation from small feasibility projects.
The next useful step for a service is to test its chosen approach in its own care pathway. Record who should be assessed, who actually is assessed, and what happens after a positive result. Review a sample of assessments for quality. Those observations can guide changes to the pathway and contribute evidence that other services can use.
Declaration of interests: I am the editor of DeliriumTools and have been involved in developing the 4AT, 4-DSD, OSLA, DelApp and Delbox. I am a co-author of several studies discussed here, including the Penfold review. The directory should apply the same standards of evidence and presentation to all tools.
The field is still developing. Better studies of routine delivery, repeated assessment, patient experience and action after a positive result should improve how we use these tools. The practical aim is earlier recognition and better care, with an assessment process that staff and patients can manage.
References
Anand, A., Cheng, M., Ibitoye, T., MacLullich, A.M.J. and Vardy, E.R.L.C. (2022) ‘Positive scores on the 4AT delirium assessment tool at hospital admission are linked to mortality, length of stay and home time: two-centre study of 82,770 emergency admissions’, Age and Ageing, 51(3), afac051. https://doi.org/10.1093/ageing/afac051.
Bellelli, G., Morandi, A., Davis, D.H. et al. (2014) ‘Validation of the 4AT, a new instrument for rapid delirium screening: a study in 234 hospitalised older people’, Age and Ageing, 43(4), pp. 496–502. https://doi.org/10.1093/ageing/afu021.
Bergeron, N., Dubois, M.J., Dumont, M., Dial, S. and Skrobik, Y. (2001) ‘Intensive Care Delirium Screening Checklist: evaluation of a new screening tool’, Intensive Care Medicine, 27(5), pp. 859–864. https://doi.org/10.1007/s001340100909.
Breitbart, W., Rosenfeld, B., Roth, A. et al. (1997) ‘The Memorial Delirium Assessment Scale’, Journal of Pain and Symptom Management, 13(3), pp. 128–137. https://doi.org/10.1016/S0885-3924(96)00316-8.
Ely, E.W., Inouye, S.K., Bernard, G.R. et al. (2001) ‘Delirium in mechanically ventilated patients: validity and reliability of the confusion assessment method for the intensive care unit (CAM-ICU)’, JAMA, 286(21), pp. 2703–2710. https://doi.org/10.1001/jama.286.21.2703.
Gaudreau, J.D., Gagnon, P., Harel, F., Tremblay, A. and Roy, M.A. (2005) ‘Fast, systematic, and continuous delirium assessment in hospitalized patients: the nursing delirium screening scale’, Journal of Pain and Symptom Management, 29(4), pp. 368–375. https://doi.org/10.1016/j.jpainsymman.2004.07.009.
Inouye, S.K., van Dyck, C.H., Alessi, C.A. et al. (1990) ‘Clarifying confusion: the confusion assessment method. A new method for detection of delirium’, Annals of Internal Medicine, 113(12), pp. 941–948. https://doi.org/10.7326/0003-4819-113-12-941.
Inouye, S.K., Kosar, C.M., Tommet, D. et al. (2014) ‘The CAM-S: development and validation of a new scoring system for delirium severity in 2 cohorts’, Annals of Internal Medicine, 160(8), pp. 526–533. https://doi.org/10.7326/M13-1927.
MacLullich, A. (2020a) ‘A classification of delirium assessment tools’, Delirium Words, 13 July. Read the article (Accessed: 4 September 2026).
MacLullich, A. (2020b) ‘Delirium detection in routine clinical care: two basic processes’, Delirium Words, 15 June. Read the article (Accessed: 4 September 2026).
MacLullich, A. (2025) ‘The CAM-Lite: why this unofficial delirium screening tool falls dangerously short’, Delirium Words, 26 April. Read the article (Accessed: 4 September 2026).
Marcantonio, E.R., Ngo, L.H., O’Connor, M. et al. (2014) ‘3D-CAM: derivation and validation of a 3-minute diagnostic interview for CAM-defined delirium: a cross-sectional diagnostic test study’, Annals of Internal Medicine, 161(8), pp. 554–561. https://doi.org/10.7326/M14-0865.
McCleary, E. and Cumming, P. (2015) ‘Improving early recognition of delirium using SQiD (Single Question to identify Delirium): a hospital based quality improvement project’, BMJ Quality Improvement Reports, 4(1), u206598.w2653. https://doi.org/10.1136/bmjquality.u206598.w2653.
National Institute for Health and Care Excellence (NICE) (2023) Delirium: prevention, diagnosis and management in hospital and long-term care. Clinical guideline CG103, recommendations 1.5–1.6, updated 18 January 2023. Read the recommendations (Accessed: 4 September 2026).
Penfold, R.S., Squires, C., Angus, A. et al. (2024) ‘Delirium detection tools show varying completion rates and positive score rates when used at scale in routine practice in general hospital settings: a systematic review’, Journal of the American Geriatrics Society, 72(5), pp. 1508–1524. https://doi.org/10.1111/jgs.18751.
Racine, A.M., D’Aquila, M., Schmitt, E.M. et al. (2019) ‘Delirium Burden in Patients and Family Caregivers: development and testing of new instruments’, The Gerontologist, 59(5), pp. e393–e402. https://doi.org/10.1093/geront/gny041.
Rohatgi, N., Weng, Y., Bentley, J. et al. (2019) ‘Initiative for prevention and early identification of delirium in medical-surgical units: lessons learned in the past five years’, The American Journal of Medicine, 132(12), pp. 1421–1430.e8. https://doi.org/10.1016/j.amjmed.2019.05.035.
Schuurmans, M.J., Shortridge-Baggett, L.M. and Duursma, S.A. (2003) ‘The Delirium Observation Screening Scale: a screening instrument for delirium’, Research and Theory for Nursing Practice, 17(1), pp. 31–50. https://doi.org/10.1891/rtnp.17.1.31.53169.
Tieges, Z., Lowrey, J. and MacLullich, A.M.J. (2021) ‘What delirium detection tools are used in routine clinical practice in the United Kingdom? Survey results from 91% of acute healthcare organisations’, European Geriatric Medicine, 12(6), pp. 1293–1298. https://doi.org/10.1007/s41999-021-00507-2.
Trzepacz, P.T., Baker, R.W. and Greenhouse, J. (1988) ‘A symptom rating scale for delirium’, Psychiatry Research, 23(1), pp. 89–97. https://doi.org/10.1016/0165-1781(88)90037-6.
Trzepacz, P.T., Mittal, D., Torres, R. et al. (2001) ‘Validation of the Delirium Rating Scale-revised-98: comparison with the delirium rating scale and the cognitive test for delirium’, The Journal of Neuropsychiatry and Clinical Neurosciences, 13(2), pp. 229–242. https://doi.org/10.1176/jnp.13.2.229.
van Eijk, M.M., van den Boogaard, M., van Marum, R.J. et al. (2011) ‘Routine use of the confusion assessment method for the intensive care unit: a multicenter study’, American Journal of Respiratory and Critical Care Medicine, 184(3), pp. 340–344. https://doi.org/10.1164/rccm.201101-0065OC.
Vasilevskis, E.E., Morandi, A., Boehm, L. et al. (2011) ‘Delirium and sedation recognition using validated instruments: reliability of bedside intensive care unit nursing assessments from 2007 to 2010’, Journal of the American Geriatrics Society, 59(Suppl. 2), pp. S249–S255. https://doi.org/10.1111/j.1532-5415.2011.03673.x.