Commentary

Tools with large-scale clinical implementation evidence

Commentary. These are the author’s interpretations of the evidence.

A validation study tells us what a tool can do under specified conditions. Routine-care studies show what happens when services try to use it.

Show these tools in the directory · Read the Penfold review summary and tables

4AT

The 4AT has been used in large routine-care datasets as well as research validation. Anand et al. (2022) analysed 82,770 older emergency admissions at two hospitals. Completion differed substantially between sites: 77% and 49%. Positive scores were associated with adverse outcomes. This supports its use as a practical source of clinical information, while also showing that introducing a tool does not ensure that everyone receives it.

Penfold et al. (2025) reported 75,221 admissions, with 4AT scores recorded in 62,188. That study examined possible dementia using recorded diagnoses as the reference. It adds implementation experience and evidence about cognitive impairment, but is not an independent diagnostic validation for delirium.

The 4AT also needs scrutiny when routine delivery is less reliable. In a cardiac surgical study, Chang et al. (2023) reported lower sensitivity in the nurse-delivered routine phase than in the research phase. The samples and delivery conditions differed. It would be wrong to assume that accuracy from a validation study transfers unchanged to every ward.

My practical conclusion is that the 4AT has substantial implementation evidence, but a service still needs to check completion, correct administration and the response to concerns. Its use should support clinical assessment. It does not replace it.

Additional source: Chang et al. (2023), The Journal of Thoracic and Cardiovascular Surgery, 165, pp. 1151–1160.e8.

Selected implementation reports

  • Anand et al. (2022). 82,770 emergency admissions. Two centres; 4AT completion 77% and 49%. Positive results were associated with outcomes. This was not a diagnostic-accuracy comparison.
  • Penfold et al. (2025). 75,221 emergency admissions. 62,188 had a 4AT recorded; admission scores were associated with recorded dementia. See the linked 2026 correction. Not a standalone dementia diagnostic test.

Open the profile and validation papers

bCAM

bCAM is a brief structured assessment with its own procedure. Its validation and implementation evidence should be read separately from that for the original CAM. Sharing an algorithm does not make different assessment methods interchangeable.

Dulin et al. (2022) reported implementation across 93,388 hospital admissions. The overall proportion screened was 19%, reaching 63% by the end of the programme. Among those screened, 15% were positive. Completion varied markedly across services.

This is useful evidence that bCAM can be introduced across clinical services. It also shows how much the result depends on delivery. The positive proportion applies to those assessed, not to all admissions, and a change in the population being assessed can change that proportion.

The report does not establish diagnostic sensitivity across the whole hospital. Nor does it show that all positive assessments led to effective care. Those are separate questions requiring other measurements.

For a service considering bCAM, I would look at whether staff can perform its attention and reasoning tasks correctly, how unable-to-assess results are handled, and whether completion is sustained. Local evaluation should include checks against an appropriate clinical assessment in a sample of patients. High completion is useful, but it is only one part of a successful delirium pathway.

Selected implementation reports

  • Dulin et al. (2022). 93,388 hospital admissions. 17,769 were screened in the staged programme (19% overall); completion reached 63% by its end. Associations with discharge destination do not prove an effect of screening.

Open the profile and validation papers

CAM

CAM has substantial experience of routine use. That experience cannot be treated as evidence that every local version works well. The original method requires an interview with cognitive testing, together with observation and history. A checkbox labelled CAM does not establish that this procedure took place.

Rohatgi et al. (2019) described a hospital programme covering 105,455 encounters. After screening was introduced, recorded completion was 98.8%, but only 2.4% of screened encounters were positive. An expert-assessed pilot found delirium in 17.3% of 278 patients. These were different samples. Their difference is concerning, but it is not a sensitivity calculation.

Corradi et al. (2016) analysed routine CAM records from 88,206 encounters. Positive and unable-to-assess records need careful interpretation. This was an analysis of clinical records, not a paired diagnostic-accuracy study.

These reports show why recording a score is insufficient as an implementation measure. Services need to establish what staff actually do and whether delirium is being detected. The finding of low sensitivity in studies of nurse recognition adds to the concern about observation alone. It does not invalidate correctly administered CAM or every named adaptation.

I would not use observation-only CAM as a substitute for the structured assessment it was designed to provide. The separate commentary sets out the evidence and its limits.

Selected implementation reports

  • Rohatgi et al. (2019). 105,455 encounters across the programme. After CAM screening was introduced, completion was 98.8% and 2.4% of screened encounters were positive. The 17.3% expert-assessed pilot was a separate sample. The study did not measure diagnostic sensitivity.
  • Corradi et al. (2016). 88,206 hospital encounters. Repeated routine CAM records; 7.9% were positive in the review extraction, with frequent unable-to-assess records. This was an EHR analysis, not paired diagnostic validation.

Open the profile and validation papers

CAM-ICU

CAM-ICU has a large validation literature and experience in routine intensive care. It was designed to allow assessment when speech is limited, including during mechanical ventilation. Its attention test and arousal requirements remain essential parts of the procedure.

Trogrlić et al. (2019) studied an ICU guideline programme involving 3,930 patients in six hospitals. The accompanying process evaluation describes the use of CAM-ICU and ICDSC. Screening increased from 35% to 96% across the programme. This is a combined programme denominator, not a separate sample size for each instrument.

The intervention also changed sedation, mobilisation and other care. Improvements in brain dysfunction cannot be attributed to CAM-ICU alone. Implementation studies of bundles answer a different question from studies of diagnostic accuracy.

Routine accuracy is variable. Van Eijk et al. (2011) reported sensitivity of 47% and specificity of 98% in their multicentre bedside study. Patients who could not be assessed because of coma were handled separately. This is a reminder that a widely validated tool still depends on correct delivery.

For a service using CAM-ICU, I would check the actual assessment procedure, staff competence, handling of coma and the clinical response. A negative result should not override a convincing concern. The relevant comparison is between well-specified procedures in the intended setting, not between tool names alone.

Additional source: van Eijk et al. (2011), American Journal of Respiratory and Critical Care Medicine, 184, pp. 340–344.

Selected implementation reports

Open the profile and validation papers

DOSS

The Delirium Observation Screening Scale, often called DOS or DOSS, brings together observations made during nursing care. Its routine-care literature is useful because it shows how an observation scale is used within a clinical service, rather than only by a research team.

Schubert et al. (2018) reported a hospital cohort of 29,278 eligible admissions. Screening was targeted according to the programme’s criteria, and the number assessed was smaller than the eligible population. These denominators should be kept separate. A programme’s total size is not the number of completed scales.

DOSS also appeared in the combined ward and intensive-care programme reported by Fuchs et al. (2020), with ICDSC used in intensive care. Results from that programme cannot be attributed to either instrument alone.

My main reservation concerns the whole pathway. Observations accumulated during a shift can be clinically valuable, but a positive score entered at handover may arrive several hours after the first change. Staff need a route to act on that change when it occurs.

A service should also specify which version and scoring rules it uses, how missing observations are handled, and who reviews a positive result. Large-scale use establishes practical experience. It does not by itself establish high sensitivity, timely treatment or better outcomes.

Selected implementation reports

  • Schubert et al. (2018). 29,278 eligible hospital patients. 10,906 had a DOSS assessment in the review extraction. Screening was recommended for a high-risk subset, so the denominator needs care.

Open the profile and validation papers

ICDSC

ICDSC combines eight features assessed in intensive care over an observation period. It is a different procedure from a brief bedside interview. Sedation, arousal and the availability of observations affect interpretation.

Fuchs et al. (2020) reported a programme involving 10,733 eligible older patients across hospital services. DOSS was used outside intensive care and ICDSC in intensive care. This is a large combined programme, not a count of 10,733 patients assessed with ICDSC alone.

The multicentre implementation work of Trogrlić et al. (2019, 2020) also included ICDSC and CAM-ICU within a broader ICU guideline programme. Screening increased alongside changes in sedation, mobilisation and other care. Those changes are important, but they prevent a simple conclusion that one screening instrument produced the outcome.

For local use, the key questions include whether staff recognise each feature, how the observation window is defined, and how a positive result changes care. A scheduled checklist should never require staff to wait before reporting a new clinical concern.

ICDSC has both validation and implementation evidence. The choice between it and another ICU assessment should consider the clinical workflow and staff competence, as well as accuracy estimates. NICE names ICDSC or CAM-ICU for critical care and post-anaesthetic recovery. That recommendation does not make every implementation equally effective.

Selected implementation reports

  • Fuchs et al. (2020). 10,733 eligible hospital patients. Combined programme: DOSS outside ICU and ICDSC in ICU. 10,261 patients were assessed. The denominator and results cannot be attributed to ICDSC alone.
  • Trogrlić et al. (2019), with process evaluation (2020). 3,930 patients across a combined six-ICU programme. CAM-ICU and ICDSC were used within the programme. This is not a separate denominator for either tool. Screening, sedation and other aspects of care changed together.

Open the profile and validation papers

NuDESC

Nu-DESC is an observational nursing scale. Its attraction is that it can draw on features seen during ordinary care. The practical question is whether staff notice the relevant changes, score them consistently and act promptly.

LaHue et al. (2021) evaluated a hospital pathway in 22,708 adults aged 50 or over across seven non-intensive-care units. Nu-DESC formed part of a broader programme that also addressed risk, prevention and management. This provides evidence of use at scale. It cannot isolate the effect of Nu-DESC from the rest of the programme.

Completion, the threshold used and the population assessed all affect the observed positive-score rate. Research estimates of sensitivity and specificity should not simply be applied to a ward’s routine results. A low score should not dismiss a new concern from the patient, family or staff.

A further issue is timing. If a scale summarises observations over a shift, staff may identify a change well before they enter the score. The care pathway should allow that concern to trigger assessment immediately. Waiting for a scheduled score adds delay without clinical benefit.

Nu-DESC is therefore relevant to services considering observational monitoring. Its implementation should be judged by recognition and response, as well as by the proportion of forms completed. The best monitoring approach remains an open question.

Selected implementation reports

  • LaHue et al. (2021). 22,708 hospitalised adults across study periods. Hospital-wide multicomponent pathway using nurse screening. Included in Penfold’s review of studies with at least 1,000 patients; Nu-DESC effects cannot be separated from the rest of the pathway.

Open the profile and validation papers

What should a service measure?

Measure who was eligible, who was assessed and when. Check whether the procedure was followed and compare a sample with an appropriate clinical assessment. Record what happened after a positive result or clinical concern. Completion, accuracy and clinical benefit answer different questions.

The field is still developing. Better studies of delivery, patient groups and clinical response will help us choose tools and pathways more confidently.

Review scope

The starting point is Penfold et al. (2024), a systematic review of general-hospital studies with at least 1,000 patients, searched to 31 December 2022. Selected later reports and ICU studies have been added. This is not a completed search of every setting.

Penfold, R.S. et al. (2024), Journal of the American Geriatrics Society, 72, pp. 1508–1524. Article and supplement.

References

Anand, A., Cheng, M., Ibitoye, T. et al. (2022). Positive scores on the 4AT delirium assessment tool at hospital admission are linked to mortality, length of stay and home time: two-centre study of 82,770 emergency admissions. Age and ageing, 51, afac051.

Chang, Y., Ragheb, S.M., Oravec, N. et al. (2023). Diagnostic accuracy of the "4 A's Test" delirium screening tool for the postoperative cardiac surgery ward. The Journal of thoracic and cardiovascular surgery, 165, 1151–1160.e8.

Corradi, J.P., Chhabra, J., Mather, J.F. et al. (2016). Analysis of multi-dimensional contemporaneous EHR data to refine delirium assessments. Computers in biology and medicine, 75, 267–274.

Dulin, J.D., Zhang, J., Marsden, J. et al. (2022). Association of delirium screening on hospitalized adults and postacute care utilization: A retrospective cohort study. The American journal of the medical sciences, 364, 554–564.

Fuchs, S., Bode, L., Ernst, J. et al. (2020). Delirium in elderly patients: Prospective prevalence across hospital services. General hospital psychiatry, 67, 19–25.

LaHue, S.C., Maselli, J., Rogers, S. et al. (2021). Outcomes Following Implementation of a Hospital-Wide, Multicomponent Delirium Care Pathway. Journal of hospital medicine, 16, 397–403.

Penfold, R.S., Bowman, E., Vardy, E.R.L.C. et al. (2025). Using scores from the 4AT delirium detection tool as an indicator of possible dementia: a study of 75 221 older adult hospital admissions. Age and ageing, 54, afaf144.

Penfold, R.S., Squires, C., Angus, A. et al. (2024). Delirium detection tools show varying completion rates and positive score rates when used at scale in routine practice in general hospital settings: A systematic review. Journal of the American Geriatrics Society, 72, 1508–1524.

Rohatgi, N., Weng, Y., Bentley, J. et al. (2019). Initiative for Prevention and Early Identification of Delirium in Medical-Surgical Units: Lessons Learned in the Past Five Years. The American journal of medicine, 132, 1421–1430.e8.

Schubert, M., Schürch, R., Boettger, S. et al. (2018). A hospital-wide evaluation of delirium prevalence and outcomes in acute care patients - a cohort study. BMC health services research, 18, 550.

Trogrlic, Z., van der Jagt, M., van Achterberg, T. et al. (2020). Prospective multicentre multifaceted before-after implementation study of ICU delirium guidelines: a process evaluation. BMJ open quality, 9, e000871.

Trogrlić, Z., van der Jagt, M., Lingsma, H. et al. (2019). Improved Guideline Adherence and Reduced Brain Dysfunction After a Multicenter Multifaceted Implementation of ICU Delirium Guidelines in 3,930 Patients. Critical care medicine, 47, 419–427.

van Eijk, M.M., van den Boogaard, M., van Marum, R.J. et al. (2011). Routine use of the confusion assessment method for the intensive care unit: a multicenter study. American journal of respiratory and critical care medicine, 184, 340–344.