Post 13 — Four Hours

England set a four-hour clock on emergency care and paid hospitals for meeting it. Patients were moved to short-stay units because the transfer stopped the clock, ambulances were held outside, and completed cases spiked in the last twenty minutes. Compliance reached 97.7 percent.

Share
Behavior Follows Rewards — Post 13: Four Hours

Every week I take one real case — a company, a government program, a league — and show how the reward shaped the behavior. Welcome to Behavior Follows Rewards.

The clock had become the patient.

In England, a government target required that ninety-eight percent of emergency patients be admitted, transferred or discharged within four hours. Financial rewards were attached to the number. Hospitals found ways to hit the target. The most reliable of them did nothing for the care.

In July 2000, the British government published the National Health Service (NHS) Plan. Among its commitments was a promise to end long emergency department waits — patients had been known to wait more than thirty-seven hours in English Accident and Emergency (A&E) departments. The Plan put a number on it: by 2004, no one should wait more than four hours in A&E from arrival to admission, transfer or discharge. No one. The target was set at one hundred percent. In 2004 it was cut to ninety-eight, to allow clinical exceptions — patients being actively resuscitated, patients who deteriorated unexpectedly.

Financial incentives were attached to make the commitment credible. Hospitals that met their targets received staged payments from the Department of Health. Management performance scores, star ratings, and hospital reputations were publicly tied to four-hour compliance. The target was real. The reward was real.

The behavior began almost immediately — and not on the four-hour target first.

In 2002, a Commission for Health Improvement review of one trust found that trolleys had been moved into hallways and recorded as beds — enough to clear the twelve-hour trolley-wait standard. The designation changed. The patient’s situation did not.

By 2004, the year the Plan had set as its deadline, the behaviors had multiplied. Trusts began moving patients to Short-Stay Units and Clinical Decision Units before the four-hour mark, because the transfer stopped the clock — a practice a later inquiry would document in detail. Ambulances were held outside departments so that patients had not technically arrived. When the numbers still did not cooperate, some trusts adjusted them directly.

The British Medical Association surveyed NHS emergency departments in this period, and the results are the fullest account we have of what the target was doing inside the departments. Sixteen percent of departments described direct manipulation of recorded waiting-time data. Eighty-two percent reported threats to patient safety as a direct result of pressure to meet the target — patients discharged before being fully assessed or stabilised, patients moved to inappropriate areas or wards, the care of the seriously ill or injured compromised. And of the departments that met the target, only twenty-six percent believed their published figures were an accurate representation of their usual performance. The figures reach us through Katy Letham and Alasdair Gray’s 2012 account of the 2005 survey in the journal Emergencias; the BMA’s own reports are no longer published online.

None of these behaviors were irrational.

The clearest evidence came from Mid Staffordshire NHS Foundation Trust. The Healthcare Commission’s 2009 inquiry found that the trust had routinely moved patients to assessment areas before they had been investigated or received a diagnosis — specifically to capture a four-hour clock stop before the deadline. Doctors at the trust admitted to the inquiry that they had been pressured to prioritize patients who were approaching the four-hour limit rather than patients who were most clinically urgent.

The clock had become the patient.

In 2010, researchers at the University of Sheffield analyzed the statistical pattern of A&E episode completions across English trusts. What they found was a spike. In the final twenty minutes before the four-hour mark, completed episodes — patients admitted and patients discharged, the admitted most of all — increased sharply, then fell just as sharply after. Writing in the British Medical Journal (BMJ), they were investigating what had already been named “last minute syndrome.” Their findings, in Letham and Gray’s summary, suggested that the four-hour rule had led to target-led rather than needs-led care. The spike was still there when the NHS’s own Strategy Unit examined six years of national data in 2019.

In 2010 the ninety-eight percent target was withdrawn. Published compliance had reached 97.7 percent in 2007, and was falling by the time the target went. What replaced it was a ninety-five percent version of the same four-hour standard, and that standard was still a key measure of hospital performance seven years later. Letham and Gray also synthesized the mathematical modeling of queuing theorists. Their summary of that modeling: meeting 98 percent compliance, even with the real improvements in management and patient flow, would not have been possible without “some form of patient re-designation or re-labelling taking place, leaving the integrity of reported performance open to question.”

Relaxing the target did not end the behavior. In 2017, Julie Eatock, Matthew Cooke and Terry Young, writing in the Future Healthcare Journal, compared two English emergency departments with similar patients and similar arrival patterns. In 2014-15 the two scored almost identically: one breached the four-hour standard on 4.68 percent of patients, the other on 4.49 percent. On the metric they were the same department. Behind the metric they were nothing alike. At the first hospital, 20.83 percent of patients left between three hours forty and four hours; at the second, 8.56 percent. Before the three-hour-twenty mark the two profiles differed by less than one percent. The gap opened only where the clock was being watched, and the number the system actually rewarded could not tell the two departments apart.

What the Department of Health had designed was a measurement. What it had attached to the measurement was a financial reward and a set of management consequences. What it had not designed was any mechanism to ensure that hitting the measurement corresponded to the thing the measurement was supposed to represent.

The result was predictable. The measurement became the goal.


Take a quick break in the waiting room. Nobody in this story set out to redefine care downward. Every hospital, every manager, every clock-watching decision was a rational response to what was actually being measured. The target didn’t fail because people gamed it. It failed because it was gameable, and gameable things get gamed.

It is worth saying plainly that the target worked at what it was first for. The thirty-seven-hour waits of the late 1990s came down — for a time. Eatock and her co-authors, who are critics of the measure, still credit it with driving enormous improvement in emergency provision and call the four-hour wait one of the shortest standards in the world. The emergency nurses Andy Mortimore and Simon Cooper interviewed called it a success with reservations. The argument here is about what the reward could see, and what it could not.

Letham and Gray’s summary of what happened has not been improved upon: “Rather than striving to provide good care within the target time, good care appears to have been redefined as achieving the target.”

That sentence describes a specific institutional failure, but the pattern is not specific to England or to healthcare. When a reward is attached to a measurement, the measurement becomes the target. The behavior the measurement was meant to represent — the outcome, the actual work, the actual care — becomes secondary. Not because anyone decided to stop caring. Because the reward structure made the number more immediately consequential than the thing the number stood for.

The four-hour target produced the measured behavior. It produced it in trusts across England, consistently, for more than fifteen years. The behavior you get is the behavior you designed the incentive for.

You have seen a version of this. The metric was real. The number was hit. The outcome was not. That gap is not only a reporting failure. It is a reward structure failure.

Where in your organization does a dashboard show green while the outcome it was designed to represent goes unexamined?

Behavior Follows Rewards. The pattern shows up wherever people are measured and rewarded.

Ask yourself: what is the difference between the metric you track and the outcome you want? The closer those two things are, the cleaner your incentive. The further apart they are, the more creative your people will be in hitting one while missing the other. The NHS set a four-hour clock. Teams hit the clock.

If you are working through an incentive design challenge in your organization — trying to understand why a policy, initiative, or team isn’t doing what you designed it to do — that is the work. Advisory Services →

Subscribe for free. Share it with someone who needs to read it.

—Wayne


Going Deeper

If you were the CEO of an NHS trust in 2004 — with mandatory targets, staged payments for hitting them, and no requirement on how you met them — and your staff had found a way to technically clear patients within four hours without improving care, would you have closed that loophole? Or documented it as a process improvement?

Leave a comment →

Last time: A professional golf league recruited the sport’s biggest stars, one of them for a reported $140 million. What its reward structure never asked for was whether any of it could pay for itself.

Next: Sixteen posts. Five continents. Five industries. One pattern. The review post looks at what the first sixteen posts proved — and what comes next. You will want to make time for that one.

In development — a university grading system where student evaluations, not academic rigor, determined a professor’s standing; a corporation whose internal ranking system made employees compete against each other instead of the competition; a UK welfare reform that moved the cost of illness off the government and onto the ill. More on the way.


Sources

Bevan, Gwyn, and Christopher Hood. “What’s Measured Is What Matters: Targets and Gaming in the English Public Health Care System.” Public Administration 84, no. 3 (2006): 517–538. DOI: 10.1111/j.1467-9299.2006.00600.x.

Department of Health. The NHS Plan: A Plan for Investment, A Plan for Reform. Cm 4818-I. London: The Stationery Office, July 2000.

British Medical Association. BMA Survey of A&E Waiting Times. London: British Medical Association, 2005. Cited in Letham and Gray (2012).

Healthcare Commission. Investigation into Mid Staffordshire NHS Foundation Trust. London: Commission for Healthcare Audit and Inspection, 2009.

Mason, S., J. Nicholl, and T. Locker. “Four Hour Emergency Target: Targets Still Lead Care in Emergency Departments.” BMJ 341 (2010): c3579. DOI: 10.1136/bmj.c3579.

Eatock, Julie, Matthew Cooke, and Terry P. Young. “Performing or not performing: what’s in a target?” Future Healthcare Journal 4, no. 3 (October 2017): 167–172. DOI: 10.7861/futurehosp.4-3-167.

Mortimore, Andy, and Simon Cooper. “The ‘4-hour target’: emergency nurses’ views.” Emergency Medicine Journal 24, no. 6 (June 2007): 402–404. DOI: 10.1136/emj.2006.044933.

Letham, Katy, and Alasdair Gray, on behalf of EMERGE. “The Four-Hour Target in the NHS Emergency Departments: A Critical Comment.” Emergencias 24, no. 1 (2012): 69–72.

Mayhew, Leslie, and David Smith. “Using queuing theory to analyse the Government’s 4-h completion time target in Accident and Emergency departments.” Health Care Management Science 11, no. 1 (2008): 11–21. DOI: 10.1007/s10729-007-9033-8.

Commission for Health Improvement. Report on the Clinical Governance Review on Surrey and Sussex Healthcare NHS Trust. London: The Stationery Office, 2002, para. 3.19, cited in Bevan and Hood (2006).

Wyatt, Steven. Waiting Times and Attendance Durations at English Accident and Emergency Departments. West Midlands: The Strategy Unit, NHS, February 2019.

BFR is written to be accessible, welcomed, and celebrated by every reader — not simplified, not elevated. Just clear.