How Outcome Scores Get Misread in a Medical Record Review

How Outcome Scores Get Misread in a Medical Record Review

Icon representing a calendar or date selection interface.
Published Date :

October 1, 2026

Icon representing a calendar or date selection interface.
Modified Date :

October 1, 2026

Home
>
Blog
>
>
How Outcome Scores Get Misread in a Medical Record Review

Outcome scores in a review, in brief:

  • A score is structured self-report. The patient answered the questions. The number is a tally of those answers, not a measurement taken from the patient.
  • Numbers get promoted. A percentage or a point total reads as objective simply because it is numeric, and that promotion happens silently in summaries.
  • The instrument has to be named. Two scores are comparable only when the same instrument was used, in the same version, with the same recall window.
  • Administration mode is rarely recorded. The same questionnaire can be completed alone in a waiting room or read aloud by staff, and the chart usually does not say which.
  • A change is not automatically meaningful. Whether a difference crosses a clinically important threshold is instrument-specific and population-dependent, and applying that threshold is a clinical judgment.
  • What the review can do: name the instrument, date every administration, quote the score, flag gaps and mode. What it cannot do is decide what the change means.

A defense summary reads: "Function improved from 62 percent disability to 28 percent over nine months."

That sentence sounds like it came from an instrument on a patient. It came from two questionnaires the patient filled out in a waiting room, eleven visits apart, with no record of who handed them over or whether anyone explained the questions.

Both facts can be true at once. The scores are real evidence, and the sentence overstates what they are. That gap is where outcome instruments cause trouble in a medical record review, and it is a gap that opens because the evidence arrives as a number.

What a Score Actually Is

A patient-reported outcome measure is a questionnaire the patient answers, converted to a score by a fixed rule. The rule is standardized. The input is the patient's account of their own function.

The ones that turn up most often in injury and disability files:

  • Oswestry Disability Index. A self-report questionnaire on how back pain affects daily activities, reported as a disability percentage.
  • Neck Disability Index. The same design adapted to the cervical spine.
  • DASH and QuickDASH. Self-report on arm, shoulder and hand function, the short form using fewer items than the full version.
  • Numeric pain rating. A single self-reported number, usually zero to ten, recorded at almost every visit.
  • PHQ-9 and GAD-7. Screening instruments for depressive and anxiety symptoms. The PHQ-9 is a nine-item instrument asking about the frequency of symptoms over the preceding two weeks, per CDC documentation of its use in national health surveys.
  • PROMIS measures. A family of instruments developed through the National Institutes of Health, reported on their own scale rather than as a raw total.

Nothing in that list is a measurement of the patient. Every one of them is a measurement of what the patient reported. That distinction sits on the same axis as separating patient reports from provider findings, with one difference that matters: these reports arrive pre-converted into numbers, and numbers survive summarization better than attribution does.

Numbers get promoted, prose does not
A note saying the patient reports difficulty climbing stairs keeps its attribution through three rounds of summarizing. The same information expressed as a 62 percent disability index loses it at the first pass, because a percentage looks like something somebody measured. The instrument is not the problem. The arithmetic of summarizing is.

Why Serial Scores Are Worth Finding

In most injury files these are the only repeated, numerically comparable measures of function that exist.

Consider what else is available. Range of motion is measured inconsistently and often estimated. Strength grading is coarse. Imaging findings describe anatomy, not function, and they do not change on the schedule symptoms do. Narrative descriptions of function are rich and not comparable across providers.

An instrument administered at intake, at three months and at discharge produces three values on one scale, collected by the same rule. That is a trend line, and it is the closest thing a typical chart holds to one. This is why they belong in a narrative summary with their dates attached, and why a review that skips them has left the most comparable evidence in the file unused.

The value and the risk come from the same property. Because they are comparable, they get compared. Because they are numeric, they get compared carelessly.

Six Ways the Comparison Breaks

Each of these produces a number that is real and a comparison that is not.

  • Different instruments, same body part. An Oswestry score at intake and a different back-specific instrument at follow-up cannot be placed side by side. Both describe back function. Neither converts into the other.
  • Different versions of the same instrument. Short forms and full versions use different item counts. A QuickDASH and a full DASH are related instruments, not interchangeable values, and a summary that lists both as DASH scores has built a trend out of two different things.
  • Recall window drift. Instruments ask about a defined period, whether that is today, the past week or the past two weeks. A score describing the last two weeks and a pain rating describing this moment are answering different questions.
  • Administration mode. The same questionnaire can be completed alone on paper, filled in on a tablet, or read aloud by staff. National survey documentation for the PHQ-9 describes it being administered by trained interviewers rather than self-completed, which shows the same instrument moves between modes. Charts almost never record which mode was used, and the review should say so rather than assume.
  • Incomplete administration. Most of these instruments have scoring rules for missing items, and those rules have limits. A questionnaire with three sections left blank may still produce a printed score, and that score is not equivalent to a complete one.
  • Transcription and scoring error. Someone tallies the items, and someone else types the total into the note. A digit lost between those steps produces a clean number that no longer matches the completed form. Where the underlying questionnaire is in the production, the score can be checked against it. Where it is not, the score is a report of a report.

Preparing an evaluation on a file with years of questionnaire scores? Get every administration named, dated and traced to its form.

When Is a Change Meaningful?

That question has a real answer, and it is not one a record review gets to give.

Outcome measurement research works with the idea of a smallest change that matters to a patient, often discussed as a minimal clinically important difference. The principle is that a difference has to exceed some threshold before it represents genuine change rather than measurement noise or ordinary day-to-day variation.

Two things follow for a reviewer. Those thresholds are specific to the instrument, and they also vary with the population studied, the condition and the method used to derive them, which is why published values for a single instrument are not uniform. And selecting a threshold and applying it to a particular claimant is a clinical determination.

So the review states the values and the dates. It does not characterize a change as significant, meaningful, or within normal variation. That belongs to the treating provider, the evaluating physician or a retained expert who can name the threshold they are using and defend it.

Ceiling, Floor and the Flat Line

An unchanged score is the most commonly overread value in the set.

A patient already answering at the most severe end of an instrument has nowhere further to go on it. Worsening function produces an identical score, because the instrument has run out of range. The same applies at the other end, where a patient near full function shows no improvement on the instrument while still reporting change.

A flat series can therefore mean stability, an instrument at its limit, a questionnaire completed without much attention, or a score copied forward from a prior visit into a template. The records sometimes distinguish these and often do not. Where they do not, a summary that reports no change has already interpreted.

Where the Instrument Came From

Not every questionnaire in a file is a validated instrument, and the record does not label the difference.

Alongside the recognized measures, injury files carry practice-designed intake forms, pain diagrams, symptom checklists and questionnaires supplied by an attorney or a vendor. Some produce numbers. Those numbers look like the others on the page.

The review's job is to name the source of each one. A published instrument administered on a date is a different record from a clinic's own form, and a form supplied by a party to the litigation is different again. None of the three is worthless. All three get read differently, and only the first supports comparison against published expectations at all.

A disability percentage is not a measurement of the patient. It is arithmetic performed on what the patient said.

quotes-icon

Four Readers, One Requirement

Everyone downstream needs the same thing from these records, and each needs it for a different reason.

  • IME and QME examiners. An examiner producing hands-on findings is often asked to reconcile them against the treating record's scores. Knowing those scores are self-report, and which instrument produced each one, is part of what medical record review before IME reports is for.
  • Plaintiff firms. Serial scores can document persistence of symptoms in the patient's own terms across years, which narrative notes do less precisely. Presented as objective measurements, the same scores invite a cross-examination the file did not need.
  • Defense counsel and carriers. An improving series is real evidence about reported function. Characterized as measured improvement, it overstates itself, and the overstatement is the part that gets challenged.
  • Life care planners. A current functional score is a data point about reported status on a date. Projecting decades of care from a trend in self-report requires the clinician's judgment about what the trend represents.

Where the Review Stops

Three layers, kept apart.

What the records document. The instrument named, the date administered, the score recorded, and the completed form where it is in the production.

What a reviewer can identify. Instrument and version changes across the series, missing or partial administrations, undated scores, a value that does not match its form, a flat series, and the absence of any recorded administration mode.

What requires a qualified professional. Whether a change is clinically meaningful, whether a score is consistent with the clinical picture, whether an instrument was appropriate for this patient, and what any of it means for causation, impairment or future need.

Instrument selection and interpretation also depend on condition, specialty and the clinical context in which the questionnaire was given. Where that bears on a case, it is a question for the evaluating clinician rather than a line in a summary.

Behind a score-level review

99.8%

Accuracy rate

Published figure, with values checked against source forms.

13+

Working with examiners, planners and litigation teams.

2M+

Records analyzed

Cumulative across medical-legal engagements.

Outcome Score Review FAQs

Are outcome scores objective findings?

Orange downward pointing arrow icon.

No. Instruments like the Oswestry Disability Index, DASH and PHQ-9 are questionnaires the patient answers, converted to a score by a fixed rule. The rule is standardized. The input is self-report, and a review should record it that way.

Why do outcome scores get misread more often than written symptom reports?

Orange downward pointing arrow icon.

Because they arrive as numbers. A note saying the patient reports difficulty climbing stairs keeps its attribution through several rounds of summarizing. The same information as a disability percentage loses it immediately, since a percentage looks measured.

When can two outcome scores be compared?

Orange downward pointing arrow icon.

When the same instrument, in the same version, with the same recall window produced both. Different instruments for the same body part, or a short form against a full version, do not convert into each other.

Can a medical record review say whether a change in score is significant?

Orange downward pointing arrow icon.

No. Whether a difference exceeds a clinically important threshold depends on the instrument, the condition and the population studied, and applying a threshold to a specific claimant is a clinical determination for the treating provider or an evaluating expert.

What does an unchanged score mean?

Orange downward pointing arrow icon.

It can mean stability, an instrument at the limit of its range, a questionnaire completed without attention, or a value copied forward into a template. The review reports what the series did and flags which of those the records can and cannot rule out.

Why does it matter how a questionnaire was administered?

Orange downward pointing arrow icon.

The same instrument can be completed alone on paper, filled in on a tablet, or read aloud by staff. National survey documentation describes the PHQ-9 being administered by trained interviewers rather than self-completed, so mode varies in practice. Charts rarely record it, and the review should note the absence.

Should the completed questionnaires be requested, not just the scores?

Orange downward pointing arrow icon.

Yes where possible. A score transcribed into a progress note is a report of a report. With the form in the production, the value can be checked against the completed items and partial administrations become visible.

Orange downward pointing arrow icon.

Orange downward pointing arrow icon.

Orange downward pointing arrow icon.

What to Request

Five things turn a scatter of numbers into a usable series.

  • Every administration listed with the instrument named in full, including the version or short form.
  • The completed questionnaires themselves, not only the scores transcribed into progress notes.
  • Each score dated to its administration rather than to the visit it was discussed in.
  • Instrument or version changes flagged wherever the series is not continuous.
  • Partial administrations, undated scores and unrecorded administration mode identified as such.

Read the Number as a Sentence

Every one of these scores started as a person answering questions about their own life. The instrument turned those answers into a figure so they could be compared, which is useful, and the figure then stopped looking like an answer, which is the problem.

Put the sentence back. Name the instrument, date it, quote it, say who completed it, and leave the meaning to the people qualified to assign it. The series still does its work, and it stops claiming more than it holds.

LezDo TechMed supports examiners, planners and litigation teams through our medical record review services. We name, date, quote and flag. The clinical and legal conclusions stay with you and your experts.

Source Credit: The description of the PHQ-9 as a nine-item instrument asking about symptom frequency over the preceding two weeks, and its administration by trained interviewers in a national health survey, is from Centers for Disease Control and Prevention National Health and Nutrition Examination Survey documentation. PROMIS measures were developed through the National Institutes of Health. Instrument descriptions here are general; no scoring thresholds or clinically important difference values are asserted, because those are instrument-specific and population-dependent. The file described in this article is a hypothetical illustration, not a client matter. Company figures are LezDo TechMed's published figures. This article is general information for medico-legal and claims professionals, not legal or medical advice.

Source Credit :  All metrics derived from LezDo TechMed’s internal project data.
Shabila Thomas

Shabila Thomas

Shabila Thomas is a Certified Legal Nurse Consultant (CLNC) and Medical-Legal Research Analyst with over two years of experience in medical record review, medico-legal research, and content development. She specializes in blogs, articles, and content that decode complex medical information, industry trends, and regulatory updates for the medico-legal field. Her clinical background and research-first approach help law firms, medical evaluators, and insurance professionals understand complex medical data, identify relevant insights, and make faster, better-informed decisions.