Research · 10 min read
What a Trial Asks When It Measures How You Feel
Some outcomes come off a machine and some come out of a person's mouth. The second kind has its own rulebook, its own failure mode, and — in this drug class — almost no presence on the labels at all.
Key takeaways
- A patient-reported outcome is defined by the FDA as a report coming directly from the patient without interpretation by anyone else, and the instrument is the questionnaire plus its supporting documentation.
- A labeling claim built on one must be consistent with the instrument's documented measurement capability, reviewed the same way as any other claim.
- The weight-reduction guidance warns that a population without enough baseline limitation cannot show a meaningful change in score, and recommends prespecifying a subgroup in that case.
- Across six prescribing documents in this class, SF-36, IWQOL, EQ-5D and the phrase quality of life appear zero times, with positive controls confirming the search worked.
- One label names patient-reported instruments — the PROMIS Short Form Sleep-Related Impairment 8a and the Epworth Sleepiness Scale — and only in its sleep apnea sections.
- That label defines remission partly by a questionnaire score: an index below 5, or 5 to 14 with an Epworth score of 10 or less, reached by 42.2 and 50.2 percent on treatment against 15.9 and 14.3 percent on placebo.
- Its sleep-related impairment result is reported as an improvement with no magnitude printed, while the machine measurement in the same section carries baselines, differences and confidence intervals.
Answer first: a questionnaire can be an endpoint, and it has to earn it
A trial can measure a person's weight on a scale, their blood pressure with a cuff, and their breathing overnight with a sleep study. It can also ask them how they are doing. The answer to that question is a real outcome, and regulators have a name and a rulebook for it.
The FDA defines a patient-reported outcome as any report of the status of a patient's health condition that comes directly from the patient, without interpretation of the patient's response by a clinician or anyone else. That last clause is the whole definition. The moment somebody else translates the answer, it stops being this kind of measurement.
The instrument is more than the questions. The agency defines it as a questionnaire plus the information and documentation that support its use — the evidence that the thing measures what it claims to measure. A company cannot invent a survey, run it, and print the result on a label. The agency reviews the evidence that a particular instrument measures the concept claimed, and a claim can appear in labeling only if it is consistent with the instrument's documented measurement capability.
The failure mode nobody expects: no room to move
The most interesting rule in this area is not about honesty. It is about arithmetic, and it is the same problem as a blood-sugar baseline that sits near the normal range.
The FDA's draft guidance on weight-reduction drugs uses physical functioning as its worked example. A sponsor wanting to demonstrate clinical benefit on a set of functional impacts should first specify and define them. They must be relevant and important to patients with obesity or overweight, and likely to demonstrate meaningful and interpretable changes in the planned trial.
Then the crucial condition. The sponsor should consider whether the target population will have sufficient limitation in their physical functioning, and how the extent of limitation at baseline may affect the instrument's ability to observe a clinically meaningful within-patient score change. And if all randomized subjects do not have sufficient limitation in physical functions, the guidance recommends prespecifying a subgroup to be analyzed, with the analysis plan submitted to the agency before the trial runs.
Translated: if the people you enrolled can already climb the stairs without difficulty, no instrument about climbing stairs can show them improving. A flat result would say nothing about the drug. It would say the ruler had no room on it.
This is why a program page reporting that its members feel better is a weaker statement than it sounds. Feeling better is measurable. Whether it was measured with an instrument capable of registering a change, in a group with somewhere to move, is the question that separates a finding from a sentiment.
Of six labels in this class, one names a patient-reported instrument
Six FDA prescribing information documents were read in full for this article: the two approved weight products, the two injectable diabetes products, the combined oral semaglutide label, and the newer oral weight product.
A full-text search across all six found no occurrence of SF-36, of IWQOL, of EQ-5D, or of the phrase quality of life. Not once, in any of them. The same search pass counted the word placebo between 134 and 514 times per label and a deliberately meaningless control string zero times, so the ruler was working and the documents were what their filenames claimed.
One label breaks the pattern, and only in one indication. The tirzepatide weight product names the Patient-Reported Outcomes Measurement Information System — PROMIS — Short Form Sleep-Related Impairment 8a, and the Epworth Sleepiness Scale. Both appear in its obstructive sleep apnea sections and nowhere else. The other five labels mention neither.
That is a striking result for a drug class sold overwhelmingly on how people expect to feel. The evidence in these labels is almost entirely about what instruments and machines recorded: kilograms, centimeters, millimeters of mercury, events per hour, percentages of hemoglobin. The subjective side of the experience is, on the page, nearly absent.
Where a questionnaire is part of the definition of success
The sleep apnea sections are worth reading closely, because they show a patient-reported score doing something more consequential than sitting in a table.
The primary endpoint in those two trials is the apnea-hypopnea index — events per hour, measured by a sleep study. That is a machine measurement, and the label reports it: baseline means near 50, falls of 25.3 and 29.3 events per hour on treatment against 5.3 and 5.5 on placebo.
But the label also reports a composite outcome it calls remission or mild non-symptomatic obstructive sleep apnea, and the definition is a hybrid. A patient counts if their index falls below 5, or if their index is between 5 and 14 and their Epworth Sleepiness Score is 10 or less. In the first trial, 15.9 percent of the placebo group and 42.2 percent of the treated group met it; in the second, 14.3 percent and 50.2 percent.
Read the definition again. For a person in the middle band, whether they count as being in remission depends on a questionnaire score. The word non-symptomatic in the endpoint's name is not decoration — it is the part the machine cannot measure, and the trial used a patient-reported instrument to measure it.
That is a good illustration of why these instruments exist. An index of events per hour describes what happened to a person's breathing. It does not describe whether they are still exhausted during the day. The composite endpoint says both matter, and defines success in terms of both.
An improvement without a magnitude
The same section carries a smaller finding that cuts the other way, and it is worth naming.
The label states that in the two sleep apnea studies, treated patients showed improvement in sleep-related impairment compared to those who received placebo, and that sleep-related impairment was assessed using the PROMIS Short Form Sleep-Related Impairment 8a. It names the instrument. It reports a direction. In the section as printed, it does not report a number.
A direction without a magnitude is not nothing — it is a statement made under the rules that govern a label, which is more than most sources offer. It is also not comparable to anything. You cannot set it beside another trial, or against placebo in a way you can check, or against the size of the change on the machine measurement in the same study.
That contrast, in one section of one label, is the whole subject in miniature. The apnea index gets baselines, changes, differences from placebo, confidence intervals, and a multiplicity-controlled significance marker. The patient's own report of how impaired their days were gets a sentence.
What this means when a program advertises how people felt
Programs in this category sell an experience, and their marketing leans hard on it — more energy, better sleep, more confidence, a life that fits again. Those are real things and they are measurable things. The question is whether they were measured.
Four questions do most of the work.
Which instrument, by name. The agency's definition is a questionnaire plus the documentation supporting its use. A claim with no named instrument has no documentation behind it, whatever else is true of it.
Asked of whom, and did that group have room to improve. The guidance's own condition is whether the population had sufficient limitation at baseline for a meaningful change in score to be observable. A group already doing well on the concept being measured cannot demonstrate improvement in it.
Compared with what. Everyone in these trials, placebo included, was on a diet and exercise program, and people who start a new health program generally report feeling better at first. Without a comparison group, a satisfaction figure measures enthusiasm as much as effect.
And who collected it. A trial instrument reviewed by the agency for whether it measures the concept claimed is a different object from a customer survey, even when both produce a percentage. Both can be honestly reported. Only one of them has a rulebook behind it.
Sources
- Patient-Reported Outcome Measures: Use in Medical Product Development to Support Labeling Claims — Guidance for IndustryThe definition of a patient-reported outcome as any report of the status of a patient's health condition coming directly from the patient without interpretation by a clinician or anyone else; the definition of a PRO instrument as a questionnaire plus the information and documentation supporting its use; the statement that a claim may be supported if it is consistent with the instrument's documented measurement capability and that the amount and kind of evidence required is the same as for any other labeling claim; and the statement that the agency reviews the evidence that a particular instrument measures the concept claimed.
- Obesity and Overweight: Developing Drugs and Biological Products for Weight Reduction — Guidance for Industry (Draft Guidance)The treatment of clinical outcome assessments as possible secondary endpoints supporting a labeling claim; the physical functioning example and the recommendation to specify functional impacts relevant and important to patients that are likely to show meaningful and interpretable change; the condition about whether the target population has sufficient limitation at baseline for a clinically meaningful within-patient score change to be observed; and the recommendation to prespecify a subgroup and submit the analysis plan for agency agreement before the trial when it does not.
- ZEPBOUND (tirzepatide) injection — full prescribing information, Section 14 Clinical Studies, obstructive sleep apnea studiesThe naming of the PROMIS Short Form Sleep-Related Impairment 8a and the statement that treated patients showed improvement in sleep-related impairment compared with placebo without a magnitude printed in that passage; the Epworth Sleepiness Scale in the table abbreviation key; the composite endpoint defined as an apnea-hypopnea index below 5, or 5 to 14 with an Epworth score of 10 or less, with proportions of 15.9 against 42.2 percent and 14.3 against 50.2 percent; and the apnea-hypopnea index baselines near 50 with falls of 25.3 and 29.3 on treatment against 5.3 and 5.5 on placebo.
Frequently asked questions
What is a patient-reported outcome?
The FDA defines it as any report of the status of a patient's health condition that comes directly from the patient, without interpretation of the patient's response by a clinician or anyone else. The instrument is defined as a questionnaire plus the information and documentation that support its use.
Can a company put a survey result on a drug label?
Only under conditions. The agency's guidance says a claim can be supported by a patient-reported instrument if the claim is consistent with that instrument's documented measurement capability. The FDA reviews the evidence that the instrument measures the concept it is claimed to measure. The amount and kind of evidence required is the same as for any other labeling claim.
Why does the guidance worry about whether patients are limited enough at baseline?
Because an instrument cannot register improvement in something the enrolled group was not struggling with. The weight-reduction guidance asks sponsors to consider whether the target population has sufficient limitation in physical functioning for a meaningful within-patient score change to be observable, and to prespecify a subgroup if not.
Do these drug labels report quality of life?
Not in the six prescribing documents read for this article. A full-text search across all six found no occurrence of SF-36, IWQOL, EQ-5D or the phrase quality of life, while control terms in the same search pass returned the expected hits. One label names patient-reported instruments, and only within its sleep apnea sections.
Which instruments does that one label name?
The tirzepatide weight label names two. The PROMIS Short Form Sleep-Related Impairment 8a was used to assess sleep-related impairment in its two obstructive sleep apnea studies. The Epworth Sleepiness Scale appears in the abbreviation keys of its sleep apnea tables, and in one endpoint definition.
How can a questionnaire be part of whether a trial counts someone as improved?
In those sleep apnea trials the label reports a composite outcome called remission or mild non-symptomatic obstructive sleep apnea. It is defined as an apnea-hypopnea index below 5, or an index of 5 to 14 together with an Epworth Sleepiness Score of 10 or less. For patients in that middle band, the questionnaire score decides whether they meet the endpoint. The reported proportions were 15.9 percent on placebo against 42.2 percent on treatment in one study, and 14.3 against 50.2 percent in the other.