Trial methodology guide
Weight Loss Injections: How the Pivotal Trials Were Designed
A methods-first reading of the injectable weight-management evidence. Who was enrolled, what ran underneath both arms, how long the trials lasted, what the withdrawal designs revealed, and which questions the protocols were never built to answer.
The short version
The evidence for weight loss injections rests on a small number of large phase 3 trials that share a common architecture: adults meeting a BMI threshold, randomised between an injectable agent and placebo, both arms placed on a diet and activity programme, a slow dose escalation, and a primary endpoint measured at 68 or 72 weeks. Almost every disputed claim in this field traces back to one of the design choices inside that architecture rather than to the drug itself.
This guide takes the design apart. It is about reading protocols and results tables, not about choosing a product. For what the outcome measures themselves mean, the companion piece is endpoints in metabolic peptide trials. For the pharmacology of the molecules involved, see the GLP-1 receptor agonist primer.

What is inside the category and what is not
The phrase covers two very different things, and separating them is the first appraisal step. On one side are injectable agents licensed for chronic weight management, each supported by a registered phase 3 programme with a published primary publication. On the other side are compounded, imported, or research-labelled preparations, which have no trial of their own. A trial result generated with an approved finished product does not describe a different preparation of nominally the same molecule, because concentration, excipients, and purity were all part of what was tested.
If you are orienting yourself to what products actually exist and how they are positioned, a consumer-facing comparison of weight loss injections is a reasonable place to start before returning here to check what sits underneath each claim. This guide covers only the licensed agents with registered trials, because those are the only ones with a design to appraise.
The anatomy of a phase 3 weight-management trial
Who gets in
Entry is normally a BMI at or above 30, or at or above 27 with at least one weight-related comorbidity. That single criterion does a great deal of work. It excludes people below the threshold entirely, so nothing in these trials describes what the drugs do at lower body weights. Protocols also routinely exclude recent major cardiovascular events, prior bariatric surgery, active malignancy, certain psychiatric histories, and in the obesity trials, type 2 diabetes, which is studied in separate trials because glycaemic status changes the weight response. SCALE, STEP 1, and SURMOUNT-1 all enrolled populations defined this way.[3],[7],[9]
What runs underneath both arms
This is the most frequently omitted element in secondary coverage. Participants in both arms, active and placebo, receive a structured diet and activity programme, typically a calorie deficit target plus a physical activity target plus regular counselling contact. That is why the placebo arms in these trials lose weight at all. It also means the reported treatment effect is the increment the drug adds on top of a supported lifestyle intervention, not the effect of the drug in isolation.
The intensity of that background programme is a design variable in its own right, and it was tested directly. STEP 3 paired semaglutide with intensive behavioural therapy rather than the lighter counselling used elsewhere in the programme, which moved both arms and changed the size of the gap between them.[5] Two trials of the same molecule with different background intensity are not interchangeable.
Titration, and why it is not a detail
These agents are escalated slowly over several weeks or months to the target dose because gastrointestinal adverse effects cluster during escalation. A trial with a longer escalation period spends more of its total duration below target dose, which affects the shape of the weight curve. It also affects how many participants ever reached the target dose and how many discontinued before they got there. When two trials report different results for similar molecules, the titration schedule is one of the first places to look.
Duration and the primary time point
The convention settled at 68 weeks for the semaglutide obesity programme and 72 weeks for the tirzepatide obesity programme.[3],[7] Those durations were chosen to satisfy the regulatory expectation of at least one year of data.[1] They are long enough to reach the plateau of the weight curve for most participants, and short enough that they say nothing about year three or year ten. A trial duration is not a statement about how long a benefit persists.
What the protocol has to show to count as a success
Regulatory guidance for weight-management products sets out a dual expectation: a mean weight loss meaningfully greater than placebo, and a significantly higher proportion of participants crossing a defined responder threshold.[1] That is why every one of these publications reports two kinds of number.
- The continuous result. The mean percent change in body weight from baseline to the primary time point, with a confidence interval on the difference between arms. STEP 1 reported 14.9 percent against 2.4 percent for placebo at 68 weeks.[3]
- The responder result. The proportion of participants achieving at least 5 percent, 10 percent, 15 percent, or 20 percent loss. In STEP 1, 86.4 percent of the semaglutide arm reached at least 5 percent against 31.5 percent on placebo.[3]
The two answer different questions. A mean tells you where the centre of the distribution sits; a responder proportion tells you how many people crossed a clinically meaningful line. Neither describes the spread, and the spread in these trials is wide. Reporting only the mean makes a heterogeneous response look like a uniform one.
Why one trial can publish two different headline percentages
Participants stop taking study drug, drop out, or start something else. The rule for handling those events is called the estimand, and it must be pre-specified.[2] Two rules are common in this literature. One asks what happened to everyone randomised, regardless of whether they stayed on treatment. The other asks what the drug does in people who take it as directed.
SURMOUNT-1 published both for its highest dose arm: approximately 20.9 percent under the first rule and approximately 22.5 percent under the second.[7] Both are correct. They answer different questions, and the gap between them is a measure of how much discontinuation cost the result. A headline percentage quoted without naming the estimand has dropped the clause that makes it interpretable.

Schematic only. The axes carry no values and the curves are not plotted from trial data; the figure shows the characteristic shape discussed in the withdrawal trials below.
The withdrawal designs, and what they settled
Two trials in this literature were built specifically to test durability, and they are the most informative designs in the whole set because they answer a question the standard placebo-controlled trial cannot.
Both used the same architecture: everyone receives active drug during an open-label run-in, then those who complete it are randomised either to continue or to switch to placebo. Because randomisation happens after the run-in, the comparison is between continuing and stopping in people who have already responded, which isolates the maintenance question cleanly.
- STEP 4. After a 20-week semaglutide run-in, continuing treatment produced roughly 7.9 percent further loss while switching to placebo produced roughly 6.9 percent regain over the following period.[4]
- SURMOUNT-4. After a 36-week tirzepatide lead-in, continuing added roughly a further 5.5 percent while those switched to placebo regained substantially.[8]
The consistent finding is that the effect is treatment-dependent. That is a pharmacological statement, not a moral one: the drug suppresses appetite while it is present at therapeutic concentration, and the physiology that defends body weight is not permanently altered by having been suppressed. Any claim that a course of injections produces durable weight loss after discontinuation is not supported by these designs.
The blinding problem nobody can design away
These trials are double-blind and placebo-controlled, which is the correct design. But gastrointestinal adverse effects are common, dose-related, and perceptible, so a substantial proportion of participants can form an accurate guess about their allocation. This is functional unblinding, and it matters most for outcomes that depend on participant behaviour, which in a weight trial includes adherence to the background diet and activity programme.
There is no clean fix. Active-comparator designs such as STEP 8, which randomised between two active injectables rather than against placebo, reduce the problem because both arms produce side effects.[6] Objective endpoints such as the adjudicated cardiovascular events in SELECT are less vulnerable than self-reported ones.[10] Our guide to bias, blinding, and why double-blind phase III matters works through the general problem, and the phase 1 to 4 guide explains where in development each design fits.
Eight questions for any weight loss injection trial
Work through these against the primary publication, not against a summary of it. The registry record answers several of them before results are even published; see how to read a ClinicalTrials.gov listing.
- What were the BMI entry criteria, and who was excluded? The exclusions define the boundary of what the result can describe.
- Was diabetes present at baseline? Weight response differs between populations with and without type 2 diabetes, which is why they are studied separately.
- What background programme ran in both arms, and how intensive was it? This sets the placebo arm and therefore the size of the gap.
- What was the comparator? Placebo, another active injectable, or an open-label run-in followed by withdrawal. These support different claims.
- How long was the escalation, and how many reached target dose? Check the disposition table, not the abstract.
- Which estimand generated the headline number? If the publication reports two, note which one is being quoted at you.
- What were the discontinuation rates, and why did people stop? A large effect in a trial with heavy attrition is a different finding from the same effect with low attrition.
- Was there any off-treatment follow-up? Without it, the trial says nothing about what happens after the last injection.
What the designs were never built to answer
- Multi-year and lifetime use. Randomised follow-up in the weight indication is measured in months to a small number of years. Chronic use over decades has not been studied under randomisation.
- Body composition. Percent body weight is the regulatory endpoint. How much of that loss is fat mass against lean mass is assessed in smaller substudies and is not what any of these trials were powered for.
- Most product pairings. Direct randomised comparisons between injectables are rare. Rankings assembled by lining up percentages from separate trials are indirect comparisons wearing the clothes of direct ones.
- Stopping strategies. The withdrawal trials establish that weight returns after abrupt discontinuation. They do not test tapering, intermittent dosing, or maintenance at a reduced dose, so no validated protocol for any of those exists in this literature.
- Populations that were excluded. Older adults with high frailty burden, people with significant psychiatric comorbidity, and those with recent cardiovascular events were commonly outside the entry criteria.
- Non-approved preparations. Compounded and research-labelled material sits entirely outside this evidence base. There is no trial to appraise.
Reading the evidence is not the same as being prescribed
Understanding how these trials were built lets you tell a supported claim from an unsupported one, and it lets you see which trial population a claim actually describes. It does not substitute for clinical assessment. Eligibility, interacting medicines, contraindications including the class boxed warning relating to thyroid C-cell tumours, tolerability, and monitoring all belong with a prescribing clinician, who is also the right person to judge which of these trial populations a given individual most resembles.
To go deeper into the specific programmes discussed here, see our spotlights on STEP, SURMOUNT, and SCALE, each of which lists the constituent trials and their primary publications.
Limitations of the evidence
This guide describes how the pivotal trials in this literature were designed and how to read them. It is not a treatment guide and contains no dosing guidance. Trial figures quoted are from the primary publications named in the references and may differ from later meta-analyses, extension reports, or longer follow-up. Design conventions described here are those used in the named phase 3 programmes and do not necessarily apply to smaller trials, open-label extensions, or observational studies of the same agents. Every pivotal trial cited was funded by the manufacturer of the drug under test. Regulatory guidance quoted reflects the referenced documents at the time of writing; thresholds and expectations are revised periodically. Nothing here applies to compounded or unapproved preparations, which do not have pivotal trial evidence of their own.
References
Citations are annotated with an evidence tier reflecting study design and replication. See Methodology for criteria.
- 1.U.S. Food and Drug Administration · Guidance for Industry: Developing Products for Weight Management · 2007Validated
- 2.International Council for Harmonisation (ICH) · ICH E9(R1) Addendum on Estimands and Sensitivity Analysis in Clinical Trials · 2019Validated
- 3.Wilding JPH, Batterham RL, Calanna S, et al. · Once-Weekly Semaglutide in Adults with Overweight or Obesity (STEP 1) · New England Journal of Medicine · 2021PMID 33567185DOI 10.1056/NEJMoa2032183NCT03548935Validated
- 4.Rubino D, Abrahamsson N, Davies M, et al. · Effect of Continued Weekly Subcutaneous Semaglutide vs Placebo on Weight Loss Maintenance in Adults With Overweight or Obesity: The STEP 4 Randomized Clinical Trial · JAMA · 2021PMID 33755728DOI 10.1001/jama.2021.3224NCT03548961Validated
- 5.Wadden TA, Bailey TS, Billings LK, et al. · Effect of Subcutaneous Semaglutide vs Placebo as an Adjunct to Intensive Behavioral Therapy on Body Weight in Adults With Overweight or Obesity: The STEP 3 Randomized Clinical Trial · JAMA · 2021PMID 33625476DOI 10.1001/jama.2021.1831NCT03611582Validated
- 6.Rubino DM, Greenway FL, Khalid U, et al. · Effect of Weekly Subcutaneous Semaglutide vs Daily Liraglutide on Body Weight in Adults With Overweight or Obesity Without Diabetes: The STEP 8 Randomized Clinical Trial · JAMA · 2022PMID 34965561DOI 10.1001/jama.2021.23619NCT04074161Validated
- 7.Jastreboff AM, Aronne LJ, Ahmad NN, et al. · Tirzepatide Once Weekly for the Treatment of Obesity (SURMOUNT-1) · New England Journal of Medicine · 2022PMID 35658024DOI 10.1056/NEJMoa2206038NCT04184622Validated
- 8.Aronne LJ, Sattar N, Horn DB, et al. · Continued Treatment With Tirzepatide for Maintenance of Weight Reduction in Adults With Obesity: The SURMOUNT-4 Randomized Clinical Trial · JAMA · 2024DOI 10.1001/jama.2023.24945NCT04660643Validated
- 9.Pi-Sunyer X, Astrup A, Fujioka K, et al. · A Randomized, Controlled Trial of 3.0 mg of Liraglutide in Weight Management · New England Journal of Medicine · 2015PMID 26132939NCT01272219Validated
- 10.Lincoff AM, Brown-Frandsen K, Colhoun HM, et al. · Semaglutide and Cardiovascular Outcomes in Obesity without Diabetes (SELECT) · New England Journal of Medicine · 2023PMID 37952131DOI 10.1056/NEJMoa2307563NCT03574597Validated