Foundational guide
GLP-1 Medications: How to Compare the Evidence Behind Each Approved Product
The approved landscape read as an evidence base rather than a product list: which trial sits under each medication, why the headline percentages cannot be lined up side by side, and which questions the trials were never designed to answer.
The short version
GLP-1 medications are usually discussed as if they were interchangeable entries in a ranked list. They are not. Each approved product carries its own trial programme, its own enrolled population, its own comparator, its own primary endpoint, and its own analysis conventions. Two products can report weight-loss percentages that look directly comparable while having been measured in different people, over different durations, against different control arms, and under different statistical rules for handling participants who stopped taking the drug.
This guide does not rank the class. It teaches the appraisal: how to find the trial that actually supports a claim, how to tell whether a comparison between two products is randomised or merely arithmetic, and how to recognise the places where the evidence base is genuinely thin. For the receptor biology and pharmacology behind the class, the companion piece is our GLP-1 receptor agonist primer. This guide picks up where that one stops, at the point where mechanism becomes marketed product.

Schematic. The four columns are the dimensions on which any claim in this class should be checked; the marks illustrate that coverage across those dimensions is uneven by trial type. They are not a rating of any individual product.
What the phrase actually covers
The everyday phrase covers three pharmacologically distinct groups that are frequently merged in coverage:
- Single GLP-1 receptor agonists. Exenatide, lixisenatide, liraglutide, dulaglutide, and semaglutide. All act at the GLP-1 receptor, but they differ in structural origin (exendin-4 derived versus human GLP-1 derived), dosing interval, and route.[1]
- Dual incretin receptor agonists. Tirzepatide activates both the GLP-1 and the GIP receptor. It is routinely grouped with the class in general coverage, and it is a different pharmacological agent with its own trial programme.
- Investigational multi-receptor agents. Molecules adding glucagon receptor agonism are in clinical development. Phase 2 signals are not approvals, and an agent in this group has no marketing authorisation to appraise.
Because the boundaries of the phrase are loose, a consumer-facing comparison list is a reasonable starting point for seeing which products exist and how they are positioned; this rundown of GLP-1 medications lays the marketed products out side by side. Use a list like that to establish what you are looking at, then come back and check each claim against the trial that generated it. The rest of this guide is about doing that second step properly.
Which trial sits under which claim
Every approved indication rests on a specific programme. The table below maps the pivotal evidence for the products that dominate current discussion. Read it as a map of what has been tested, not as a ranking.
| Molecule | Pivotal programme | Population enrolled | Comparator | What it does not establish |
|---|---|---|---|---|
| Liraglutide | LEADER (T2DM, CV outcomes); SCALE (weight) | T2DM at high cardiovascular risk; separately, adults with obesity | Placebo on top of standard care | Relative standing against any other GLP-1 medication in diabetes |
| Semaglutide (injectable) | SUSTAIN, STEP, SELECT | T2DM; adults with obesity without T2DM; obesity with established CVD | Placebo; liraglutide in STEP 8 | Outcome benefit in people without established cardiovascular disease |
| Semaglutide (oral) | PIONEER | T2DM across background therapies | Placebo and active oral agents, varying by trial | Equivalence of the oral and injectable forms at any given milligram figure |
| Tirzepatide | SURPASS (T2DM); SURMOUNT (weight) | T2DM; adults with obesity with and without T2DM | Placebo, insulin comparators, and semaglutide 1 mg in SURPASS-2 | Cardiovascular outcome superiority over placebo, which was not the design of its outcomes trial |
T2DM = type 2 diabetes mellitus. CVD = cardiovascular disease. Programme pages on this site list the constituent trials, sponsors, and primary publications for each entry.
Why the league table is the wrong instrument
The most common error in coverage of GLP-1 medications is to place headline percentages from separate trials into a single ranked table. STEP 1 reported a mean body-weight reduction of 14.9 percent with semaglutide 2.4 mg at 68 weeks against 2.4 percent with placebo.[6] SURMOUNT-1 reported approximately 22.5 percent with tirzepatide 15 mg at 72 weeks against 2.4 percent with placebo.[7] Those two numbers are not a comparison. They differ on at least five axes at once:
- Duration. 68 weeks against 72 weeks. Weight-loss curves in this class are still flattening at that stage, so four extra weeks are not neutral.
- Enrolled population. Baseline body weight, baseline BMI, sex distribution, comorbidity burden, and geography all differ between programmes, and percentage change from baseline is sensitive to all of them.
- Background intervention. The intensity of the diet and activity programme running underneath both arms differs across trials, and it moves the placebo arm as well as the active arm.
- Titration schedule. How quickly participants reached the target dose affects both tolerability and the proportion who reached full dose at all.
- Analysis convention. The rule for handling participants who stopped treatment or started a rescue intervention differs between analyses, and it moves the headline number by a meaningful margin.
An indirect comparison across trials can be done, but it is a formal statistical exercise with explicit assumptions about the similarity of the trial populations, not a subtraction of one press release from another.
One trial, two legitimate headline numbers
The clearest demonstration that a percentage is not a fixed property of a drug comes from inside a single trial. SURMOUNT-1 reported its 15 mg arm under two different estimands, both pre-specified and both correct answers to different questions.[7] Under the analysis that asks what happened to everyone randomised regardless of whether they kept taking the drug, the reduction was approximately 20.9 percent. Under the analysis that asks what the drug does in people who take it as directed, the figure was approximately 22.5 percent. The gap between them is not error. It is the cost of discontinuation and rescue therapy, made visible.
The ICH E9(R1) addendum formalised this: a treatment effect is only defined once you state how so-called intercurrent events, such as stopping the drug or adding another one, are handled.[10] When a source quotes a single percentage for a GLP-1 medication without naming the estimand, it has dropped the part of the sentence that gives the number meaning. Our guide to primary, secondary, and exploratory endpoints works through the estimand framework in detail.

The randomised comparisons that do exist
Only a handful of trials in this class randomised participants between two active GLP-1 medications. These are the only comparisons that carry the protection of randomisation, and they are correspondingly more informative than any cross-trial arithmetic.
SURPASS-2: tirzepatide against semaglutide in type 2 diabetes
Participants with type 2 diabetes were randomised to tirzepatide 5, 10, or 15 mg or to semaglutide 1 mg, with HbA1c change at 40 weeks as the primary endpoint. The tirzepatide arms produced least-squares mean reductions of 2.01, 2.24, and 2.30 percentage points against 1.86 for semaglutide, with all between-group differences statistically significant.[8] Two constraints matter when reading this. The semaglutide comparator was dosed at 1 mg, which was the highest approved diabetes dose at the time the trial was designed but is not the highest dose in use for weight management. And the endpoint was glycaemic, so weight results are secondary outcomes from a trial powered for HbA1c.
STEP 8: semaglutide against liraglutide in obesity
Adults with overweight or obesity without diabetes were randomised to weekly semaglutide 2.4 mg or daily liraglutide 3.0 mg. Semaglutide produced approximately 15.8 percent mean weight reduction at 68 weeks against approximately 6.4 percent with liraglutide.[9] This is a within-class randomised comparison on a weight endpoint, which makes it one of the more directly usable results in the entire literature.
Beyond these, most product-versus-product claims in circulation are indirect. That does not make them worthless, but it does change what they can support. Absence of a head-to-head trial is not evidence of equivalence, and it is not evidence of difference. It is absence of the comparison.
Surrogate endpoints against hard outcomes
HbA1c and percent body weight are surrogate endpoints. They are validated enough to support approval, but they are measurements of intermediate biology, not of events patients care about directly. The cardiovascular outcomes trials are where the class moves from surrogate to hard endpoint, and this is where the evidence separates sharply between products.
The 2008 FDA guidance requiring cardiovascular safety assessment for new type 2 diabetes drugs is what generated these trials in the first place.[2] Several then went beyond safety and demonstrated benefit:
- LEADER. Liraglutide reduced three-component major adverse cardiovascular events in type 2 diabetes at high cardiovascular risk, hazard ratio 0.87 (95% CI 0.78 to 0.97).[3]
- SUSTAIN-6. Injectable semaglutide reduced the same composite, hazard ratio 0.74 (95% CI 0.58 to 0.95), in a trial designed principally to establish safety rather than superiority.[4]
- SELECT. Semaglutide 2.4 mg reduced the composite in people with obesity and established cardiovascular disease but without diabetes, hazard ratio 0.80 (95% CI 0.72 to 0.90).[5] This extended outcome evidence to a population defined by weight rather than glycaemia.
Two cautions apply to all three. First, the enrolled populations were at elevated cardiovascular risk, so the absolute risk reduction in a lower-risk population would be smaller for the same hazard ratio. Second, and more important for appraisal, a cardiovascular result belongs to the molecule and the population that were tested. It does not transfer to other members of the class by association. The class also includes outcomes trials of lixisenatide, once-weekly exenatide, and dulaglutide, which did not all produce the same answer; anyone treating cardiovascular benefit as a class property should read each primary publication rather than assume. Our guide to reading PubMed covers how to retrieve them.
Tirzepatide is the instructive case for how an outcomes question can be left open. Its dedicated cardiovascular outcomes trial was designed against an active GLP-1 comparator rather than placebo. An active-comparator design can establish that a drug is not worse than an established agent, which is a genuinely useful finding, but it cannot produce the same statement as a placebo-controlled superiority trial. When a claim about cardiovascular benefit is made for a product, the first question is which trial, with which control arm. The ClinicalTrials.gov reading guide explains how to check a registry record for the comparator and the primary outcome measure before the results are published.
A five-question appraisal checklist
Apply these to any specific claim about any product in the class. If a source cannot answer question one, the claim is not yet appraisable.
- What was the comparator? Placebo, an active drug, or nothing at all. An improvement against placebo and an improvement against another GLP-1 medication are different findings with different weight.
- Who was enrolled? Check baseline weight or HbA1c, BMI entry criteria, whether diabetes was present, and the exclusions. A result in a trial that excluded people with recent cardiovascular events says little about people with them.
- Which endpoint was primary? Everything else in the paper is secondary or exploratory and was not the basis for the statistical design. A weight figure quoted from a trial powered for HbA1c is a secondary result.
- How long, and what happened afterwards? Most pivotal trials in this class ran 40 to 72 weeks. Effects in the withdrawal trials reverse substantially once treatment stops, so duration of the trial is not duration of the benefit.
- Who funded and analysed it? Manufacturer funding is standard for phase 3 and is not disqualifying. It is a reason to check whether the protocol and statistical analysis plan were registered before the data were unblinded, and whether the pre-specified endpoints match the reported ones.
For the broader methodology behind questions three and five, see bias, blinding, and why double-blind phase III matters.
Where the evidence is genuinely thin
Naming the gaps is part of describing the evidence base honestly. As of the last review date, the following are not well answered by the published trial literature:
- Long-horizon data. The longest randomised follow-up in the weight indication is measured in years, not decades. Chronic use across a working lifetime has not been studied under randomisation.
- Most product pairings. The randomised comparisons are few. Nearly every ranking that circulates online rests on indirect comparison.
- Body composition. Percent body weight is the regulatory endpoint; the proportion of that loss which is fat mass against lean mass is measured in substudies with smaller samples, and is not the basis of any approval.
- Discontinuation strategy. Trials establish that weight is regained after stopping. What they do not establish is any validated protocol for tapering, intermittent dosing, or maintenance at reduced dose.
- Compounded and unapproved preparations. These have no pivotal trial behind them at all. A trial result generated with an approved finished product does not describe a different preparation of nominally the same molecule, because concentration, excipients, and purity are all part of what was tested.
What this means if a prescription is being considered
Reading the evidence well does not make anyone their own prescriber. The trial data describe average effects in enrolled populations under protocol conditions, and the decision for any individual involves comorbidities, contraindications, interacting medicines, tolerability, and monitoring that no article can assess. The class carries a boxed warning relating to thyroid C-cell tumours and specific contraindications, and gastrointestinal adverse effects are the dominant reason for discontinuation across the programmes. Those belong in a conversation with a prescribing clinician, who is also the right person to weigh which trial population a given individual most resembles.
What appraisal skill does provide is the ability to tell a supported claim from an unsupported one before that conversation, and to arrive with the right question. For the molecule-level detail behind the products discussed here, see the hub pages for semaglutide, tirzepatide, and liraglutide.
Limitations of the evidence
This guide is about appraising evidence, not about selecting a drug. Approval status, licensed indications, and labelling differ by regulator and change as new indications are granted; the regulatory statements here reflect the position at the last review date and should be checked against current prescribing information. Trial results quoted are from the primary publications named in the references and may differ from later meta-analyses, extension data, or longer follow-up reports. Every pivotal trial cited was funded by the manufacturer of the drug under test, which is normal for phase 3 development but relevant to interpretation. The absence of a head-to-head trial between two products is not evidence that they perform equivalently, and it is not evidence that they differ; it means the comparison has not been made under randomisation. Nothing here is a recommendation for or against any product for any individual.
References
Citations are annotated with an evidence tier reflecting study design and replication. See Methodology for criteria.
- 1.Drucker DJ. · Mechanisms of Action and Therapeutic Application of Glucagon-like Peptide-1 · Cell Metabolism · 2018PMID 29617641DOI 10.1016/j.cmet.2018.03.001Validated
- 2.U.S. Food and Drug Administration · Guidance for Industry: Diabetes Mellitus, Evaluating Cardiovascular Risk in New Antidiabetic Therapies to Treat Type 2 Diabetes · 2008Validated
- 3.Marso SP, Daniels GH, Brown-Frandsen K, et al. · Liraglutide and Cardiovascular Outcomes in Type 2 Diabetes (LEADER) · New England Journal of Medicine · 2016PMID 27295427DOI 10.1056/NEJMoa1603827NCT01179048Validated
- 4.Marso SP, Bain SC, Consoli A, et al. · Semaglutide and Cardiovascular Outcomes in Patients with Type 2 Diabetes (SUSTAIN-6) · New England Journal of Medicine · 2016PMID 27633186DOI 10.1056/NEJMoa1607141NCT01720446Validated
- 5.Lincoff AM, Brown-Frandsen K, Colhoun HM, et al. · Semaglutide and Cardiovascular Outcomes in Obesity without Diabetes (SELECT) · New England Journal of Medicine · 2023PMID 37952131DOI 10.1056/NEJMoa2307563NCT03574597Validated
- 6.Wilding JPH, Batterham RL, Calanna S, et al. · Once-Weekly Semaglutide in Adults with Overweight or Obesity (STEP 1) · New England Journal of Medicine · 2021PMID 33567185DOI 10.1056/NEJMoa2032183NCT03548935Validated
- 7.Jastreboff AM, Aronne LJ, Ahmad NN, et al. · Tirzepatide Once Weekly for the Treatment of Obesity (SURMOUNT-1) · New England Journal of Medicine · 2022PMID 35658024DOI 10.1056/NEJMoa2206038NCT04184622Validated
- 8.Frías JP, Davies MJ, Rosenstock J, et al. · Tirzepatide versus Semaglutide Once Weekly in Patients with Type 2 Diabetes (SURPASS-2) · New England Journal of Medicine · 2021PMID 34170647DOI 10.1056/NEJMoa2107519NCT03987919Validated
- 9.Rubino DM, Greenway FL, Khalid U, et al. · Effect of Weekly Subcutaneous Semaglutide vs Daily Liraglutide on Body Weight in Adults With Overweight or Obesity Without Diabetes: The STEP 8 Randomized Clinical Trial · JAMA · 2022PMID 34965561DOI 10.1001/jama.2021.23619NCT04074161Validated
- 10.International Council for Harmonisation (ICH) · ICH E9(R1) Addendum on Estimands and Sensitivity Analysis in Clinical Trials · 2019Validated
- 11.U.S. Food and Drug Administration · Guidance for Industry: Developing Products for Weight Management · 2007Validated