Article
Every sentence true
A trial write-up can be accurate in every particular and still leave an expert reader with the wrong impression. This is not a theory: someone ran it as a randomised trial, on 300 clinicians, using real abstracts and their own de-spun rewrites. It is the one failure a citation check cannot catch.
spin in trial reporting, and the failure mode no citation check can catch · 2026-08-06 · 8 sources
Here is the position a reader is actually in.
You are handed the abstract of a randomised trial. It runs about 300 words. Every sentence in it is true: the numbers are right, the statistics are correctly computed, nothing is invented, and if you looked up every claim it would check out. You have perhaps four minutes. Decide what you think of the drug.
The uncomfortable finding of the last twenty years is that this is not enough, and not because the abstract is too short. A set of true sentences can be arranged to leave a reader with an impression the data do not support, and the arrangement leaves no trace in any individual sentence.
The literature has a name for the arrangement. It calls it spin.
What spin is, and where it collects
The definition worth using is from the study that first counted it: reporting strategies, whatever their motive, that emphasise the experimental treatment as beneficial despite a statistically nonsignificant result on the primary outcome, or that distract the reader from that result.1Boutron, I., Dutton, S., Ravaud, P., & Altman, D. G. (2010). "Reporting and interpretation of randomized controlled trials with statistically nonsignificant results for primary outcomes." JAMA 303(20), 2058-2064. DOI 10.1001/jama.2010.651. PMID 20501928. 616 reports screened, 72 appraised; spin in the title of 13 (18.0%; 95% CI 10.0-28.9%), abstract Results of 27 (37.5%), abstract Conclusions of 42 (58.3%; 95% CI 46.1-69.8%).Primary Note what the definition does not contain. No claim of dishonesty, no allegation of error. Spin is a property of emphasis.
That 2010 survey screened 616 published trial reports and appraised the 72 with a clearly identified primary outcome that came out nonsignificant. Spin appeared in the title of 13 (18.0%; 95% CI 10.0-28.9%), in the abstract's Results section in 27 (37.5%), and in the abstract's Conclusions in 42 -- 58.3% (95% CI 46.1-69.8%).
Read that against the population it is drawn from. These are trials that did not meet their primary endpoint. In more than half, the conclusion a hurried reader is most likely to reach -- and often the only part they will read -- was written to emphasise benefit anyway.
The optimism is not spread evenly across the literature. It collects where the result was disappointing.
It travels, and it grows on the way
For anyone whose working diet is company announcements rather than journals, the next finding is the one that matters.
A 2012 cohort study took every press release indexed in EurekAlert! over four months, kept the 70 that described two-arm randomised trials, retrieved the underlying paper for each, and then tracked the resulting news coverage.2Yavchitz, A., Boutron, I., Bafeta, A., Marroun, I., Charles, P., Mantz, J., & Ravaud, P. (2012). "Misrepresentation of randomized controlled trials in press releases and news coverage: a cohort study." PLoS Medicine 9(9), e1001308. DOI 10.1371/journal.pmed.1001308. PMID 22984354. 498 press releases screened, 70 two-arm RCTs included; spin in 28 abstract conclusions (40%) and 33 press releases (47%); the only factor associated with spin in the press release was spin in the abstract conclusion (RR 5.6; 95% CI 2.8-11.1; P < 0.001); findings overestimated in 19 reports (27%).Primary Spin appeared in 28 of the abstract conclusions (40%) and in 33 of the press releases (47%).
They then tested what predicted spin in a press release, against journal type, funding source, sample size, treatment type and the significance of the primary outcome. One factor survived: spin in the article's own abstract conclusion, with a relative risk of 5.6 (95% CI 2.8-11.1; P < 0.001). And the findings of 19 reports (27%) were overestimated relative to the trial when read through the press release.
So this is not a self-contained problem inside the academic literature. It is a chain, and each link is further from the data than the last:
primary result -> abstract conclusion -> press release -> news
fixed 40% spun 47% spun 27% overestimated
The reader at the far end is making a decision about a drug from a document two removes from the trial, in which nothing false has been written at any step.
The part that makes this different: somebody ran the experiment
Surveys establish that spin exists and is common. They cannot establish that it works, because trials that get spun differ from those that do not in every other respect too.
In 2014 a group including Isabelle Boutron, Douglas Altman and Ian Tannock closed that gap in the only way it closes. They ran a randomised controlled trial on spin itself.3Boutron, I., Altman, D. G., Hopewell, S., Vera-Badillo, F., Tannock, I., & Ravaud, P. (2014). "Impact of spin in the abstracts of articles reporting results of randomized controlled trials in the field of cancer: the SPIIN randomized controlled trial." Journal of Clinical Oncology 32(36), 4120-4126. DOI 10.1200/JCO.2014.56.7503. PMID 25403215. Two-arm parallel-group RCT; 300 clinicians randomised, 150 per arm, blinded to the hypothesis. Spun version: treatment rated more beneficial (mean difference 0.71; 95% CI 0.07-1.35; P = .030), trial rated less rigorous (mean difference -0.59; P = .034), greater interest in the full text (mean difference 0.77; P = .029).Primary
The design deserves care, because its elegance is the whole argument. They took real published cancer trials with a nonsignificant primary outcome and spin in the abstract conclusion. For each, they produced a second version: same trial, same data, same results, rewritten without the spin. Then they randomised 300 clinicians -- corresponding authors of trial reports, trial investigators, grant reviewers -- to read one version or the other, blinded to the hypothesis. The only difference between arms was the prose.
Figure 1. Why the comparison is clean. Everything upstream of the write-up is held
fixed -- the same trials, the same nonsignificant primary outcomes, the same numbers.
Only the arrangement of true sentences differs, and readers are randomised between the
two versions. Any difference in what they conclude is therefore attributable to the
prose and to nothing else. Drawn by fig-1.py.
Clinicians reading the spun version rated the experimental treatment as more beneficial: mean difference 0.71 on a 0-10 scale (95% CI 0.07-1.35; P = .030). They were also more interested in reading the full article (mean difference 0.77; P = .029).
A modest effect, and worth reporting as one. But consider who moved. Not students, not the public. These were clinicians who write trial reports and review grants for a living, whose professional competence is precisely the reading of documents like this. Expertise did not neutralise it. There is no reason to think seniority is the missing ingredient -- the same conclusion the fabrication literature reached by a different road.
One result from the same study deserves more attention than it gets. Clinicians reading the spun version also rated the trial as less rigorous (mean difference -0.59; P = .034). Something in the writing registered as off. Their assessment of the methods moved -- and their assessment of the treatment moved favourably anyway. Noticing that a document is working on you turns out not to be the same as being unaffected by it.
Sixteen years on, in myeloma
The obvious question is whether any of this survived the intervening decade of reform -- CONSORT, trial registration, structured abstracts, the whole apparatus.
A cross-sectional analysis in The Oncologist in June 2026 looked at myeloma specifically: 71 randomised trials that began enrolling in 2015 or later.4Singstock, M., Abdul Khader, A. H. S., Wayant, C., Lemieux, M., Mainou, M., & Mohyuddin, G. R. (2026). "Reporting of primary endpoint and associated spin in randomized myeloma trials." The Oncologist, published 6 June 2026. DOI 10.1093/oncolo/oyag221. PMID 42231132. Cross-sectional analysis of myeloma RCTs enrolling from January 2015; final search March 2025; methods registered on the Open Science Framework. 82 screened, 71 included; primary endpoint significant in 46 (64.8%); sample size calculations reported in 51 (71.8%); conclusions aligned with the primary endpoint in 63 (88.7%); spin in 25 (35.2%); meeting the prespecified endpoint associated with lower odds of spin (adjusted OR 0.15; 95% CI 0.03-0.70; P = .016).Primary The prespecified primary endpoint was statistically significant in 46 (64.8%). Two reviewers screened the reporting for spin and found it in 25 trials -- 35.2%, about one in three.
Then the finding that reorganises everything above. In multivariable analysis, meeting the prespecified primary endpoint was associated with substantially lower odds of spin: adjusted odds ratio 0.15 (95% CI 0.03-0.70; P = .016).
Inverted, which is the form worth carrying:
missing the primary endpoint is the strongest available predictor
that the write-up will read optimistically
The paper makes no claim about anyone's motives and neither does this. It is a statement about where the optimism collects, and it points the same way as the 2010 data. One more figure from the same analysis sharpens it: conclusions aligned with the primary endpoint result in 63 of 71 trials (88.7%) -- so roughly one trial in nine reached a conclusion its own prespecified primary result did not support.
Nor is this an oncology problem. Two other 2026 analyses, in fields with no relation to myeloma or to each other, found spin in 40.2% of randomised trials of digital implant surgery and in 88.4% of pilot and feasibility trials in hip and knee arthroplasty.5"'Spin' in randomized controlled trials of digital implant surgery: a meta-epidemiologic study." Evidence-Based Dentistry, 12 June 2026. PMID 42286167. Spin identified in 51 abstracts (40.2%).Primary 6"Spin reporting is common in pilot and feasibility trials in hip and knee arthroplasty: a methodological analysis." Pilot and Feasibility Studies, 30 May 2026. PMID 42216065. Spin appeared in 88.4% of studies.Primary The regularity is not a specialty's culture. It appears wherever a disappointing result meets someone who has to write it up.
And it is not obscure. The 2010 paper has been cited 826 times and the 2014 experiment
362 times; the surrounding meta-research literature runs to tens of thousands of
works.7Citation counts and corpus magnitude retrieved from OpenAlex at authoring time (6 August 2026) via openalex_get_work_by_doi and openalex_search_works: 826 citations for the 2010 JAMA report and 362 for the 2014 SPIIN trial. The corpus figure is a full-text search total, reported as a magnitude only -- it is not a systematic count and should not be read as one. (Primary, transport; the underlying authority is each indexed work.) This is among the best-documented problems in clinical publishing,
and the 2026 rate is 35.2%. Whatever the remedy is, being widely known about is not it.
Why nothing catches this
Now the part that matters for anyone who audits documents for a living.
Take a spun abstract and run every available check. Does each cited study exist? Yes. Does each identifier resolve? Yes. Is each number accurately transcribed? Yes. Are the statistics correctly computed? Yes. Does every claim trace to a document a sceptic can open? Yes.
The document passes. All of it passes, because none of it is wrong.
This is the third and hardest member of a family, and the three are worth seeing together:
| failure | what is wrong | what catches it |
|---|---|---|
| fabrication | the claim is false | the identifier does not resolve |
| omission | the claim is absent | nothing -- absence has no shape |
| spin | nothing. every claim is true | nothing -- emphasis has no locator |
Each is invisible to the check that catches the one before it. A citation gate was built for the first; it says nothing about the second, and it cannot in principle address the third -- because spin is not a property of any sentence. It is a property of which true sentences were selected, and in what order. There is no field to validate, no identifier to resolve, no quote to match. The unit of the error is the document.
That is a real limit on provenance, and better stated plainly than worked around. A perfectly sourced report can be a misleading report.
The one thing that does not move
There is a fixed point, and the 2026 analysis uses it without quite naming it as the remedy: the trial's own registry entry.
Take a myeloma trial currently enrolling, NCT06464991.8ClinicalTrials.gov, NCT06464991 (FUMANBA-03), retrieved 6 August 2026 via search_clinical_trials and get_clinical_trial_outcome_measures. Registered primary outcome measure and description quoted verbatim; overall status RECRUITING, start date 2024-03-27, primary completion date 2027-08. Cited as an illustration of what a prespecified endpoint looks like on the public record before a result exists -- not as evidence about this trial or its sponsor.Primary Its registered primary
outcome measure reads:
Progression-Free Survival (PFS) as assessed by Independent Review Committee (IRC) -- the time from randomization to the first documented disease progression as determined by IRC or death due to any cause. Time frame: up to 5 years from randomization.
That trial started on 27 March 2024 and its primary completion date is August 2027. The endpoint above is on the public record now, and the answer does not exist yet. It was chosen before anyone could know whether it would be met, which is exactly what makes it useful: it cannot have been selected to flatter a result nobody has seen. It is used here for that reason and no other -- the trial has not reported, so no judgment about it is possible or implied.
That is the property spin cannot touch. Emphasis is chosen after the result; the registered endpoint was fixed before it.
What this has to do with a diligence report
Which says what a read has to add once the sourcing is done.
The 2026 analysis contains its own remedy, in its methods rather than its conclusions. It asked of every trial what the prespecified primary endpoint was, whether that endpoint was met, and whether the stated conclusion aligned with the answer. Three questions. All answerable from public documents. All independent of how the abstract is worded -- which is precisely why they survive spin.
So the useful thing a report carries is not a better-calibrated adjective about a trial's promise. It is that comparison, made explicit and made every time: here is what was prespecified, here is what happened, here is whether the conclusion follows. A reader who has that does not need to trust the write-up's tone, which is the only defence against tone that has ever worked.
None of which makes provenance less necessary. It makes it insufficient -- a different claim, and a more useful one. Grading sources rules out the first two failures in the table. The third was never a sourcing problem, and a practice that treats a complete bibliography as the end of the work will keep passing documents in which every sentence is true.
The alternative to saying this is to imply a guarantee nobody can give. What can be given is narrower: the comparison the spun abstract leaves out, on every claim, every time, in a form you can check without taking anyone's word for the tone.
Appendix A -- provenance and method
How every number above was pulled, graded, and made checkable -- not asserted.
Every figure in this piece was read from a primary record at authoring time, through the ToolUniverse CLI (v1.1.11) against PubMed, OpenAlex and ClinicalTrials.gov. The site is a static export, so this is an authoring-time act, not a live fetch: what follows is the audit trail, reproducible, so a sceptic can re-run it and land on the same numbers.
| Step | Tool / API call | What it verified |
|---|---|---|
| Literature discovery | tu run PubMed_search_articles x 4 (spin / reporting / trial-registry terms) | What the meta-research record contains, including the three 2026 studies. Deliberately broad -- an over-specified first query returned zero. |
| Abstract capture | tu run PubMed_get_article x 6 (PMIDs 20501928, 22984354, 25403215, 42231132, 42216065, 42286167) | Every quoted figure, interval and P value read from the abstract of record rather than a summary of it. |
| Citation reach | tu run openalex_get_work_by_doi x 2 (10.1001/jama.2010.651, 10.1200/JCO.2014.56.7503) | 826 and 362 citations respectively -- the basis for the claim that this problem is well known rather than obscure. |
| Literature size | tu run openalex_search_works (spin / RCT / abstracts) | The surrounding meta-research corpus, tens of thousands of works. Reported as a magnitude, never as a precise count -- a full-text search total is not a census. |
| Breadth check | The two non-oncology 2026 studies above | That the pattern is not a myeloma artefact: 40.2% and 88.4% in unrelated specialties. |
| The fixed point | tu run search_clinical_trials -> get_clinical_trial_outcome_measures -> get_clinical_trial_status_and_dates (NCT06464991) | The registered primary endpoint quoted verbatim, plus start date 2024-03-27 and primary completion 2027-08 -- establishing that the endpoint is on record while the result does not yet exist. |
Two things this appendix deliberately does not claim. The OpenAlex total is a search result, not a systematic count, and is described that way in the text. And the ClinicalTrials.gov entry is an illustration of a mechanism, not evidence about that trial or its sponsor -- it was chosen because it has not reported.
Appendix B -- reasoning, run four ways
How the conclusion was reached, run four ways -- so you can find the seam if there is one.
This piece rests on one claim: a document can be true in every sentence and still mislead, so grading sources is necessary and not sufficient. Feynman's first rule of honest thinking is that you must not fool yourself, and you are the easiest person to fool. The cure is independence -- if several routes that do not lean on each other reach the same place, the odds you fooled yourself on every one grow small. Four routes, and where each lands:
| Route | Where it lands |
|---|---|
| Follow the evidence | Spin is common, it travels, and it moves experts |
| Assume it, then check | All three conditions hold; the claim refuses to break |
| The recurring pattern | The same rate appears in fields with nothing in common |
| From a basic principle | No per-claim check can detect a between-claim property |
Follow the evidence
Start from the record and trace it forward.
Just follow the trail. In trials that missed their primary endpoint, 58.3% of abstract conclusions were written to emphasise benefit anyway. Those conclusions predict spin in the press release at a relative risk of 5.6, and press-release readers overestimated the finding in 27% of cases. Randomise clinicians onto spun versus de-spun versions of the same trial and the spun readers rate the treatment higher. In current myeloma trials the rate is 35.2%. The trail ends somewhere uncomfortable: every document in that chain can be factually accurate, so accuracy is not the property that was protecting anyone.
Assume it, then check
Suppose the claim is true; confirm each condition it would require (working backward).
Try to break it rather than build it. If "true documents can mislead, so sourcing is insufficient" were true, three things would have to hold. Spin would have to be common -- it is, at 35.2-88.4% depending on the field. It would have to actually change expert judgment, not merely annoy people -- it does, demonstrated by randomisation rather than inferred. And no citation-level check could catch it -- none can, because every claim in a spun document resolves to its source. Each condition is independently documented. The claim refuses to break.
The recurring pattern
One regularity across many independent cases (induction).
Step back and watch the same thing recur. Myeloma randomised trials: 35.2%. Digital implant surgery: 40.2%. Pilot trials in hip and knee arthroplasty: 88.4%. Nonsignificant trials across all fields in 2010: 58.3% of abstract conclusions. These specialties share no journals, no investigators, no professional culture and no funding structure. The common factor is not a field. It is the situation -- a disappointing result meeting someone who has to write it up -- and the rate tracks the situation, not the discipline. (This is induction: a strong pattern, not a law, so a field with genuinely different incentives could defy it. None of the ones measured has.)
From a basic principle
From a premise almost no one rejects, the conclusion follows of necessity (deduction).
Finally, reason it out. A citation check validates each claim against the source it cites -- that is what it is and all it does; nobody disputes this. Spin is not a property of any claim. It is a property of which true claims were selected, and in what order, relative to the ones that were available. It therefore exists only between claims, never within one. A procedure that examines claims one at a time cannot detect a property that no single claim has. This is not a limitation of current tools that a better tool would fix. It is a structural consequence of what per-claim checking is, and it holds however good the checker becomes.
Four roads, one destination. Any single road you might doubt -- perhaps I selected the evidence, perhaps the pattern is a coincidence of the fields that happened to be studied. But the fourth road needs no evidence at all: it follows from the definition of the check. That is the strongest of the four, and it is why the conclusion here is a concession rather than a boast -- sourcing is necessary, it is not sufficient, and no amount of it will become sufficient.
Footnotes
-
Boutron, I., Dutton, S., Ravaud, P., & Altman, D. G. (2010). "Reporting and interpretation of randomized controlled trials with statistically nonsignificant results for primary outcomes." JAMA 303(20), 2058-2064. DOI 10.1001/jama.2010.651. PMID 20501928. 616 reports screened, 72 appraised; spin in the title of 13 (18.0%; 95% CI 10.0-28.9%), abstract Results of 27 (37.5%), abstract Conclusions of 42 (58.3%; 95% CI 46.1-69.8%). (Primary) ↩
-
Yavchitz, A., Boutron, I., Bafeta, A., Marroun, I., Charles, P., Mantz, J., & Ravaud, P. (2012). "Misrepresentation of randomized controlled trials in press releases and news coverage: a cohort study." PLoS Medicine 9(9), e1001308. DOI 10.1371/journal.pmed.1001308. PMID 22984354. 498 press releases screened, 70 two-arm RCTs included; spin in 28 abstract conclusions (40%) and 33 press releases (47%); the only factor associated with spin in the press release was spin in the abstract conclusion (RR 5.6; 95% CI 2.8-11.1; P < 0.001); findings overestimated in 19 reports (27%). (Primary) ↩
-
Boutron, I., Altman, D. G., Hopewell, S., Vera-Badillo, F., Tannock, I., & Ravaud, P. (2014). "Impact of spin in the abstracts of articles reporting results of randomized controlled trials in the field of cancer: the SPIIN randomized controlled trial." Journal of Clinical Oncology 32(36), 4120-4126. DOI 10.1200/JCO.2014.56.7503. PMID 25403215. Two-arm parallel-group RCT; 300 clinicians randomised, 150 per arm, blinded to the hypothesis. Spun version: treatment rated more beneficial (mean difference 0.71; 95% CI 0.07-1.35; P = .030), trial rated less rigorous (mean difference -0.59; P = .034), greater interest in the full text (mean difference 0.77; P = .029). (Primary) ↩
-
Singstock, M., Abdul Khader, A. H. S., Wayant, C., Lemieux, M., Mainou, M., & Mohyuddin, G. R. (2026). "Reporting of primary endpoint and associated spin in randomized myeloma trials." The Oncologist, published 6 June 2026. DOI 10.1093/oncolo/oyag221. PMID 42231132. Cross-sectional analysis of myeloma RCTs enrolling from January 2015; final search March 2025; methods registered on the Open Science Framework. 82 screened, 71 included; primary endpoint significant in 46 (64.8%); sample size calculations reported in 51 (71.8%); conclusions aligned with the primary endpoint in 63 (88.7%); spin in 25 (35.2%); meeting the prespecified endpoint associated with lower odds of spin (adjusted OR 0.15; 95% CI 0.03-0.70; P = .016). (Primary) ↩
-
"'Spin' in randomized controlled trials of digital implant surgery: a meta-epidemiologic study." Evidence-Based Dentistry, 12 June 2026. PMID 42286167. Spin identified in 51 abstracts (40.2%). (Primary) ↩
-
"Spin reporting is common in pilot and feasibility trials in hip and knee arthroplasty: a methodological analysis." Pilot and Feasibility Studies, 30 May 2026. PMID 42216065. Spin appeared in 88.4% of studies. (Primary) ↩
-
Citation counts and corpus magnitude retrieved from OpenAlex at authoring time (6 August 2026) via
openalex_get_work_by_doiandopenalex_search_works: 826 citations for the 2010 JAMA report and 362 for the 2014 SPIIN trial. The corpus figure is a full-text search total, reported as a magnitude only -- it is not a systematic count and should not be read as one. (Primary, transport; the underlying authority is each indexed work.) ↩ -
ClinicalTrials.gov, NCT06464991 (FUMANBA-03), retrieved 6 August 2026 via
search_clinical_trialsandget_clinical_trial_outcome_measures. Registered primary outcome measure and description quoted verbatim; overall status RECRUITING, start date 2024-03-27, primary completion date 2027-08. Cited as an illustration of what a prespecified endpoint looks like on the public record before a result exists -- not as evidence about this trial or its sponsor. (Primary) ↩