A claim that an app has been “tested in a trial” leaves several questions unanswered. Who designed the comparison? Who handled the data? Who decided which outcome mattered? Independence is one part of that picture, alongside the quality and transparency of the study itself.
- Publication
- npj Digital Medicine · 17 September 2026
- Version
- Peer-reviewed accepted article; final editing pending
Original publication: 17 Sep 2026 · The date above refers to this brief.
Developer involvement in the reviewed trials
Share of trials (%)
What the review found
Zhou and colleagues examined 229 trials from 29 systematic reviews covering nutrition, maternal health, mental health and sleep. They report developer involvement in 73% of trials and independent conduct in 27%. In a sample-size-weighted analysis, developer-involved trials had higher odds of statistically significant results: odds ratio 1.23 (95% confidence interval 1.16–1.31). They were also more likely to be preregistered (odds ratio 2.47). This is an association across studies, not random assignment of developer involvement.
Source 1 ↗The number is not a measure of patient benefit
“Statistically significant” describes a result under an analysis, not the size or importance of a health improvement. An odds ratio about the reporting of significant results is therefore not a percentage improvement in the people using an app.
Our suggested reading exercise is to rewrite any numerical claim as a sentence with a subject and an outcome. For example: “Trials in one category had different odds of reporting significance.” If the headline instead says “users improved by that amount,” the subject has silently changed. Keep those two statements separate.
Independence and quality answer different questions
Developer involvement alone does not establish that a result is wrong. Independent conduct alone does not establish that a result is reliable. A reader still needs a suitable comparison, a prespecified outcome, adequate follow-up and a clear account of missing data.
For a hypothetical app study, imagine that one team compares its product with no additional support, while another compares it with an established active service. A difference in the resulting headlines could reflect the comparison, even before considering who funded the work. That example illustrates why an observational comparison cannot isolate every possible explanation.
How to read the disclosure and registration together
Our worksheet below treats disclosure as a map of responsibilities. Look for who designed the study, who could access the data, who performed the analysis, and whether investigators could publish an unfavorable result. A statement of funding is useful, but it may not answer all four questions.
Then follow the registration link. Compare the planned primary outcome, time point and analysis with the report. Preregistration gives you something concrete to compare; the presence of a registration number alone cannot demonstrate that the plan was followed. Our trial-registration guide walks through that comparison.
A constructive question for the next generation of studies
Our proposed follow-up would compare clearly defined forms of developer involvement, while making trial size, comparator, outcomes and reporting practices visible. An independent team could repeat an evaluation of a fixed software version with the same prespecified outcomes.
The goal would be to learn which findings survive a new test and why estimates differ. This review does not supply a ranking of specific apps. Readers can use it to demand an inspectable evidence trail, while leaving product-level judgments to evidence about the actual product and intended use.
Four responsibilities to trace
| Question | Where to look | Why it matters |
|---|---|---|
| Who chose the comparison? | Protocol and methods | The alternative defines the question answered |
| Who controlled the data? | Data-access and sponsor-role statements | Others need a way to check the analysis |
| What was planned in advance? | Dated registration and analysis plan | Later changes should be identifiable |
| Who could publish the result? | Sponsor-role and publication-rights statements | An unfavorable finding must be reportable |
Our reading worksheet. It is not an assessment of every trial in the review.
Your questions, answered
Does developer involvement prove that a trial is biased?
No. It identifies a relationship worth examining. Assess what people actually controlled, how the trial was designed, and whether the reporting supports the claim.
Is a registered trial automatically trustworthy?
Registration supports scrutiny. Its value depends on timing, detail, and whether the reported outcomes and analyses can be compared with the original plan.
Does statistical significance tell me whether an app is useful?
Not on its own. Look for the size and uncertainty of the effect, the outcome measured, the comparison and the burdens of using the product.
How does this connect to the chatbot articles?
They ask a complementary question: what did the evaluation measure? Combine that with who designed and analyzed it. Neither a well-known developer nor an impressive benchmark replaces an appropriate test.
Limits of this interpretation
- This source check covers the accepted-article abstract, not the complete methods or trial-level records.
- Associations between study characteristics do not by themselves explain their causes.
- No app-specific clinical recommendation or pooled treatment-effect estimate is provided.
Sources & transparency
- Zhou, Kohli and Tiemeier (2026): Study independence and the developer effect
Publisher-indexed abstract and publication-status notice checked. Peer-reviewed accepted article, released before the final Version of Record. Direct full-page retrieval failed; full methods, supplementary files and individual trials were not independently assessed. We use a short factual report and an original chart of reported numbers, not a reproduction of a journal figure. · Accessed 29 Sep 2026
DOI: 10.1038/s41746-026-03234-9
Prepared and source-checked with AI on 29 September 2026. Press-news Team is our collective publication byline, not a claim of medical credentials or human review. No human editorial or clinical review has taken place. This is educational reporting on research methods, not medical advice. Access limitations are listed with each source. Reading frameworks and proposed follow-up tests are our commentary, not additional study findings. The photograph is illustrative.
Source check: AI source check — selected primary research and reported numbers
Clinical review: Not applicable to this educational guide
Suggest a correction