Dissertations

How two systematic reviews can ask the same question and reach opposite conclusions

Systematic reviews sit at the top of the evidence pyramid - yet reviews of the same question routinely disagree. A worked example shows how seven ordinary decisions can flip a conclusion.

Scales representing conflicting evidence and conclusions
Image by LoggaWiggler from Pixabay

Imagine you're writing an essay on whether online learning works as well as face-to-face teaching at university. You do the responsible thing and go looking not for individual studies but for systematic reviews - the gold standard, the top of the evidence pyramid, the thing every research-methods lecture has told you to trust. You find two, both recent, both in respectable peer-reviewed journals, both following the same reporting guidelines. One concludes that online learning produces outcomes "not significantly different" from in-person teaching. The other concludes that online delivery is associated with "meaningfully lower attainment and higher withdrawal." Same question. Opposite answers.

Your first instinct will be that one of them must be wrong - badly done, biased, perhaps funded by someone with a stake in the answer. Sometimes that's true. But far more often it isn't, and understanding why is one of the most useful things a student can learn about how evidence actually works. Two systematic reviews can be honest, competent, transparent and PRISMA-compliant and still contradict each other, because a systematic review is not a neutral photograph of "the literature." It is the product of dozens of decisions - about dates, databases, search terms, what counts as a study worth including, how quality is judged and how results are combined - and each decision, made reasonably, nudges the final answer. Stack enough nudges in one direction and you have one conclusion; stack them the other way and you have its opposite.

This article walks through exactly how that happens, using a hypothetical pair of reviews on online learning as a worked example. It then looks at what the research says about how common the problem is, how to read two discordant reviews intelligently, and why all of this makes a systematic review a genuine research method - a thing you do - rather than an unusually long piece of reading.

The problem is real, and it's growing

The phenomenon has a name in the methodological literature: discordant reviews. In 1997, Alejandro Jadad and colleagues published a guide in the Canadian Medical Association Journal for clinicians facing exactly this situation, opening with the observation that "the very tools that have been promoted as arbiters have, at times, confused rather than clarified." Their example was laser therapy for musculoskeletal pain: one systematic review said it worked, another said it didn't, and a health-service manager deciding whether to fund it had no way to choose.

That was when systematic reviews were relatively rare. They no longer are. John Ioannidis's 2016 analysis in The Milbank Quarterly found that annual publication of systematic reviews had grown by over 2,700% between 1991 and 2014, that most topics addressed by meta-analyses of trials now have multiple overlapping reviews - "same-topic meta-analyses may exceed 20 sometimes" - and that 185 separate meta-analyses of antidepressants for depression were published in the eight years to 2014 alone. With that many reviews of the same evidence, disagreement is a statistical certainty.

And the disagreements aren't marginal. When Klaus Linde and Stefan Willich compared systematic reviews of the same questions in complementary medicine, they found seventeen topics that had each been reviewed between two and five times. The number of primary studies included "varied greatly within most topics"; results and conclusions differed frequently; and the single most common reason was neither fraud nor incompetence but different inclusion criteria, which explained the discrepancies in thirteen of the seventeen topics. Their closing sentence should be pinned above every dissertation student's desk: "apparently minor decisions in the review process can have major impact."

A worked example: two reviews of online learning

Let's build the two reviews from our opening paragraph and watch them diverge. Call them Review A and Review B. Both teams are competent academics. Both register a protocol in advance, follow the PRISMA reporting guideline, screen studies in duplicate and publish their full search strategies. Neither has any commercial interest in the answer. Both begin from what looks like the same question: Do university students learn as well online as they do face-to-face?

Here is how seven ordinary decisions pull them apart.

1. Definitions: what counts as "online," and what counts as "learning"?

Before a single database is opened, the question has to be operationalised, and here the reviews part company almost immediately.

Review A defines online learning as courses designed from the outset for online delivery - planned curricula with purpose-built materials, trained instructors and established platforms. It excludes "blended" courses (mixing online and in-person elements) as a different intervention, and it excludes "emergency remote teaching" - the rapid, improvised shift to video calls during the 2020–21 pandemic - on the grounds that a lecture hastily streamed over Zoom is not what anyone means by an online course.

Review B takes a pragmatic view. It defines online learning as any delivery in which the majority of instruction occurs remotely via digital technology, because that is what students actually experienced and what universities actually offered. It includes blended courses where more than half the contact was online, and it includes pandemic-era studies, reasoning that excluding the largest natural experiment in the history of the subject would be perverse.

The outcome definitions diverge too. Review A's primary outcome is assessed attainment - grades or test scores - because that is what most studies measure and what can be pooled statistically. Review B's primary outcomes are course completion and progression, on the grounds that a student who gets the same grade but is twice as likely to drop out has not had an equivalent education; attainment is a secondary outcome.

Notice that neither definition is wrong. Both are defensible, both are stated openly in the protocol - and they are already reviewing different bodies of evidence.

2. Search dates: when the clock stopped

Review A was registered in 2019 and its search window closes in December 2019. Review B searched to December 2022. Three years' difference - but those three years contain hundreds of studies of remote university teaching, produced under conditions of enormous stress, in which students were isolated, anxious, often without adequate equipment, and taught by staff with days rather than months to prepare. Review B's evidence base is dominated by them; Review A's contains none.

Search dates matter far more than readers assume. Kaveh Shojania and colleagues' survival analysis of 100 high-quality systematic reviews found that the median time before new evidence emerged that was significant enough to warrant updating was just 5.5 years; for 23% a signal appeared within two years, and for 7% the evidence had already changed by the time the review was published. In a fast-moving field, a review's search date is not a footnote - it is a statement about which world the review describes.

3. Databases: where you look determines what you find

Review A searches ERIC (the main education database) and Web of Science. Review B searches those plus Scopus, PsycINFO, ProQuest Dissertations & Theses, and the grey-literature repositories of national education agencies, and hand-searches the reference lists of every included study.

The difference is not cosmetic. Wichor Bramer and colleagues' prospective study of 58 systematic reviews found that 16% of all included references were found in only one database; that even the best single database recovered only 86% of relevant studies; and - most strikingly - that an estimated 60% of published systematic reviews fail to retrieve 95% of the relevant literature because they search too few sources. Every database has its own coverage, its own indexing quirks and its own blind spots. Review A will simply not see some of the studies Review B includes - and, as we'll see, the studies that are hardest to find tend not to be a random sample.

4. Search terms: the vocabulary problem

Even within the same database, two teams can retrieve different studies because they type different things. Review A's search string centres on "online learning," "online course," "distance education" and "e-learning." Review B adds "remote teaching," "emergency remote," "virtual classroom," "hybrid," "blended," "flipped," "MOOC," "synchronous," "asynchronous" and "web-based instruction" - terminology that shifted markedly across the two decades of the literature and again during the pandemic.

Review A also restricts to English-language publications; Review B includes Spanish, Portuguese and Mandarin, and had a large body of Chinese-language remote-teaching studies from 2020–21 translated. Language restriction is common, convenient, and - as several methodological studies have shown - a source of systematic distortion in some fields, because positive results from non-English-speaking countries are more likely to be published in English-language journals than null results are.

5. Eligibility rules: which studies are allowed in

This is where Linde and Willich located the biggest source of discordance, and our example is no exception.

Review A admits only randomised controlled trials and rigorous quasi-experiments with a matched control group, a minimum of 50 students per arm, and publication in a peer-reviewed journal. Its logic is impeccable: these are the designs least vulnerable to bias, and the review wants to say something causal.

Review B admits any comparative study with a face-to-face comparison group, with no minimum sample size, and includes doctoral dissertations, institutional evaluations and government reports. Its logic is also impeccable: randomised trials in education are rare and tend to occur in unusually well-resourced settings, so restricting to them describes a privileged corner of the world; and excluding unpublished work invites publication bias - the well-documented tendency for studies with positive or striking results to be published and studies with null or negative results to sit in a drawer.

That last point deserves a number. In 2008, Erick Turner and colleagues compared the published literature on twelve antidepressants with the complete set of trials registered with the US Food and Drug Administration. Of 74 registered trials, 31% were never published; among those published, 94% appeared positive - but the FDA's own analysis of all trials showed only 51% were, and the published literature overstated effect sizes by around a third. A review that includes only published, peer-reviewed studies is not reviewing the evidence. It is reviewing the evidence that someone chose to publish.

The result: Review A ends with 23 studies, mostly from well-funded North American and Australian universities. Review B ends with 187, spanning six continents and every kind of institution. They now overlap in about fifteen studies. In any meaningful sense, they are no longer reviewing the same literature.

6. Quality thresholds: what to do with weak studies

Both teams appraise the quality of their included studies with a standard risk-of-bias tool. Both find that a substantial fraction of studies have serious problems - students self-selecting into online sections, no control for prior attainment, outcomes measured inconsistently. The teams handle this differently.

Review A excludes every study rated "high risk of bias" from the main analysis, on the grounds that combining unreliable results with reliable ones contaminates the pool. That removes eight of its 23 studies. Review B keeps all 187 in the main analysis but runs a sensitivity analysis showing what happens when the high-risk studies are removed, arguing that discarding two-thirds of the evidence base is a stronger distortion than including it with appropriate caveats.

Both approaches appear in methods textbooks. Both are defended in the Cochrane Handbook under different circumstances. And they leave the two teams with radically different evidence bases entering the final step.

7. Synthesis: how the studies are combined - and how the answer is worded

Review A, with fifteen reasonably homogeneous, well-conducted studies all measuring attainment, runs a meta-analysis: it converts every result to a standardised mean difference and pools them with a random-effects model. The pooled estimate is a tiny difference favouring face-to-face teaching, with a confidence interval that comfortably spans zero. Using the GRADE framework for rating certainty, the team judges the evidence "moderate." Conclusion: "Purpose-designed online courses produce assessed learning outcomes not significantly different from face-to-face delivery."

Review B, with 187 studies measuring wildly different things in wildly different contexts, judges statistical pooling inappropriate for its primary outcome and instead uses structured narrative synthesis with vote counting by direction of effect for completion, plus a meta-analysis of the attainment subset. It finds that in 118 of 160 studies reporting completion, online students were less likely to finish; that attainment differences were small overall but substantially larger in the pandemic-era subgroup and in studies of first-year students; and that heterogeneity was very high. It rates the certainty "low." Conclusion: "Online delivery is associated with meaningfully lower completion and, in several populations, lower attainment; effects vary considerably with context."

Read the two conclusions again. They are both accurate descriptions of what each team found. Review A asked whether a well-built online course teaches a self-selected, supported student as well as a lecture theatre does, and found that it roughly does. Review B asked whether the online provision that universities actually delivered, to the students they actually enrolled, produced the same outcomes, and found that it often didn't. A reader who encounters only the two headline conclusions sees a contradiction. A reader who traces the seven decisions sees two different questions answered honestly.

Neither fraud nor incompetence: the garden of forking paths

It's tempting to think that better training or stricter rules would eliminate this. The evidence suggests otherwise. In 2018, Raphael Silberzahn and colleagues gave the same dataset and the same research question to 29 independent teams of analysts - 61 researchers in total - and asked whether football referees were more likely to give red cards to dark-skinned players. The teams used 21 different combinations of statistical controls. Their effect estimates ranged from essentially zero to a near-tripling of the odds. Twenty teams found a significant effect; nine found none. Crucially, neither the analysts' expertise, nor their prior beliefs, nor peer ratings of the quality of their analyses explained the variation. As the authors put it, "significant variation in the results of analyses of complex data may be difficult to avoid, even by experts with honest intentions."

The statisticians Andrew Gelman and Eric Loken call this the garden of forking paths: at each step in an analysis there are several reasonable choices, each path is individually defensible, and the destination depends on the path. A systematic review is a garden with more forks than almost any other method - a team makes decisions at the level of question, definition, dates, sources, terms, eligibility, quality and synthesis before it reaches a single number. Two teams starting from the same gate and walking honestly can emerge in different places, and it is no one's fault.

This does not mean anything goes. Reviews do differ in quality, and a bad review is still bad. It means that discordance alone is not evidence that anyone cheated or bungled - and that the interesting question is never "which review is right?" but "which review answers the question I actually have, and how confident should I be in it?"

How to read two reviews that disagree

Jadad and colleagues built a decision algorithm for exactly this in 1997, and it translates well into a student's checklist. When you find discordant reviews, work through the following, in order:

  1. Are both reviews methodologically sound? Use a recognised appraisal tool (AMSTAR 2 is the current standard) to check for a registered protocol, a comprehensive search, duplicate screening, risk-of-bias assessment and appropriate synthesis. If one review is seriously flawed, prefer the other and stop.
  2. Do they actually ask the same question? Compare populations, interventions, comparators and outcomes - the PICO elements. In our example, Review A's "purpose-designed online courses" and Review B's "any majority-online delivery" are different interventions, and "attainment" and "completion" are different outcomes. If the questions differ, choose the review whose question is closer to yours and stop.
  3. Do they include the same studies? Compare the inclusion lists. Wildly different lists mean the discordance lives in the search and eligibility stages - check dates, databases, language and study-design rules, and ask which set of choices better fits your purpose.
  4. If the studies overlap substantially, look at extraction and synthesis. Did they extract the same outcomes at the same time points? Did they pool with different models? Did one exclude high-risk studies and the other include them? The Jadad guide notes that reviews which explicitly test whether pooling is appropriate are "probably more credible" than those that pool by default.
  5. Read the certainty ratings, not just the conclusions. "Moderate certainty of no difference" and "low certainty of harm" are not as contradictory as they sound; the second is telling you it isn't sure.
  6. Report the discordance rather than hiding it. In an essay or dissertation, the strongest move is not to pick a side silently but to write: "Two recent reviews reach differing conclusions; the difference appears to arise from X and Y, and for the purposes of this question the more relevant evidence is…" That sentence demonstrates exactly the critical judgement examiners are looking for.

Students who have read our article on the citation echo will recognise a familiar danger here: a systematic review's headline conclusion is precisely the kind of confident summary that gets quoted onward without anyone checking what sat behind it. "A systematic review found no difference between online and face-to-face learning" is a sentence that will be repeated in a hundred essays, stripped of the definitions, dates and exclusions that made it true.

Why this makes a systematic review a research method, not a reading exercise

Everything above points to a conclusion that surprises many students when they first encounter it: a systematic review is empirical research. Its "participants" are studies; its "data collection" is the search; its "measurement" is data extraction and risk-of-bias assessment; its "analysis" is the synthesis. Every stage involves decisions that shape the result, which is exactly why the method demands that those decisions be made in advance, written down, justified and reported - so that a reader can see the garden of forking paths the reviewers walked, and judge it.

In practice, that means a proper systematic review involves:

  • A registered protocol (on PROSPERO or a similar registry) specifying the question, eligibility criteria, search strategy, outcomes and planned synthesis before the search begins - so that the criteria can't quietly drift to fit the results.
  • A reproducible search across multiple databases, with the full search strings published, and supplementary methods such as citation chasing and grey-literature searching to counter publication bias.
  • Duplicate, independent screening of titles, abstracts and full texts, with disagreements resolved and recorded, and a PRISMA flow diagram showing exactly how thousands of hits became dozens of included studies.
  • Structured data extraction using piloted forms, and formal risk-of-bias assessment with a validated tool appropriate to the study designs included.
  • A synthesis method chosen to fit the evidence - meta-analysis where studies are similar enough to pool, structured narrative synthesis where they aren't - with heterogeneity examined rather than ignored, and sensitivity analyses showing how robust the result is to the choices made.
  • An explicit certainty rating (GRADE or equivalent) that tells the reader how much to trust the conclusion.
  • Complete reporting against the 27-item PRISMA 2020 checklist.

None of that is quick. Studies of the time required to complete a systematic review estimate an average of around 67 weeks from registration to publication for a team, and an older but still-cited estimate put the labour at over a thousand person-hours. A master's dissertation compresses this into a few months and a single pair of hands, which is exactly why students undertaking one so often underestimate it - and why the difference between a "systematic review" that is genuinely systematic and one that is a literature review wearing a lab coat is so visible to examiners.

It is also a genre almost no one has seen done properly before being asked to do it. A search strategy that can be reproduced, a PRISMA diagram that accounts for every excluded record, a risk-of-bias table, a defensible argument for pooling or not pooling - these are technical artefacts with specific conventions, and the gap between reading about them and producing them is wide. The approach this series has recommended throughout applies with particular force here: study a well-constructed example of the genre before attempting your own. UKEssays' systematic literature review service provides model reviews written by academics trained in evidence-synthesis methods - showing how a protocol is constructed, how a multi-database search is documented, how eligibility decisions are justified and how heterogeneous evidence is synthesised into a conclusion with an honest certainty rating - alongside feedback on your own draft. Used as intended, it lets you see what each fork in the garden looks like when handled properly, before you have to choose your own path. The usual rule holds: the model is a scaffold to learn from, and the review you submit must be your own - not least because, as this article has shown, a systematic review is its decisions, and the examiner will want to hear you defend every one of them.

If you're doing one: seven habits that prevent avoidable discordance

  1. Write the protocol first and register it. Decide your definitions, dates, sources, terms, eligibility rules and synthesis plan before you search. This is the single biggest difference between a systematic review and a reading list.
  2. Search more sources than feels necessary. Bramer's data suggest most reviews under-search; four databases plus supplementary methods is a reasonable minimum, with subject-specific databases added where relevant.
  3. Build the search string with a librarian. Controlled vocabulary, synonyms, truncation and field tags are a specialist skill, and subject librarians possess it.
  4. Justify every eligibility decision in writing - and think about what each one systematically excludes. "Peer-reviewed only" excludes null results. "English only" excludes half the world. "RCTs only" excludes most of education.
  5. Keep a decision log. Every judgement call during screening and extraction goes in a file with a date and a reason. It becomes your methods chapter and your defence in the viva.
  6. Test your conclusion's robustness. Run sensitivity analyses: does the finding survive removing the high-risk studies, the pandemic studies, the largest study? Report what happens either way.
  7. Rate your certainty honestly and word the conclusion to match. "Low-certainty evidence suggests" is not weakness. It is the sentence that separates a researcher from an advocate.

The bottom line

Systematic reviews earned their place at the top of the evidence pyramid by replacing the expert's selective memory with an explicit, reproducible method. But "explicit and reproducible" was never the same as "unique and inevitable." Every review is a path through a garden of reasonable choices - about what counts, when to stop looking, where to look, what to include, what to trust and how to combine - and different honest teams will walk different paths. Review A and Review B in our example are not a fraud and a truth; they are two accurate answers to two questions that only looked identical. The student who understands this stops asking "which review is right?" and starts asking "what did each review decide, and which decisions matter for what I need to know?" - which is, not coincidentally, exactly the question a systematic reviewer must answer about every study they include. Learn to read reviews that way, and you are already halfway to being able to write one.

← Back to the blog