Opening Remarks
Screening Criteria
Your search returned hundreds of records. Most of them do not belong in your review. The question is whether you can say why, in writing, before you decide.
A review earns its claim to be systematic by making every keep-or-discard decision against explicit, written rules, applied consistently, rather than by feel, one abstract at a time; anyone can search a database, but not every search produces a systematic review.
Learning Outcomes
By the end of this unit, you will be able to:
- Write inclusion and exclusion criteria specific enough that a stranger applying them would reach the same decision you did.
- Run two-stage screening with an interrater comparison, and use disagreement to sharpen an ambiguous criterion rather than split remaining work by convenience.
- Match a quality-appraisal approach to your review type, rather than importing a generic checklist built for a different research tradition.
Inclusion and Exclusion Criteria
Question-Derived Criteria
A criterion should follow from your research question, not from what happens to be easy to check.
“Only English-language papers” is a convenience criterion, defensible for resource reasons, but it should be named as a resource constraint, not dressed up as a methodological choice. “Only papers that empirically study the phenomenon in an organizational setting, not consumer settings” is a question-derived criterion, it follows directly from what your research question is actually about.
Write criteria specific enough that a second person, applying them to a paper they haven’t seen before, would make the same inclusion or exclusion decision you would. If two reasonable people could read your criterion and disagree about a given paper, the criterion is not yet specific enough.
Two-Stage Screening
Screening Sequence
- Stage 1: screen titles and abstracts against your criteria. Fast, low-information, err toward inclusion when uncertain, you can still exclude later with more information.
- Stage 2: retrieve and read the full text of everything that survived stage 1, apply the same criteria with the fuller picture a complete paper provides.
The two-stage structure exists because an abstract is not enough information to make a confident exclusion decision, but reading every paper’s full text at the search-result stage is not feasible at the volumes systematic search typically returns. Stage 1 is deliberately permissive: when in genuine doubt, include and let stage 2 make the final call with better information, a paper wrongly excluded at stage 1 never gets a second look.
Interrater Work
At least a sample of your screening should be done independently by two team members, then compared.
Have two people independently screen the same sample of, say, 50 records, without discussing them first. Then compare: where do you agree, and where do you disagree? Treat disagreement as useful information rather than as a failure. It shows you exactly where a criterion is ambiguous. For each disagreement, discuss until you reach consensus, and use what you learn to sharpen a criterion that turned out to be more ambiguous than it looked on paper.
If your independent agreement rate on the sample is low, revise your criteria before you screen the remaining records rather than pushing through and splitting the rest of the work by convenience.
Quality Appraisal
Appraisal Judgment
Checklists built for randomized trials do not transfer cleanly to conceptual or qualitative work.
A great deal of published quality-appraisal guidance assumes you are appraising quantitative, often experimental, studies, randomization, blinding, sample size calculations. If your review, per your SLR I decision, is synthesizing conceptual, qualitative, or design-oriented work, those specific checklist items simply do not apply, and forcing them onto a paper that was never designed to be evaluated that way produces a score with no evaluative meaning.
Boell & Cecez-Kecmanovic (2014) argue for judging a study’s contribution on its own methodological terms: is the argument’s logic coherent, is the evidence, whatever form it takes, appropriate to the claim being made, is the paper’s own stated method actually followed. Match your appraisal approach to your review type rather than importing a generic checklist from a different research tradition.
Worked Example: Appraisal
Two invented papers, appraised side by side against different criteria.
Quantitative
Survey study (N=412), moderated-mediation model of AI transparency and employee trust across three industries.
- Sample size adequate for the model’s power?
- Scales previously validated?
- Statistical assumptions checked?
- Response-bias addressed?
Qualitative
Interview study (24 informants, four organizations), process model of trust repair after an AI failure.
- Sampling purposive and justified?
- Coding process documented, audit trail?
- Claims triangulated across informants or sources?
- Researcher’s own role reflected on?
Both papers above are invented, illustrative only, meant to show how the appraisal question changes shape depending on what kind of paper you are holding.
For the quantitative paper, appraisal centers on statistical adequacy: is the sample large enough to detect the effects the model claims, are the measurement instruments established rather than improvised for this study, are the assumptions behind the statistical test actually met, and has the study addressed the possibility that people who responded differ systematically from people who didn’t.
For the qualitative paper, none of those questions transfer directly. Appraisal instead centers on interpretive rigor: was the sample selected for a defensible theoretical reason rather than convenience, can a reader trace how raw interview text became the reported categories, does the paper triangulate its claims across multiple informants or data sources rather than resting on one account, and does the author acknowledge their own position and its possible influence on the interpretation.
Running both appraisals side by side makes the underlying point concrete: neither list of questions is the “real” quality checklist and the other a lesser substitute. Each fits the kind of evidence its paper offers.
Reporting Transparently
The PRISMA Flow Diagram
One diagram, tracking every record from search to final inclusion.
The PRISMA reporting standard (Page et al., 2021) popularized a specific transparency artefact: a flow diagram showing how many records were identified, how many were removed as duplicates, how many were screened at each stage, how many were excluded at each stage (with reasons), and how many remain in the final synthesis. You do not need to follow PRISMA’s full checklist, it was built primarily for health-sciences systematic reviews and meta-analyses, but the flow diagram itself is a genre-independent, extremely effective transparency device, and you should produce one regardless of your specific review type.
Transparency Boundary
Transparency lets a reader verify that your conclusions actually follow from your search and screening process.
Templier & Paré (2018) found substantial variation in how consistently top IS journals report their review methods, search strategies, criteria, and screening numbers, some reviews report enough to reconstruct the process, many do not, showing that peer-reviewed status alone does not guarantee a review’s process was transparent. Because you are held to a standard that a meaningful share of published work does not fully meet, the reproducibility test in Peer-Review Workshop II requires that someone else can actually verify your process, not merely take your word for it.
This also connects directly back to replicability, the first of the four principles from Knowledge, Science and Epistemology: a screening process that cannot be reconstructed from what you wrote down fails the same test a non-replicable experiment would.
Homework
Produce your screening artefacts.
Work on your final inclusion and exclusion criteria, your interrater comparison results and how disagreements were resolved, your quality-appraisal approach (and why it fits your review type), and a PRISMA-style flow diagram of your full screening process.
Upload your team’s current search protocol, databases, search strings, and inclusion and exclusion criteria, to Moodle by Mon, 07.12, so the team reproducing it at Peer-Review Workshop II has something to execute. Reproduction pairings, who reproduces whom, will be published on Moodle before the session.