Opening Remarks
Search Reproducibility
Today’s task: execute another team’s search protocol, and get feedback to yours.
Knowledge, Science and Epistemology named replicability as one of four principles a claim needs to count as scientific knowledge: the extent to which a procedure is documented well enough that someone outside the team could repeat it and obtain similar results. Today you find out, concretely, whether your own search protocol clears that bar, by handing it to a team that was not in the room when you wrote it.
Learning Outcomes
By the end of this unit, you will be able to:
- Execute another team’s documented search protocol exactly as written, and diagnose where and why your resulting paper set diverges from theirs.
- Distinguish four sources of divergence, ambiguous terms, undocumented filters, database differences, silent judgement calls, all of which trace back to what was not written down.
- Write a structured, transparency-focused review of a search protocol, and revise your own protocol against every gap a reproducing team identified.
The Reproducibility Test
Reproduction Task
- Download another team’s documented search protocol: databases, search strings, inclusion and exclusion criteria.
- Execute it, exactly as written, no clarifying questions to the original team.
- Compare your resulting set of papers against theirs.
- Where the sets diverge, diagnose why.
The instruction “exactly as written, no clarifying questions” mirrors a real review, where nobody can ask the original authors what they meant, so an unwritten clarification your protocol needs to be repeatable is exactly the gap this exercise is designed to surface.
Divergence Sources
Divergence is the learning object of this exercise. It does not mean something went wrong.
- Ambiguous terms: a search term that could reasonably be interpreted two ways, and was.
- Undocumented filters: a date range, language restriction, or document type filter that was applied but never written down.
- Database version differences: the same database, searched on a different day, with a different index.
- Silent judgement calls during screening: an inclusion or exclusion decision that made sense to the original team in the moment, but was never turned into a written rule.
Three of these four causes stem from what was not written down, not from anything wrong with the search itself, which is the practical meaning of transparency as the boundary between a review and an opinion: your reasoning must be inspectable by someone who was not there, not just correct.
Reviewing the Protocol
Recap: Reviewer Craft
Apply the same rubric-driven, constructive review from Peer-Review Workshop I, just to a different artefact.
The general craft, be polite and specific, search for the gem before you search for flaws, be concise, be consistent between your comments and your verdict, still applies. What changes is the object under review: a protocol’s reproducibility rather than a research question’s argument chain.
Example Review
How specific issues raised in a review could look like in practice:
- Databases and search strings: “You state Scopus and Web of Science, and the search string in your protocol reads
('AI agent' OR 'autonomous agent') AND trust. Your appendix search log shows an addedAND organizationalfilter that never made it into the protocol text. Add that filter to the written string itself.” - Inclusion and exclusion criteria: “Your criterion ‘sufficiently related to the topic’ is not a rule a stranger could apply. Working from the protocol alone, I could not tell why paper X was excluded. Replace it with the concrete test you actually used to exclude it.”
- Screening process: “No flow diagram, and the protocol reports only a final count of 34 papers. I cannot verify how many records were removed at each stage, or why. Add a count at each step, from initial search to final inclusion.”
- Reproducibility: “Working only from what is written, I arrived at 29 of your 34 papers. The remaining 5 likely fall under the undocumented ‘organizational’ filter above, but I cannot confirm that without asking you.”
Notice the pattern across all four comments: each one names a specific line, quotes or paraphrases it, and states precisely what is missing rather than a general “make this clearer.” Calibrating on a model review before writing your own serves the same function here that Peer-Review Workshop I’s “Additional advice” does for the argument-chain rubric: it sets the expected level of specificity before, not after, you have written a review that falls short of it.
Structured Review
Assess the protocol you executed against these transparency criteria:
- Are the databases and the exact search strings stated in full, not summarized?
- Are inclusion and exclusion criteria written as rules a stranger could apply, not as impressions?
- Is the screening process (how many records at each stage) reported, ideally as a flow diagram?
- Would you, personally, have arrived at the same set of papers using only what is written down?
When authors revise a paper after review, a common and effective practice is to tabulate every reviewer comment in a table with three columns: the comment, the response, and internal notes. It forces a response to every point, not just the convenient ones. Apply the same discipline here: tabulate every gap the reproducing team found, and respond to each one specifically in your revised protocol, rather than doing a general tidy-up and hoping the specific issues got covered along the way.
Homework
Revise your own Review Protocol in light of the review you received.
For every gap the reproducing team identified, either close it (make the decision explicit and written) or justify why it is left as a judgement call.