Literature review · Screening
How to write inclusion and exclusion criteria you can apply
Criteria fail at the screening desk, not on the page. A rule that reads well but cannot be decided from a title and abstract turns every record into a judgement call, and two screeners into two different reviews.

In this article
Eligibility criteria are usually written once, early, and in prose. Then screening begins and the ambiguities surface all at once: does a mixed-age cohort count, is a pilot study a trial, what happens to the paper that measures the right thing in the wrong population. By then the criteria are load-bearing and changing them is expensive.
This describes how to write criteria that survive contact with a real record set. It produces a criteria matrix and a screening rule set. It does not tell you what your review should be about, and it does not replace your protocol.
1. Build the criteria from the review PICO
The Cochrane Handbook's chapter on defining criteria for including studies sets the structure directly: the population, intervention and comparison components of the review question, plus a specification of which study designs will be included, form the basis of the pre-specified eligibility criteria.
The useful discipline is to write each criterion twice — once as an inclusion rule and once as its exclusion counterpart — and check that they partition the space rather than overlap or leave a gap. Most ambiguity at the screening desk comes from a dimension where only the inclusion side was written down, leaving the screener to infer what the exclusion should be.
2. Outcomes are rarely eligibility criteria
This is the rule most often broken, and Cochrane states it directly: it is rare to use outcomes as eligibility criteria. Studies should be included irrespective of whether they report outcome data, though they may legitimately be excluded if they do not measure outcomes of interest at all, or if they explicitly aim to prevent the outcome you are studying.
The distinction is between measuring and reporting. A trial that measured your outcome but reported it selectively, or reported a null result in a single line, is still eligible — and excluding it is how reporting bias enters a review through the front door. If you screen on reported results, your included set is systematically enriched for studies that found something.
3. Make each rule decidable at the stage it is applied
A criterion has to be answerable from the information available at the point it is used. At title-and-abstract stage the screener has perhaps 250 words. A rule like “studies with adequate follow-up” cannot be applied there, and it cannot be applied consistently anywhere, because “adequate” is doing undefined work.
Rewrite each criterion until it is a question with a determinate answer, and assign it to a stage:
- Title and abstract: broad population, obvious design exclusions, clearly off-topic records. When in doubt, promote the record rather than excluding it — a false exclusion here is invisible later.
- Full text: anything needing the methods section — comparator description, follow-up duration stated as a number, whether adult data are separable.
- Never: anything requiring a judgement about quality. Risk of bias is an assessment of included studies, not a filter for getting in.
Replace vague quantifiers with values. “Adequate follow-up” becomes “follow-up of at least 30 days, stated explicitly.” “Recent studies” becomes a year with a justification in the protocol. If a threshold is arbitrary, say so in the text rather than hiding it behind an adjective.
4. Pilot the criteria before screening starts
Cochrane recommends piloting criteria to check that categories are sufficiently distinct to allow classification without being so narrow that they fragment. The cheapest version takes an afternoon: pull 30 to 50 records from the actual search, have both screeners apply the criteria independently, and compare.
Read the disagreements rather than counting them. Each one points at a specific rule that two competent readers interpreted differently, and that rule needs rewording before it is applied to fifteen hundred records. Agreement statistics tell you there is a problem; the disagreement list tells you which sentence caused it.
The Handbook's chapter on selecting studies asks for at least two people working independently on eligibility decisions, with title-and-abstract screening ideally done in duplicate too. Piloting is where you find out whether your criteria can actually support that, or whether the second screener is just reproducing the first one's guesses.
5. Record the reason, not just the outcome
Every full-text exclusion needs a reason drawn from a fixed list, because those reasons have to be reported by category in the flow diagram. Free-text reasons that vary by screener cannot be totalled, and totalling them later means re-reading papers you have already rejected.
Keep one column for the rule that triggered the exclusion and one for the specific fact that made it apply — wrong-comparator alongside “single-arm, no control described.” The category supports the count; the note is what lets you defend an individual decision when a reader asks about a specific paper.
Keep a record you could not decide as a record you could not decide. A study you were unable to assess is awaiting classification, not excluded, and that distinction has to survive into the diagram. The counting rules are in building a PRISMA flow diagram that reconciles, and the retained set carries forward into your evidence matrix.