Methodology
Published in full, including section 12, which is the register of what this review does not yet meet. A review that publishes its shortfalls can be argued with, which is the point.
Programme Threshold, brain injury and justice Maintained by Stan Gilmour, Oxon Advisory Academic lead Professor Huw Williams, University of Exeter Version 1.1 Issued 13 September 2026 Status Living document. Supersedes the method section of the Threshold Baseline Scoping Plan.
1. Purpose and Standing of This Document
This document records the methodology of the Threshold Baseline review: the standards it works to, the decisions taken and why, what has actually been done, and where current practice falls short of the standard it is aiming at. It exists so that any claim in the baseline document can be traced back to a stated method, and so that a reviewer, a commissioner or a journal editor can judge the work without taking anything on trust.
It is maintained continuously, and revised whenever the method changes. Every change to search, eligibility, appraisal or synthesis is recorded in section 13 with a date and a reason, and no method change is made silently. Section 12 holds the register of known departures from the standards named here, which is the section a critical reader should turn to first.
When this document is updated
- Any change to eligibility criteria, information sources or the search strategy.
- Any addition to or change in the appraisal instruments used.
- Any new release of a standard named here, checked at each version and at least every six months.
- Completion of a harvest phase, recording what that phase actually achieved.
- Any correction arising from verification that bears on method rather than on a single record.
Maintenance dependency
The standards cited here are versioned and several changed during 2024 to 2026. Section 14 lists each with its current version and the date it was last checked. Citing a superseded instrument is a substantive error in a methodology document, so the check is scheduled rather than occasional.
2. Review Type and Reporting Standard
The review type is a scoping review
The Threshold Baseline is a scoping review, not a systematic review. The decision follows the indications set out by Munn and colleagues (2018), whose criteria for preferring a scoping review are met on four counts: the purpose is to identify the types of evidence available in a field rather than to answer a single question of effectiveness; the field's key characteristics and concepts need clarifying, which the definitional problem described in section 4 makes acute; the intention includes identifying and analysing knowledge gaps; and the corpus deliberately spans study designs and document types that no single effect estimate could combine.
A systematic review remains the right instrument for particular questions inside the field, and two are candidates: the association between brain injury and reoffending, and the diagnostic accuracy of justice-setting screening instruments. Both would be separate pieces of work with their own protocols.
Standards adopted
| Function | Standard | Citation |
|---|---|---|
| Conduct | JBI scoping review methodology | Peters et al. (2024), in Aromataris et al. (eds), JBI Manual for Evidence Synthesis, 2024 edition |
| Reporting | PRISMA-ScR, 20 essential and 2 optional items | Tricco et al. (2018), doi:10.7326/M18-0850 |
| Search reporting | PRISMA-S, 16 items | Rethlefsen et al. (2021), doi:10.1186/s13643-020-01542-z |
| Search peer review | PRESS 2015 evidence-based checklist | McGowan et al. (2016) |
| Synthesis reporting where no meta-analysis is performed | SWiM | Campbell et al. (2020), doi:10.1136/bmj.l6890 |
| Use of artificial intelligence | RAISE, as adopted in the joint position statement of Cochrane, the Campbell Collaboration, JBI and the Collaboration for Environmental Evidence | Flemyng et al. (2025), doi:10.1002/cl2.70074 |
PRISMA 2020 (Page et al., 2021, doi:10.1136/bmj.n71) remains the parent standard and is not superseded as at September 2026. PRISMA-ScR sits on top of it and governs this review's reporting.
Registration
PROSPERO does not register scoping reviews. The Open Science Framework is the appropriate route and registration will be made there before the definitive search is run, with the protocol deposited and timestamped. Registration is not mandatory for a consultancy asset, and it is being done because a timestamped protocol is the cheapest available defence against the charge that eligibility was decided after the results were seen.
Relationship to an evidence and gap map
The corpus as structured, with a controlled theme vocabulary crossed against thirteen strands and a tier assignment, is close to an evidence and gap map in the sense described by White and colleagues (2020) for Campbell and 3ie. Producing a formal map from the same corpus is a derived output and is recorded as such rather than as a separate review.
3. Review Question and Eligibility
Question, in PCC form
PRISMA-ScR frames scoping review questions as Population, Concept, Context rather than PICO.
| Element | Definition adopted |
|---|---|
| Population | People who come into contact with the criminal justice system as suspects, defendants, people serving sentences, or victims and complainants, together with the justice workforce as an occupationally exposed group |
| Concept | Acquired and traumatic brain injury: its prevalence, ascertainment, consequences, management and legal treatment |
| Context | Any jurisdiction, any point in the justice pathway, 1988 to the present |
Inclusion criteria
- Empirical studies of any design reporting on brain injury in a population defined above.
- Systematic reviews, meta-analyses and scoping reviews of the same.
- Clinical and practice guidelines governing what a health or justice agency does about brain injury.
- Policy documents, inspection reports and official statistics from a named issuing body.
- Legal instruments, statutory provisions and reported case law bearing on the treatment of cognitive impairment in justice process.
- Qualitative and lived-experience research.
- Studies of sport-related, military and occupational brain injury where they supply mechanism evidence, workforce evidence, or a directly drawn comparison with a justice population.
Exclusion criteria
- Wider neurodisability as a literature in its own right, meaning foetal alcohol spectrum disorder, attention deficit hyperactivity disorder, autism, learning disability, and speech, language and communication needs. Studies reporting these alongside brain injury, or using a composite neurodisability measure as the only available ascertainment, are included and flagged so the distinction survives into synthesis.
- Acute clinical management of head injury, except where a guideline governs a justice agency's duty.
- Animal studies, and human studies with no justice, victimisation or justice-workforce component.
- Non-English language sources, which is a resource limitation recorded as a deviation in section 12.
Handling three classes of document
The corpus contains three kinds of material and they are not equivalent. The distinction is carried in the record itself and into the synthesis.
Empirical research is evidence. It is appraised, tiered and synthesised.
Policy, guidance and inspection material is treated as evidence of what a system has committed to or observed, not as evidence of effect. Where such a document reports an original finding it is appraised as grey literature under AACODS (Tyndall, 2010). Where it restates a finding from a primary study, it is recorded as restating and the primary study is cited instead. This rule was applied during harvest v1 after reviewers noted advocacy reports recycling the same headline figures.
Legal instruments and case law are context. No recognised appraisal instrument exists for statutes, international conventions or judgments within evidence synthesis methodology, and none has been invented for this review. They are cited to their official source, used to establish the normative and procedural framework, and are neither tiered nor counted as evidence of effect. This is a bespoke convention, adopted because the alternative was to exclude the legal framework from a review about rights and legal process, and it is declared as bespoke rather than presented as standard practice.
4. Case Definition and Ascertainment
This section carries more weight in this field than in most, because the largest single source of variation in published findings is how brain injury was defined and detected. Applying five current definitions of mild traumatic brain injury to the same 596 blunt trauma patients produced sensitivities ranging from zero to 100 per cent, in a single-centre retrospective study (Harris et al., 2024). Within Scottish prisons, hospital-record ascertainment gives a prevalence of 24.7 per cent while self-report interview in comparable establishments gives 86 per cent (McMillan et al., 2019; McMillan et al., 2025). The meta-analysis of Hunter and colleagues (2023) found definition and detection method to be significant moderators of its pooled 45.8 per cent estimate.
Reference standards adopted
The review adopts a tiered case definition rather than a single one, because the available schemes were built for different purposes.
| Ascertainment route in the source study | Reference standard applied | Why |
|---|---|---|
| Prospective, interview-based or clinically assessed | American Congress of Rehabilitation Medicine 2023 diagnostic criteria for mild traumatic brain injury (Silverberg et al., 2023, Archives of Physical Medicine and Rehabilitation, 104(8), pp. 1343-1355, doi:10.1016/j.apmr.2023.03.036) | A Delphi-consensus revision of the 1993 ACRM definition reaching 90.7 per cent final agreement, unifying criteria across sport, civilian and military contexts |
| Retrospective or administrative record | Mayo classification system (Malec et al., 2007) | Designed to classify incomplete records and maps onto diagnostic coding |
| Coded health data | Coding system stated, ICD-10 or ICD-11 | The transition between revisions changes what is counted |
Ascertainment recorded as a structured field
Every included record carries an exposure_measure field naming how brain injury was established:
the named instrument, self-report, structured interview, clinical assessment, or administrative
record. No prevalence figure enters the synthesis without its ascertainment method attached. Pooled or
compared figures are only presented within ascertainment strata.
Common data elements
Included studies are mapped, where they report enough to allow it, onto the National Institute of Neurological Disorders and Stroke Common Data Elements for traumatic brain injury. Mapping is partial by necessity and its coverage is reported rather than assumed.
4a. The Three Epochs
The review is organised into three epochs, so that the report can show how the field's position was reached rather than only where it now stands. Epoch is derived from the year of publication and is never hand-entered.
| Code | Years | Working name | Works |
|---|---|---|---|
| e0 | before 1996 | Antecedents, context only | 4 |
| e1 | 1996 to 2005 | Establishing the question | 91 |
| e2 | 2006 to 2015 | Establishing the prevalence | 116 |
| e3 | 2016 to 2026 | Testing the causal claim | 296 |
Even decades were chosen over boundaries drawn at events in the field. An event-based cut would tell a better story and would require the review to defend the boundary as well as the finding. The working names describe what the harvests found and are revised if the evidence says otherwise.
Additional fields for the two earlier epochs
Works in e1 and e2 carry four fields that e3 works do not, because the question asked of them is different. The 2016 to 2026 literature is asked what it found. The earlier literature is also asked what it did.
epoch_rolerecords the work's function at the time: originating, establishing, consolidating, instrumenting, challenging, applying or context.standing_nowrecords where its claim sits today: holds, refined, contested, overturned, untested or historic.superseded_bylinks to later works in the corpus that superseded, replicated or overturned it. A link is recorded only where it has been confirmed. An unverified lineage is worse than none, and the build rejects a link whose target is not a corpus record.why_it_matteredstates in one or two sentences what the work did to the field, which is a different claim from what it found.
Historic works are recorded as they were written. A finding is never corrected to later standards in
key_findings; standing_now and why_it_mattered place it instead. Where the prevailing belief of
an epoch has since been overturned, both the work that established the belief and the work that
overturned it are harvested, and they are linked.
Depth of the earlier epochs
The earlier epochs were harvested to a targeted depth rather than a symmetrical one. A work earns its place if it shaped what the field believed at the time, or if its fate since is instructive. Completeness for 1996 to 2015 was not attempted and is not claimed. This is a deliberate limitation and is recorded in section 12.
The 2006 to 2015 epoch is fuller than its targeted harvest alone would suggest, because 86 works of that decade were already in the corpus from the thematic strands. The targeted harvest added 49 more after deduplication removed those already held.
5. Information Sources and Search
Minimum database set
Harvest v1 searched a narrower set than the standard requires. The definitive search will cover the following, selected on the recall evidence of Bramer and colleagues (2017, doi:10.1186/s13643-017-0644-y), whose tested minimum of Embase, MEDLINE, Web of Science and Google Scholar reached 98.3 per cent recall in health topics, extended for criminology and law on the Campbell Collaboration information retrieval guidance (MacDonald et al., 2024, doi:10.1002/cl2.1433).
| Source | Reason |
|---|---|
| MEDLINE | Core biomedical coverage of brain injury, prevalence and screening |
| Embase | Near-essential complement to MEDLINE, adds European and clinical coverage |
| PsycINFO | Behavioural, cognitive, offending trajectory and desistance literature |
| CINAHL | Nursing and allied health delivery of screening, linkworker and rehabilitation services |
| Web of Science Core Collection | Multidisciplinary reach into criminology, plus citation chasing |
| Scopus | Complementary multidisciplinary and citation coverage with imperfect overlap with Web of Science |
| Criminal Justice Abstracts | The specialist criminology database |
| ProQuest Dissertations and Theses | Unpublished postgraduate research, which the qualitative and causation literatures rely on |
| NCJRS Abstracts | US justice policy and evaluation reports, searched as a grey literature source |
| Google Scholar | Supplementary only, for grey literature triangulation and forward citation searching, first 200 results screened and reported |
HeinOnline, Westlaw and Lexis are used for targeted retrieval of named legal instruments. No recall claim is made for legal searching, because no recall study for legal databases was found.
Search structure and peer review
Search concepts are built on the PCC frame for the scoping question and on a PECO frame, meaning population, exposure, comparator, outcome, for the exposure-outcome sub-questions. The full Boolean strategy for at least one database is reproduced in the review, with all others available in the supplementary record, as PRISMA-S requires.
The strategy is peer reviewed against the PRESS 2015 checklist (McGowan et al., 2016) by an information specialist who is not the person who wrote it. This has not yet been done and is recorded as an outstanding action.
Grey literature
Grey literature is searched to a documented method, using the CADTH Grey Matters checklist as the source frame and a documented, time-boxed website search of named issuing bodies: gov.uk, HM Inspectorate of Prisons, HM Inspectorate of Probation, HMICFRS, the Prisons and Probation Ombudsman, the Scottish Government, the Scottish Prison Service, Healthcare Improvement Scotland, the Welsh Government, the Department of Justice Northern Ireland, the Irish Prison Service, the US Centers for Disease Control and Prevention, the National Institute of Justice, SAMHSA, Public Safety Canada, the Office of the Correctional Investigator, the Australian Institute of Criminology, Ara Poutama Aotearoa, the World Health Organization and UNODC. For each, the date searched, the terms used and the number of items screened are recorded.
Grey literature materially changes conclusions in this field. The frequently quoted United Kingdom figures of £43 billion in annual acquired brain injury cost and a 16:1 return on rehabilitation appear only in grey literature and are modelled rather than measured, which is exactly the kind of claim that enters policy unchallenged when grey literature is either excluded or accepted uncritically.
Supplementary methods
Backward and forward citation chasing is run on every tier one record. Snowballing supplements rather than substitutes for database searching. Reference lists of included systematic reviews are screened. Named authors with substantial bodies of work in the field are searched directly.
Limits, deduplication and currency
English language only, 1988 onwards. Deduplication is performed on normalised digital object identifier, then PubMed identifier, then normalised title and year, with the merge rule preferring the richer record and preserving strand membership; the script that does this is version controlled and its output audited. The search is re-run before the baseline document is finalised if more than six months have elapsed, and the re-run date is reported.
6. Selection and Charting
Screening configuration
Titles and abstracts are screened by a single experienced reviewer, with a second reviewer verifying all full-text exclusions. This follows the configuration Cochrane's Rapid Reviews Methods Group permits for resource-constrained reviews (Garritty et al., 2021, doi:10.1016/j.jclinepi.2020.10.007; superseded for rapid review conduct generally by Garritty et al., 2024, doi:10.1136/bmj-2023-076335).
Single screening is not equivalent to double screening. Waffenschmidt and colleagues (2019, doi:10.1186/s12874-019-0782-0) found single screening misses a median 5 per cent of studies, rising to 13 per cent for less experienced reviewers. The configuration above is a considered trade-off and is declared as such rather than described as best practice. A minimum 20 per cent dual screen of titles and abstracts is applied as a calibration check, and the disagreement rate is reported.
Charting
Data are charted into a fixed 31-field schema covering bibliographic identity, verification route, population, jurisdiction, setting, design, sample size, ascertainment method, outcomes, findings, effect sizes with intervals, stated limitations, an evidential weight note, a tier, an era, a controlled theme vocabulary, contested status, and free notes. The schema is reproduced in Appendix B and is version controlled; changes to it are logged in section 13.
Findings are charted as the source states them. Numbers are transcribed, never inferred, and where a source gives a range the range is recorded. No finding is paraphrased into a stronger claim than the source makes.
7. Critical Appraisal
Critical appraisal is optional under PRISMA-ScR. It is being done here because the baseline is intended to support commissioning decisions, and a map of a field that does not distinguish a national register cohort from a service audit is not fit for that purpose.
Instrument by design
| Design in the corpus | Instrument | Version |
|---|---|---|
| Prevalence and cross-sectional prevalence studies | Hoy et al. risk of bias tool | 2012, doi:10.1016/j.jclinepi.2011.11.014 |
| Analytical cross-sectional studies | JBI checklist for analytical cross-sectional studies | Current JBI suite |
| Cohort and register-based exposure-outcome studies | ROBINS-E | Higgins et al., 2024, doi:10.1016/j.envint.2024.108602 |
| Evaluations of a defined intervention or programme | ROBINS-I | Sterne et al., 2016 |
| Randomised trials | RoB 2 | Sterne et al., 2019 |
| Diagnostic accuracy and screening instrument validation | QUADAS-3 | Whiting et al., 2026, Annals of Internal Medicine, 179(4), pp. 548-555, doi:10.7326/ANNALS-25-02104 |
| Comparative diagnostic accuracy | QUADAS-C | Yang et al., 2021 |
| Measurement properties of a screening instrument | COSMIN guideline v2.0 | Mokkink, Elsman and Terwee, 2024, doi:10.1007/s11136-024-03761-6 |
| Systematic reviews whose findings are used as evidence | ROBIS | Whiting et al., 2016 |
| Systematic reviews rated for conduct | AMSTAR 2 | Shea et al., 2017 |
| Qualitative studies | CASP qualitative checklist or JBI qualitative checklist | Current |
| Mixed methods studies | MMAT | Hong et al., 2018, Education for Information, 34(4), pp. 285-291 |
| Clinical guidelines, development process | AGREE II | Brouwers et al., 2010 |
| Clinical guidelines, specific recommendations | AGREE-REX | Brouwers et al., 2020 |
| Grey literature and policy reports | AACODS | Tyndall, 2010 |
| Legal instruments and case law | None; treated as context | Not applicable |
QUADAS-3 supersedes QUADAS-2 for study-level diagnostic accuracy appraisal. QUADAS-C remains a separate live tool for comparative accuracy and has not been folded into QUADAS-3. Where a corpus study reports its own appraisal under QUADAS-2, that is recorded as historical rather than corrected.
The Newcastle-Ottawa Scale is used only where no alternative fits, with its documented reliability and validation problems stated at the point of use.
Qualitative appraisal produces a structured trustworthiness judgement and never a numeric threshold for exclusion, following the objection to checklist scoring in qualitative synthesis set out by Dixon-Woods and colleagues (2006).
The tier system, and what it is not
Every record carries a tier: one for landmark or definitive work, two for solid contributory evidence, three for context and commentary. The tier orders records for navigation and prioritisation. It is not a risk of bias assessment, it is not a certainty rating, and it should not be reported as either.
Its weaknesses are known and stated. Tiers in harvest v1 were assigned by the single reviewer who harvested the record, were not moderated across strands, and have no published validation. Ad hoc hierarchies of this kind are criticised in the methodological literature for exactly this reason. Formal appraisal under section 7 replaces the tier as the basis for any evidential claim, and the tier survives only as a way of ordering a search interface. Until formal appraisal is complete, the baseline document does not describe a study as strong or weak on the basis of its tier alone.
Certainty of evidence
GRADE is applied to causal and prognostic claims, using the prognostic factor guidance of Foroutan and colleagues (2020) and the overall prognosis guidance of Iorio and colleagues (2015).
GRADE is not applied to prevalence estimates. The GRADE Working Group has published no guidance article addressing certainty of evidence for prevalence or burden estimates, and GRADE's core domains, which turn on effect size and comparator-based imprecision, do not map onto a single-group proportion. Certainty in prevalence findings is described narratively, in terms of the number and quality of contributing studies, their ascertainment methods, and the spread of estimates within ascertainment strata.
GRADE-CERQual (Lewin et al., 2018) is applied to findings from qualitative synthesis.
8. Synthesis
Synthesis is narrative and structured, reported against SWiM (Campbell et al., 2020). No meta-analysis is performed, since the corpus was assembled to map a field rather than to pool a homogeneous set of estimates, and the ascertainment heterogeneity described in section 4 would make a pooled prevalence figure misleading at the level of precision it would appear to offer.
The synthesis groups by strand for navigation and by question for argument. Within each question the order is: what the evidence establishes, what it disputes, and what it does not address. Numerical findings are reported with their intervals and their ascertainment method. Where the corpus contains an existing meta-analysis, that meta-analysis is reported rather than recalculated.
Contested findings are handled explicitly. A record carries a contested flag and a contested_note
naming who disputes the finding and on what grounds. In the synthesis, a contested question is
presented with the strongest case on each side and the studies that carry it, and is not resolved by
authorial preference. The association between brain injury and offending is the principal instance and
is treated this way throughout.
9. Use of Artificial Intelligence
Harvest v1 was assembled with substantial assistance from large language models. This section is the declaration required by RAISE, as adopted in the joint position statement of Cochrane, the Campbell Collaboration, JBI and the Collaboration for Environmental Evidence (Flemyng et al., 2025, doi:10.1002/cl2.70074). Campbell's status as a co-signatory makes this the applicable standard for criminological as well as health synthesis. The RAISE guidance documents themselves are hosted on the Open Science Framework rather than published as journal articles, and are cited accordingly.
What was used, and for what
| Task | Tool | Human role |
|---|---|---|
| Search formulation and execution across bibliographic APIs and the web | Claude Opus 5, thirteen parallel agent instances, September 2026 | Search domains, inclusion boundaries and strand definitions specified by the reviewer |
| Record identification and screening | As above | No independent human screen of all identified records was performed at harvest v1 |
| Data charting into the 31-field schema | As above | Schema designed by the reviewer; field values not independently re-extracted |
| Adversarial verification of bibliographic identity | Claude Opus 5, two further agent instances, against Crossref, PubMed and issuing-body websites | Verification protocol specified by the reviewer; verdicts and corrections logged and reviewable |
| Drafting of this document and the scoping plan | Claude Opus 5 | Authored, edited and taken responsibility for by the reviewer |
What was not used
No tier assignment, eligibility decision or synthesis conclusion in the baseline document rests on an unreviewed model output. No formal risk of bias assessment has been produced by a model, and none will be: the evidence below makes that indefensible.
The performance evidence this declaration rests on
Model performance varies sharply by task, and the declaration is calibrated to it rather than to a general claim of reliability.
Screening performs comparatively well. Xie and colleagues (2026) pooled 18 studies published between 2023 and 2025 and reported sensitivity of 0.92, with a 95 per cent confidence interval of 0.81 to 0.96, and specificity of 0.94, with an interval of 0.90 to 0.97, at title and abstract stage. Against that, Khraisha and colleagues (2024, Research Synthesis Methods) found GPT-4 screening performance fell to chance level under balanced datasets once chance agreement was corrected for, and rated data extraction only moderate.
Risk of bias assessment performs badly. Reported sensitivity for detecting high or low risk trials has ranged from 0.46 to 0.53, with kappa between 0.14 and 0.51, across evaluations by Gandhi and colleagues (2026) and Taneri (2025), and the former concludes explicitly that current models are unsuitable as a sole assessor.
Citation fabrication is the specific failure mode this project was designed against, and it is why the verification protocol in section 10 exists.
Responsibility
Responsibility for every claim in the Threshold Baseline rests with its named authors. The models used are tools within a method specified, supervised and corrected by the reviewer, and no part of this declaration transfers responsibility to them.
10. Verification and Error Correction
The two-pass protocol
Every record must have been confirmed to exist in a bibliographic database or on the publisher's or
issuing body's own website, and the record names where. This is stored as a verified_via field and is
a condition of entry rather than a later check.
Two adversarial verification passes are then run by reviewers instructed to assume nothing is real until confirmed.
Pass one takes every record carrying neither a digital object identifier nor a PubMed identifier, which is the highest-risk category and consists mainly of policy reports, guidelines and legal instruments. Each is checked for existence under the stated title, issuing body and year; for a live and correct source location; and for supersession. Verdicts are confirmed, confirmed with correction, confirmed at a corrected location, superseded, unconfirmed, or likely fabricated.
Pass two takes a random sample of records carrying identifiers and resolves each against Crossref and PubMed, comparing title, first author, year and journal, and checking that identifiers agree with each other where both are present.
Results at harvest v1
Pass one covered all 73 records without identifiers: 53 confirmed, 13 confirmed with a field correction, 4 confirmed at a corrected location, 3 unconfirmed, none fabricated. Pass two covered a random sample of 60 tier one records: 57 exact matches, 3 with a minor field error, no identifier resolving to a different work, none fabricated.
Pass three covered the 178 records added by the epoch harvests, being all 14 without an identifier and a random sample of 50 carrying one: 60 confirmed, 2 confirmed with a field correction, 1 minor correction, and 1 identifier that resolved to a different work and has been removed from the record. No fabrication was found. The reviewer's judgement, which matters for whether the earlier epochs can carry weight, was that pre-2006 material is not materially less reliable than the rest.
Twenty five corrections were applied and logged against the records they affect, in a corrections_log
table recording the field, the old value, the new value and the pass that found it. The three
unconfirmed records carry a flag in their notes field and are excluded from citation without a manual
check.
Standing rules arising
- Records held at abstract level are marked as such and their findings are attributed to the abstract.
- Brain Injury assigns articles to print issues asynchronously from their online-first date. Any citation to that journal with a December online date has its year checked against the issue record.
- A grey literature document that restates a primary study's finding is recorded as restating, and the primary study is cited instead.
Verification coverage target
Harvest v1 and the epoch harvests verified 196 of 507 records, being all of the highest-risk category and a sample of tier one. The target before the baseline document is published is 100 per cent identifier resolution for every record cited in the text, and continued sampling for the remainder.
11. Ethics, Terminology and Framing
The review concerns people in custody and people with cognitive impairment, and the way it describes them is a methodological matter rather than a stylistic one.
Person-first language is used throughout: people with brain injury, people in custody, children in the youth justice system. Vulnerability is treated as produced by circumstances rather than as a property of a person. Neurodisability describes a condition and never a person. Statutory categories are used where the statutory category is what is meant, and descriptive terms are preferred elsewhere.
No individual case material, service record or operational example enters the review. All included material is already published or publicly issued.
Where the review reports a finding that could be read as attributing offending to a characteristic of a group, the finding is stated with its confidence interval, its adjustment set and its contested status, because a partial statement of that literature does foreseeable harm.
12. Known Deviations from Standard
This section is the honest register. Harvest v1 does not yet meet the standard set out above, and the gap is stated rather than managed.
- Harvest v1 is a structured expert search, not a PRISMA-ScR compliant search. It has no reported Boolean strategy, no reported screening flow, no PRESS peer review and no dual screening calibration. It should not be described as systematic, and the baseline document must not claim PRISMA-ScR compliance until the definitive search has been run and reported.
- Database coverage in harvest v1 was narrower than the minimum set in section 5. PubMed, Semantic Scholar and Consensus were searched, supplemented by targeted web searching. Embase, CINAHL, Criminal Justice Abstracts, Scopus, Web of Science and ProQuest were not.
- Formal critical appraisal has not been performed. The tier assignment is a reviewer judgement, unmoderated across strands, and is not a substitute.
- Screening was not independently performed by a human at harvest v1. Verification addressed whether records exist and are correctly described, and did not address whether records that should have been found were missed.
- Approximately 40 records are held at abstract level. Their charted findings reflect the abstract and not the full text.
- Three records remain unconfirmed and are flagged in place.
- English language only, which introduces language bias of unknown size in a review covering Nordic register studies and Gulf epidemiology.
- Grey literature coverage is deliberately uneven, being thorough for the United Kingdom, the United States, Canada, Australia and New Zealand, and thin elsewhere. No justice health policy naming brain injury was found in Northern Ireland, the Republic of Ireland or the Gulf, which is reported as a finding rather than a coverage failure, with the caveat that thin searching cannot fully distinguish the two.
- No published precedent was found for an AI-assisted evidence harvest documented with an adversarial verification pass in either criminology or health. The protocol in section 10 is therefore original and untested against an external standard.
- The two earlier epochs were harvested to a targeted depth, not a complete one. The 1996 to 2005 and 2006 to 2015 records show what shaped the field, and a prevalence or intervention study of those decades may be absent without that absence meaning anything. No completeness claim is made for either epoch, and no count from them should be read as a count of the literature.
- The reported rise in fabricated references in the academic literature, cited in press coverage of a 2026 Lancet analysis, could not be traced to its primary source and its figures are therefore not reproduced here.
13. Version History
| Version | Date | Change | Reason |
|---|---|---|---|
| 1.1 | 13 September 2026 | Added section 4a, the three-epoch design, with the four additional fields for the earlier epochs and the rule that epoch is derived from year. Recorded the targeted depth of the earlier harvests as a limitation. Added verification pass three, covering the 178 records the epoch harvests added. | Stan asked for two further epochs covering the two prior decades, so the report can show the research journey rather than only the settled position. |
| 1.0 | 13 September 2026 | First issue. Records harvest v1 method, adopts scoping review type, PRISMA-ScR reporting, JBI conduct, the appraisal instrument set, and the RAISE declaration. Establishes the deviations register. | Methodology required as a maintained document rather than a section of the scoping plan. |
14. Standards Register and Next Check
Each standard below was checked against its own source or publishing journal on 13 September 2026. Next scheduled check: 13 March 2027, or on notice of a new release.
| Standard | Current version | Date | Verified |
|---|---|---|---|
| PRISMA 2020 | 2020 statement | 2021 | Confirmed current, not superseded |
| PRISMA-ScR | 2018 | 2018 | Confirmed, no revision issued |
| PRISMA-S | 2021 | 2021 | Confirmed, no revision issued |
| PRISMA-LSR (living reviews, not adopted here) | Akl et al., 2024, doi:10.1136/bmj-2024-079183 | 2024 | Confirmed |
| JBI Manual for Evidence Synthesis | 2024 edition | 2024 | Confirmed |
| JBI scoping review methodology | Peters et al. | 2024 | Confirmed |
| Cochrane Handbook | Version 6.5 | 27 November 2024 | Confirmed current |
| Campbell conduct standards | Aloe et al., doi:10.1002/cl2.1445 | 2024 | Confirmed |
| Rapid review guidance | Garritty et al., doi:10.1136/bmj-2023-076335 | 2024 | Confirmed, supersedes 2021 interim guidance |
| SWiM | Campbell et al. | 2020 | Confirmed, no update |
| RAISE joint position statement | Flemyng et al., doi:10.1002/cl2.70074 | November 2025 | Confirmed; RAISE 1 to 3 are OSF-hosted documents, not journal articles |
| ROBINS-E | Higgins et al., doi:10.1016/j.envint.2024.108602 | 2024 | Confirmed as the citation the tool's own site requests |
| QUADAS-3 | Whiting et al., doi:10.7326/ANNALS-25-02104 | Online February 2026, print April 2026 | Confirmed; supersedes QUADAS-2 |
| COSMIN | Guideline v2.0, doi:10.1007/s11136-024-03761-6 | 2024 | Confirmed |
| Hoy prevalence risk of bias tool | doi:10.1016/j.jclinepi.2011.11.014 | 2012 | Confirmed, nothing supersedes it |
| AACODS | Tyndall, Flinders University | November 2010 | Confirmed; frequently mis-cited as 2008 |
| ACRM mild TBI criteria | Silverberg et al., doi:10.1016/j.apmr.2023.03.036 | 2023 | Confirmed |
| GRADE for prevalence | Does not exist | n/a | No GRADE Working Group guidance article for prevalence or burden certainty was found |
| PROSPERO eligibility | Does not register scoping reviews | 2026 | Confirmed via secondary source only; manual check outstanding |
Citation status of this document
Every source cited here has been resolved against Crossref, PubMed or the issuing body's own site. Of 41 authored sources, 39 resolved exactly as cited and 2 required a correction, both applied above. No source cited in this document was found to be non-existent. The full Harvard reference list is at Appendix C.
Appendix A. Outstanding Methodological Actions
- Run and report the definitive search across the ten-source minimum set, with the Boolean strategy reproduced and PRESS peer review completed by an information specialist.
- Register the protocol on the Open Science Framework before that search runs.
- Complete formal critical appraisal by design, replacing the tier as the basis for evidential claims.
- Retrieve full text for records currently held at abstract level and re-chart them.
- Resolve identifiers for all remaining records, and clear or withdraw the three unconfirmed items.
- Confirm PROSPERO's current eligibility wording directly.
- Record the dual screening calibration rate once the definitive search is screened.
Appendix B. Charting Schema
The 31-field record schema and the controlled theme vocabulary are held in the corpus repository at
schema/record-schema.md and schema/themes.md, version controlled, and reproduced in full in the
baseline document's methodological annex. Fields fall into four groups: bibliographic identity and
verification route; population, setting and design; findings, effect sizes and stated limitations; and
reviewer judgements comprising evidential weight, tier, era, themes, jurisdictional relevance and
contested status.
Appendix C. References
Harvard, author-date. Every entry resolved against Crossref, PubMed or the issuing body's own site on 13 September 2026.
Akl, E.A., Meerpohl, J.J., Elliott, J., Kahale, L.A., Schünemann, H.J. and the Living Systematic Review Network (2024) 'Living systematic reviews: PRISMA-LSR extension', BMJ, 387, e079183. doi:10.1136/bmj-2024-079183.
Aloe, A.M., Bangpan, M., Dickson, K., Miguel, C. and Pigott, T.D. (2024) 'Campbell Collaboration conduct standards', Campbell Systematic Reviews. doi:10.1002/cl2.1445.
Bramer, W.M., Rethlefsen, M.L., Kleijnen, J. and Franco, O.H. (2017) 'Optimal database combinations for literature searches in systematic reviews: a prospective exploratory study', Systematic Reviews, 6, p.245. doi:10.1186/s13643-017-0644-y.
Brouwers, M.C., Kho, M.E., Browman, G.P., Burgers, J.S., Cluzeau, F., Feder, G., Fervers, B., Graham, I.D., Grimshaw, J., Hanna, S.E., Littlejohns, P., Makarski, J. and Zitzelsberger, L., for the AGREE Next Steps Consortium (2010) 'AGREE II: advancing guideline development, reporting and evaluation in health care', CMAJ, 182(18), pp.E839–E842. doi:10.1503/cmaj.090449.
Brouwers, M.C., Spithoff, K., Kerkvliet, K., Alonso-Coello, P., Burgers, J., Cluzeau, F., Férvers, B., Graham, I., Grimshaw, J., Hanna, S., Kastner, M., Kho, M., Kunnamo, I., Vandvik, P.O., Vernooij, R., Zitzelsberger, L. and the AGREE Next Steps Consortium (2020) 'Development and validation of a tool to assess the quality of clinical practice guideline recommendations', JAMA Network Open, 3(5), e205535. doi:10.1001/jamanetworkopen.2020.5535.
Campbell, M., McKenzie, J.E., Sowden, A., Katikireddi, S.V., Brennan, S.E., Ellis, S., Hartmann-Boyce, J., Ryan, R., Shepperd, S., Thomas, J., Welch, V. and Thomson, H. (2020) 'Synthesis without meta-analysis (SWiM) in systematic reviews: reporting guideline', BMJ, 368, l6890. doi:10.1136/bmj.l6890.
Dixon-Woods, M., Bonas, S., Booth, A., Jones, D.R., Miller, T., Sutton, A.J., Shaw, R.L., Smith, J.A. and Young, B. (2006) 'How can systematic reviews incorporate qualitative research? A critical perspective', Qualitative Research, 6(1), pp.27–44. doi:10.1177/1468794106058867.
Flemyng, E., Moher, D., Marshall, I.J., Cumpston, M., Aromataris, E., Grant, S., Wilson, C., Foxlee, R., Higgins, J.P.T. and the RAISE consortium (2025) Recommendations for the Appropriate use of Artificial Intelligence in Systematic Evidence syntheses (RAISE), joint position statement of Cochrane, the Campbell Collaboration, JBI and the Collaboration for Environmental Evidence. doi:10.1002/cl2.70074. [Hosted on the Open Science Framework rather than published as a journal article.]
Foroutan, F., Guyatt, G., Zuk, V., Vandvik, P.O., Alba, A.C., Mustafa, R., Vernooij, R., Arevalo-Rodriguez, I., Munn, Z., Roshanov, P., Riley, R., Schandelmaier, S., Kuijpers, T., Siemieniuk, R., Canelo-Aybar, C., Schünemann, H. and Iorio, A. (2020) 'GRADE Guidelines 28: use of GRADE for the assessment of evidence about prognostic factors: rating certainty in identification of groups of patients with different absolute risks', Journal of Clinical Epidemiology, 121, pp.62–70. doi:10.1016/j.jclinepi.2019.12.023.
Gandhi, S., Shokravi, A., Chelliahpillai, Y. and Balas, M. (2026) 'Evaluating large language model performance in Risk of Bias assessments: a cross-sectional validation study', PLoS One, 21(7), e0353155. doi:10.1371/journal.pone.0353155.
Garritty, C., Gartlehner, G., Nussbaumer-Streit, B., King, V.J., Hamel, C., Kamel, C., Affengruber, L. and Stevens, A. (2021) 'Cochrane Rapid Reviews Methods Group offers evidence-informed guidance to conduct rapid reviews', Journal of Clinical Epidemiology, 130, pp.13–22. doi:10.1016/j.jclinepi.2020.10.007. [Superseded for general rapid review conduct by Garritty et al., 2024, below; retained here as the citation for the single-reviewer screening configuration.]
Garritty, C., Tsertsvadze, A., Hamel, C., Devane, D., Miller, D.W.J., Skidmore, B., Nikolic, D., Fadeeva, A., Mikita, J., Featherstone, R., Vafaei, A., Motilall, A., Skoetz, N., Stevens, A. and the 2024 Rapid Review Guidance Working Group (2024) 'Updated recommendations for the conduct, reporting, and appraisal of rapid reviews: a consensus-based methods guidance paper', BMJ, 384, e076335. doi:10.1136/bmj-2023-076335.
Harris, K., Brusnahan, A., Shugar, S. and Miner, J. (2024) 'Defining mild traumatic brain injury: from research definition to clinical practice', The Journal of Surgical Research, 298, pp.101–107. doi:10.1016/j.jss.2024.03.006.
Higgins, J.P.T., Morgan, R.L., Rooney, A.A., Taylor, K.W., Thayer, K.A., Silva, R.A., Lemeris, C., Akl, E.A., Bateson, T.F., Berkman, N.D., Glenn, B.S., Hsu, S.A., LaKind, J.S., Lam, J., Masuda, Y.J., Radke, E.G., Rooney, A.A., Schünemann, H.J., Thayer, K.A., Vandenberg, J.J., Vesterinen, H.M., Wattam, S., Whaley, P. and Wolffe, T.A.M. (2024) 'A guideline for assessing the certainty in evidence for observational and non-randomized studies of exposures (ROBINS-E) and its extensions to cover epidemiological and environmental health studies', Environment International, 186, 108602. doi:10.1016/j.envint.2024.108602.
Hong, Q.N., Fàbregues, S., Bartlett, G., Boardman, F., Cargo, M., Dagenais, P., Gagnon, M.P., Griffiths, F., Nicolau, B., O'Cathain, A., Rousseau, M.C., Vedel, I. and Pluye, P. (2018) 'The Mixed Methods Appraisal Tool (MMAT) version 2018 for information professionals and researchers', Education for Information, 34(4), pp.285–291. doi:10.3233/EFI-180221.
Hoy, D., Brooks, P., Woolf, A., Blyth, F., March, L., Bain, C., Baker, P., Smith, E. and Buchbinder, R. (2012) 'Assessing risk of bias in prevalence studies: modification of an existing tool and evidence of interrater agreement', Journal of Clinical Epidemiology, 65(9), pp.934–939. doi:10.1016/j.jclinepi.2011.11.014.
Hunter, S., Bunn, F., Isham, L., Rawnsley, N., Barnicoat, K. and Bramwell, C. (2023) 'The prevalence of traumatic brain injury (TBI) among people impacted by the criminal legal system: an updated meta-analysis and subgroup analyses', Law and Human Behavior, 47(5), pp.539–565. doi:10.1037/lhb0000543.
Iorio, A., Spencer, F.A., Falavigna, M., Alba, C., Lang, E., Burnand, B., McGinn, T., Hayden, J., Williams, K., Shea, B., Wolff, R., Kujpers, T., Perel, P., Vandvik, P.O., Glasziou, P., Schünemann, H. and Guyatt, G. (2015) 'Use of GRADE for assessment of evidence about prognosis: rating confidence in estimates of event rates in broad categories of patients', BMJ, 350, h870. doi:10.1136/bmj.h870.
Johnson, S.D., Tilley, N. and Bowers, K.J. (2015) 'Introducing EMMIE: an evidence rating scale to encourage mixed-method crime prevention synthesis reviews', Journal of Experimental Criminology, 11(3), pp.459–473. doi:10.1007/s11292-015-9238-7.
Khraisha, Q., Put, S., Kappenberg, J., Warraitch, A. and Hadfield, K. (2024) Research Synthesis Methods. [Supplied as pre-verified without a DOI or full title; formatted as given and not re-verified in this pass.]
Lewin, S., Booth, A., Glenton, C., Munthe-Kaas, H., Rashidian, A., Wainwright, M., Bohren, M.A., Tunçalp, Ö., Colvin, C.J., Garside, R., Carlsen, B., Langlois, E.V. and Noyes, J. (2018) 'Applying GRADE-CERQual to qualitative evidence synthesis findings: introduction to the series', Implementation Science, 13(Suppl 1), p.2. doi:10.1186/s13012-017-0688-3.
MacDonald, M., Kavanagh, J. and Featherstone, R. (2024) 'Guidance on information retrieval for Campbell systematic reviews', Campbell Systematic Reviews. doi:10.1002/cl2.1433.
Malec, J.F., Brown, A.W., Leibson, C.L., Flaada, J.T., Mandrekar, J.N., Diehl, N.N. and Perkins, P.K. (2007) 'The Mayo classification system for traumatic brain injury severity', Journal of Neurotrauma, 24(9), pp.1417–1424. doi:10.1089/neu.2006.0245.
McGowan, J., Sampson, M., Salzwedel, D.M., Cogo, E., Foerster, V. and Lefebvre, C. (2016) 'PRESS Peer Review of Electronic Search Strategies: 2015 guideline statement', Journal of Clinical Epidemiology, 75, pp.40–46. doi:10.1016/j.jclinepi.2016.01.021.
McMillan, T.M., Graham, L., Pell, J.P., McConnachie, A. and Mackay, D.F. (2019) 'The lifetime prevalence of hospitalised head injury in Scottish prisons: a population study', PLoS One, 14(1), e0210427. doi:10.1371/journal.pone.0210427.
McMillan, T.M., Aslam, H., McGinley, A., Walker, V. and Barry, S.J.E. (2025) 'Associations between significant head injury and cognitive function, disability, and crime in adult men in prison in Scotland UK: a cross-sectional study', Frontiers in Psychiatry, 16, 1544211. doi:10.3389/fpsyt.2025.1544211.
Mokkink, L.B., Elsman, E.B.M. and Terwee, C.B. (2024) 'COSMIN guideline for systematic reviews of patient-reported outcome measures, version 2.0', Quality of Life Research, 33(11), pp.2929–2939. doi:10.1007/s11136-024-03761-6.
Munn, Z., Peters, M.D.J., Stern, C., Tufanaru, C., McArthur, A. and Aromataris, E. (2018) 'Systematic review or scoping review? Guidance for authors when choosing between a systematic or scoping review approach', BMC Medical Research Methodology, 18, p.143. doi:10.1186/s12874-018-0611-x.
Munn, Z., Moola, S., Lisy, K., Riitano, D. and Tufanaru, C. (2015) 'Methodological guidance for systematic reviews of observational epidemiological studies reporting prevalence and cumulative incidence data', International Journal of Evidence-Based Healthcare, 13(3), pp.147–153. doi:10.1097/XEB.0000000000000054.
Page, M.J., McKenzie, J.E., Bossuyt, P.M., Boutron, I., Hoffmann, T.C., Mulrow, C.D., Shamseer, L., Tetzlaff, J.M., Akl, E.A., Brennan, S.E., Chou, R., Glanville, J., Grimshaw, J.M., Hróbjartsson, A., Lalu, M.M., Li, T., Loder, E.W., Mayo-Wilson, E., McDonald, S., McGuinness, L.A., Stewart, L.A., Thomas, J., Tricco, A.C., Welch, V.A., Whiting, P. and Moher, D. (2021) 'The PRISMA 2020 statement: an updated guideline for reporting systematic reviews', BMJ, 372, n71. doi:10.1136/bmj.n71.
Peters, M.D.J., Godfrey, C., McInerney, P., Khalil, H., Larsen, P., Marnie, C., Pollock, D., Tricco, A.C. and Munn, Z. (2024) 'Chapter 11: Scoping reviews', in Aromataris, E., Lockwood, C., Porritt, K., Pilla, B. and Jordan, Z. (eds) JBI Manual for Evidence Synthesis. JBI, 2024 edition.
Rethlefsen, M.L., Kirtley, S., Waffenschmidt, S., Ayala, A.P., Moher, D., Page, M.J., Koffel, J.B. and the PRISMA-S Group (2021) 'PRISMA-S: an extension to the PRISMA statement for reporting literature searches in systematic reviews', Systematic Reviews, 10, p.39. doi:10.1186/s13643-020-01542-z.
Shea, B.J., Reeves, B.C., Wells, G., Thuku, M., Hamel, C., Moran, J., Moher, D., Tugwell, P., Welch, V., Kristjansson, E. and Henry, D.A. (2017) 'AMSTAR 2: a critical appraisal tool for systematic reviews that include randomised or non-randomised studies of healthcare interventions, or both', BMJ, 358, j4008. doi:10.1136/bmj.j4008.
Silverberg, N.D., Iverson, G.L., Cogan, A., Dams-O'Connor, K., Delmonico, R., Graf, M.J.P., Iaccarino, M.A., Kajankova, M., Kamins, J., McCulloch, K.L., McCrea, M., Panenka, W.J., Rabinowitz, A.R., Reyes, J. and Wethe, J.V. (2023) 'The American Congress of Rehabilitation Medicine diagnostic criteria for mild traumatic brain injury', Archives of Physical Medicine and Rehabilitation, 104(8), pp.1343–1355. doi:10.1016/j.apmr.2023.03.036.
Sterne, J.A.C., Hernán, M.A., Reeves, B.C., Savović, J., Berkman, N.D., Viswanathan, M., Henry, D., Altman, D.G., Ansari, M.T., Boutron, I., Carpenter, J.R., Chan, A.W., Churchill, R., Deeks, J.J., Hróbjartsson, A., Kirkham, J., Jüni, P., Loke, Y.K., Pigott, T.D., Ramsay, C.R., Regidor, D., Rothstein, H.R., Sandhu, L., Santaguida, P.L., Schünemann, H.J., Shea, B., Shrier, I., Tugwell, P., Turner, L., Valentine, J.C., Waddington, H., Waters, E., Wells, G.A., Whiting, P.F. and Higgins, J.P.T. (2016) 'ROBINS-I: a tool for assessing risk of bias in non-randomised studies of interventions', BMJ, 355, i4919. doi:10.1136/bmj.i4919.
Sterne, J.A.C., Savović, J., Page, M.J., Elbers, R.G., Blencowe, N.S., Boutron, I., Cates, C.J., Cheng, H.Y., Corbett, M.S., Eldridge, S.M., Emberson, J.R., Hernán, M.A., Hopewell, S., Hróbjartsson, A., Junqueira, D.R., Jüni, P., Kirkham, J.J., Lasserson, T., Li, T., McAleenan, A., Reeves, B.C., Shepperd, S., Shrier, I., Stewart, L.A., Tilling, K., White, I.R., Whiting, P.F. and Higgins, J.P.T. (2019) 'RoB 2: a revised tool for assessing risk of bias in randomised trials', BMJ, 366, l4898. doi:10.1136/bmj.l4898.
Taneri, P.E. (2025) 'Human versus artificial intelligence: comparing Cochrane authors' and ChatGPT's risk of bias assessments', Cochrane Evidence Synthesis and Methods, 3(5), e70044. doi:10.1002/cesm.70044.
Tricco, A.C., Lillie, E., Zarin, W., O'Brien, K.K., Colquhoun, H., Levac, D., Moher, D., Peters, M.D.J., Horsley, T., Weeks, L., Hempel, S., Akl, E.A., Chang, C., McGowan, J., Stewart, L., Hartling, L., Aldcroft, A., Wilson, M.G., Garritty, C., Lewin, S., Godfrey, C.M., Macdonald, M.T., Langlois, E.V., Soares-Weiser, K., Moriarty, J., Clifford, T., Tunçalp, Ö. and Straus, S.E. (2018) 'PRISMA extension for scoping reviews (PRISMA-ScR): checklist and explanation', Annals of Internal Medicine, 169(7), pp.467–473. doi:10.7326/M18-0850.
Tyndall, J. (2010) AACODS Checklist. Flinders University.
Waffenschmidt, S., Knelangen, M., Sieben, W., Bühn, S. and Pieper, D. (2019) 'Single screening versus conventional double screening for study selection in systematic reviews: a methodological systematic review', BMC Medical Research Methodology, 19, p.132. doi:10.1186/s12874-019-0782-0.
White, H., Albers, B., Gaarder, M., Kornør, H., Littell, J., Marshall, Z., Mathew, C., Pigott, T., Snilstveit, B., Waddington, H. and Welch, V. (2020) 'Guidance for producing a Campbell evidence and gap map', Campbell Systematic Reviews. doi:10.1002/cl2.1125.
Whiting, P., Savović, J., Higgins, J.P.T., Caldwell, D.M., Reeves, B.C., Shea, B., Davies, P., Kleijnen, J. and Churchill, R. (2016) 'ROBIS: a new tool to assess risk of bias in systematic reviews was developed', Journal of Clinical Epidemiology, 69, pp.225–234. doi:10.1016/j.jclinepi.2015.06.005.
Whiting, P.F., Tomlinson, E., Rutjes, A.W.S., Davenport, C.F., Mallett, S., Takwoingi, Y., Deeks, J.J. and the QUADAS-3 Group (2026) 'QUADAS-3: a revised tool for the quality assessment of diagnostic test accuracy studies', Annals of Internal Medicine, 179(4), pp.548–555. doi:10.7326/ANNALS-25-02104.
Xie, C., Kong, W., Pi, L., Qi, D., Yang, Y., Wang, B. and Liang, H. (2026) 'Performance of large language models in automated medical literature screening: a systematic review and meta-analysis', Journal of Evidence-Based Medicine, advance online publication. doi:10.1111/jebm.70166.
Yang, B., Mallett, S., Takwoingi, Y., Davenport, C.F., Hyde, C.J., Whiting, P.F., Deeks, J.J., Leeflang, M.M.G. and the QUADAS-C Group (2021) 'QUADAS-C: a tool for assessing risk of bias in comparative diagnostic accuracy studies', Annals of Internal Medicine, 174(11), pp.1592–1599. doi:10.7326/M21-2234.
Note on section 9 figures. The claimed "sensitivity for detecting high or low risk trials has ranged from 0.46 to 0.53 with kappa between 0.14 and 0.51" is corroborated by the two sources newly attached to it: Gandhi et al. (2026) report sensitivity 0.46 for high-risk detection and kappa as low as 0.14 (versus original review authors); Taneri (2025) reports sensitivity 53% for identifying high-risk studies and a weighted kappa of 0.51. Both figures in the range are now traceable to named, real sources and the document should cite them explicitly rather than leaving the claim unattributed.
Compiled for Oxon Advisory. Methodology v1.0, 13 September 2026. Licensed CC BY 4.0. Stan Gilmour, ORCID 0000-0002-6755-6842.