Preprint

Proposed framework sets stricter rules for AI in evidence reviews

Preprint: ARISMA treats AI as an inspected, benchmarked, logged and reversible assistant, with humans retaining control over consequential review decisions.

A preprint proposes a framework for using artificial intelligence (AI) in evidence reviews while keeping scientific responsibility with people. Called ARISMA, it treats AI as an inspected, benchmarked, logged and reversible assistant rather than an autonomous reviewer.

That approach sets a clear hierarchy. ARISMA distinguishes assistive, adjudicative, meaning involved in settling a review decision, and generative uses of AI. As an AI task carries more potential to shape scientific judgment, the framework requires stricter validation and human oversight; generative output is treated only as draft material grounded in verified data.

The proposal is about governance and reporting. It does not report a prospective evaluation of whether ARISMA is effective. Instead, it offers a structure for making AI use inspectable, documented and accountable.

A process, not a promise

ARISMA was developed through literature-informed design, combining established evidence-synthesis guidance, literature on AI-assisted review and structured expert consultation. The consultation involved 45- to 60-minute video discussions with 21 researchers and information specialists experienced in evidence synthesis. Participants were selected purposively for methodological expertise, and the feedback was formative rather than statistically representative.

The framework organizes evidence synthesis into six linked phases: framing; protocolization, or setting the review plan; retrieval; selection and enrichment; evidence structuring; and synthesis and reporting. The sequence ties AI use to a particular point in the review and to the checks required there.

Consultation feedback tightened the operating rules. It expanded calibration and pilot-testing requirements, added records of the model and prompts used, clarified when an AI system should be revised or disabled, and strengthened human review responsibilities.

Rules at the points that matter

Screening is one area where those rules become concrete. AI use must be calibrated for the task against a pilot set labeled by humans and containing inclusions, likely exclusions and ambiguous borderline records. For reviews where an erroneous exclusion could have especially serious consequences, AI may assist as a second reviewer, prefilter or prioritization system. It should not be the sole excluding reviewer without domain-specific validation and an explicit justification.

Data extraction follows the same principle. AI may prepopulate extraction forms, but every numeric datum used in aggregation or synthesis must be checked against the source by a human reviewer.

At the reporting stage, the paper turns ARISMA into submission-ready reporting and validation instruments. Its validation matrix links every workflow step to a potential failure, the validation required before an AI output can be trusted and a minimum audit trail. Together, those tools are intended to show not only where AI was used, but what could go wrong, what was checked and what record remains for later review.

Testing still lies ahead

The framework's own evidence base comes with a clear caveat. The consultation was purposive and formative, not a formal consensus exercise or an effectiveness evaluation. Its role was to refine the proposal, while broader prospective evaluation remains necessary.

The manuscript states that it was accepted for AGENTICS 2026 presentation and Springer proceedings publication. It is the authors' manuscript version, with a final authenticated publication to come via Springer.

For review teams, editors and peer reviewers, the proposal offers a way to set boundaries around AI, record its use and preserve human control over consequential scientific judgments. On the evidence presented, however, ARISMA is a governance foundation rather than evidence that review outcomes improve. That question remains open until the framework is tested prospectively.

Paper data and sources

Original title: ARISMA: Guidelines for AI- and LLM-Assisted Systematic Reviews, Scoping Reviews, and Mapping Studies
Authors: Mahyar Tourchi Moghaddam, Mina Alipour
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-25
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.