The one rule to start with
An AI chatbot is not a search engine. Ask it to "find papers on" a topic and it will often invent plausible-looking references that do not exist.
The safe pattern is the opposite of how most people first use these tools. Use proper search tools and databases to find papers; then use an AI to help you work through papers you have actually got in front of you. Find with tools built for finding; synthesise with AI on verified material.
The current evidence supports exactly this split. A 2025 review in Research Synthesis Methods (Clark et al.) found that while AI can assist humans in most steps of a review, it cannot yet be relied on to run the process on its own. The strongest approach pairs traditional database searching for completeness with AI for speed.
What AI does well here
- Screening at speed. Tools that learn from your include/exclude decisions can prioritise the likely-relevant papers, cutting screening time substantially.
- Pulling information into a table. Extracting study details (population, intervention, outcome) into a consistent format, with each claim linked back to the source text.
- Plain-language synthesis of papers you provide. Paste in verified abstracts and ask for a structured summary.
- Scoping and mapping. Seeing how papers connect, and getting a quick read on whether a body of evidence leans one way.
Across a review, these tools typically save in the order of 30 to 40 per cent of total time, with the biggest gains in screening and data extraction. For small, stretched research teams that is meaningful.
The tools, by what they are for
You do not need all of these. This is a map so you can pick the right one for a task.
Finding and discovery
- Traditional databases (PubMed, Embase, Scopus, Web of Science) remain the primary way to search comprehensively. Nothing below replaces them.
- Semantic Scholar: free, over 200 million papers, AI summaries and citation markers.
- Research Rabbit / Connected Papers / Litmaps: visual citation mapping to discover related work.
Screening
- Covidence: Melbourne-based, Cochrane-endorsed, the standard platform for PRISMA-compliant reviews in Australia; many universities hold licences (check your library).
- Rayyan: free AI-assisted screening that learns from your decisions.
- ASReview: free, open-source, fully auditable; good when reproducibility matters.
Data extraction
- Elicit: strong at pulling data into tables with sentence-level links to the source, reported at high extraction accuracy. Best used as a second extractor alongside a person.
Quick evidence checks
- Consensus: shows the balance of studies supporting or opposing a claim; useful for a fast read, not for a formal review.
- Perplexity (Deep Research): produces cited multi-page summaries quickly; good for scoping, not for formal reviews.
The catch with AI search tools
The most important number on this page: when one leading AI tool was tested as a primary search method, its search sensitivity averaged around 40 per cent, against around 95 per cent for traditional database searching (Lau and Golder, 2025, Cochrane Evidence Synthesis and Methods). In plain terms, used on its own to search, it missed most of the relevant papers.
So AI search tools supplement, never replace a proper search. Use them to catch extra papers and to speed up screening and extraction, on top of a comprehensive database search, not instead of one. Responsible vendors take the same view: one screening platform deliberately withheld an automatic-exclusion feature after finding it dropped recall too far.
If a review missing relevant studies would matter, and in health research it usually does, the database search is not optional. AI sits alongside it.
A note on Indigenous and community knowledge
Published databases capture a particular slice of knowledge. Much Aboriginal and Torres Strait Islander knowledge, and a lot of community-held and grey-literature evidence, was never published in indexed journals, or was written about communities by outsiders. An AI summarising "the literature" inherits those gaps and silences without flagging them. Treat an AI synthesis as a summary of what is in the indexed record, not a summary of what is known. The data sovereignty and ICIP page covers the wider questions this raises.
Documenting AI use in a review
If you use AI anywhere in an evidence review, record it so your method is transparent: the tool, its version, the date, and what you used it for. Emerging reporting guidance (the PRISMA-trAIce checklist and RAISE guidelines) exists for exactly this. The NHMRC treats using AI to search, summarise or code literature as routine research use, but transparency about how you used it is still expected.
Glossary
- Systematic review
- a structured, comprehensive review of all the evidence on a question, following a documented method.
- Screening
- deciding which search results are relevant enough to include.
- Data extraction
- pulling the key details out of each included study into a consistent format.
- Sensitivity (of a search)
- the share of relevant papers a search actually finds; low sensitivity means missed studies.
- Recall
- another word for how much of the relevant material a tool retrieves.
- PRISMA
- the standard reporting method for systematic reviews; PRISMA-trAIce extends it to AI use.
- Grey literature
- reports, theses and material not published in indexed journals; often where community evidence sits.
Take it with you: a one-page cheat sheet of the tools-by-task and the search rule is available to download from the button at the top of this page.
