Screen arXiv papers for a systematic literature review

By General Input

Pull candidates from arXiv into a screening queue, call Include, Exclude or Maybe with an AI first pass, and export the included set to a sheet.

Integrations

  • arXiv
  • Google Sheets

Type

App

Categories

  • Operations

Build me a literature screening workbench I can work out of for weeks at a time while I run a systematic review. arXiv is read only, so the app itself is the system of record: reviews, candidate papers and every screening decision persist in the app, and Google Sheets is only the export target at the end.

Review setup. The home screen lists my reviews with a progress bar on each, plus a New review form. A review has a topic name, a free text inclusion criteria block, a free text exclusion criteria block, an editable list of named exclusion reasons (seed it with off topic, wrong method, wrong population, not an empirical result, superseded by a later version, and other), and one or more sources. A source is either a search string run through arXiv Search Papers, using the Lucene style field prefixes such as ti:, abs:, au: and cat: with AND, OR and ANDNOT, or a subject category run through List Recent Papers in a Category, such as cs.LG or q-bio.NC. Each source can carry an earliest submission date and a cap on how many candidates it pulls. Opening a review takes me into the workbench for that review.

Pulling candidates. A Fetch candidates button on the review runs every source and stores each paper as a candidate row keyed by its arXiv ID, capturing title, authors, abstract, primary subject category, all categories, submission and last updated dates, and the link to the abstract page. Page through results with the start offset and page size rather than asking for everything at once, and strip the version suffix from the ID before comparing so a paper matched by two sources, or matched again on a later refresh, only ever enters the queue once. arXiv asks for no more than one request every three seconds on a single connection, so space the requests out, send a descriptive user agent, back off and retry when it responds with a rate limit, and cache every page you pull so the app never re-fetches a paper it already holds. Responses come back as Atom XML rather than JSON, so parse them accordingly and watch for the case where a valid looking feed contains a single error entry. Run the fetch as a background job and report how many papers were pulled, how many were new and how many were duplicates. I can re-run it later to top the queue up without touching decisions already made.

AI pre-read. When new candidates land, use AI Generation to read each abstract against that review's inclusion and exclusion criteria and store a suggested call of include, exclude or maybe, a one line rationale, and, where the suggestion is exclude, whichever of my exclusion reasons it best matches. Store the suggestion on the candidate so it is produced once per paper and never regenerated when I open the screening screen. Make the pre-read something I can switch off per review.

The screening screen. This is the main surface and it shows one paper at a time from the signed in reviewer's queue: title, authors, primary subject category, the full abstract, the submission date, and a link out to the arXiv abstract page. Under it, show the AI suggestion and its rationale as a suggestion I can accept with one click or simply ignore. Three decision buttons: Include, Exclude and Maybe. Excluding requires a reason chosen from the review's exclusion reason list, with an optional free text note; Include and Maybe take an optional note. Give me keyboard shortcuts for the three calls and for picking a reason. Saving a decision advances straight to the next undecided paper, with an undo on the decision I just made, and a counter showing how many are left in my queue.

Blind, per reviewer decisions. Decisions persist per reviewer, so a second person opening the same review gets the same queue and screens it independently. While screening, never surface another reviewer's call, reason or note for the paper on screen. Each reviewer's queue contains only the papers that reviewer has not yet decided on.

Conflicts tab. List every paper where two or more reviewers have both decided and their calls differ, counting any difference between include, exclude and maybe as a disagreement. Each row shows the paper alongside each reviewer's call, reason and note side by side, and a resolve control that records a final agreed decision, a resolution note and who resolved it. The final agreed decision is what counts downstream. Add a filter for unresolved conflicts and a count badge on the tab so I can see at a glance what is still open.

Summary view. Keep running counts for the review: papers found, duplicates removed, screened, included, excluded, maybes still open, and conflicts open versus resolved, plus a breakdown of exclusions by reason and a per reviewer row showing screened, remaining and how often that reviewer agreed with the other one. These are the numbers a PRISMA style flow diagram is built from, so make them easy to read off and copy.

Export. An Export button writes the included set to a Google Sheet using Append Values against a spreadsheet ID and tab name I choose. One row per included paper carrying the arXiv ID, title, authors, primary category, submission date, abstract link, final decision, each reviewer's call and reason, the resolution note where there was a conflict, the AI suggested call, and the date it was screened. Write a header row first if the tab is empty, append rather than overwrite so an earlier export survives, use the user entered input mode so links stay clickable, and tell me how many rows were written. Give me the option to export the full screened set instead of only the included papers, with excluded rows carrying their exclusion reason.

General behaviour. I can have several reviews running at once and switch between them. Nothing runs on a schedule: candidates are pulled when I press the button, and everything I have done survives closing the tab, because a real review runs over weeks.

Related prompts

Explore more prompts
Call overdue Xero customers with an AI collections agentLocal listing health board for every location you manageLet support send one-off Loops emails without an engineerStop cold emails to anyone with a live deal in PipedriveiMessage campaign console with pre-flight checks and delivery boardLinkedIn Ads budget pacing dashboard for every client accountFront desk appointment confirmation board for the next 3 daysGive your team Looker numbers without buying more seatsBuild audience segments from product usage and push to LoopsTurn the people who engage with your posts into Pipedrive leads