Skip to content
PDFSifter
Built for research teams

Extract structured data from papers with AI.

Show PDFSifter’s AI three examples of what you need. It reads every PDF in the folder, pulls out those values in minutes, and shows you the exact quote behind each one.

Free plan: 200 pages a month · No credit card needed

Trial facts · 48 papers · 47 done

How it works

From a folder of PDFs to a table you can trust

No prompts to write and no rules to maintain. You describe the data once, check a few examples, and the AI does the rest.

  1. 1

    Describe what you need

    Write a short brief in plain English and list your fields: text, numbers, percents or dates.

    Sample size · numberFollow-up · text
  2. 2

    Train the AI on three examples

    Add PDFs with the answers you expect. “Label in PDFSifter” drafts them for you to check. The AI writes its own instructions and shows how often it matched you.

    Measured on your examples
  3. 3

    Run, review, export

    Drop in a folder or a zip. Progress updates live, uncertain values are queued for review, and results export in one click.

    CSVExcelJSONPDF report
See every step →

Features

AI extraction you can check

Every value shows where it came from. Checking what the AI found shouldn't mean re-reading every paper: PDFSifter tells you which values to trust and which ones to look at.

  • Quote, page and confidence

    Each value links to the sentence it came from and the page it’s on, with a confidence score from 0 to 100.

  • Review only what’s uncertain

    Values below your threshold are flagged. Step through them with J and K, edit with E, accept with Enter.

  • Scans and charts, too

    Pages without a text layer, like scans and photos of charts, are read with OCR. Nothing to switch on.

  • Retry only what failed

    If a few documents fail, run just those again in the same job. You only pay for their pages.

  • Gets better as you review

    Turn a reviewed document into a training example in one click. Retrain, then compare every version in History.

  • Exports that keep their types

    Excel keeps numbers, percents and dates typed and highlights values still waiting for review.

For teams and developers

Fits the way your lab already works

Invite your team, give everyone the right role, and plug PDFSifter into your own pipeline when you’re ready.

  • Roles and permissions for every workspace
  • An audit log of who changed what, and when
  • Search names, values and the full text of every PDF
  • Resumable uploads for big folders and zips
  • API keys and signed webhooks on Lab and Research
# 1. Upload PDFs with any tus client
tus upload papers/*.pdf  →  /files/

# 2. Start a job with a trained template
POST /api/v1/workspace/jobs
Authorization: Bearer sk_live_…

# 3. Hear back when it's done
←  job.finished
   signed with HMAC-SHA256

# 4. Read the records, or export them
CSV · Excel · JSON · PDF report

Security and privacy

Your papers stay yours

Unpublished manuscripts, patient reports, internal studies: PDFSifter is built to handle documents you can’t afford to leak.

  • Files stored in the EU

    Uploaded PDFs live in EU-jurisdiction storage.

  • Zero data retention for AI

    Model requests go only to providers that don’t keep your data.

  • Deleted on your schedule

    Source PDFs are removed after your retention period. Results and search keep working.

  • Two-step verification

    Authenticator-app codes and one-time recovery codes for every account.

  • Sign in the way you already do

    Google, Microsoft or ORCID, next to email and password.

  • A 30-day safety net

    Deleted jobs and templates wait in the Trash for 30 days before they’re gone.

Read how we handle your data →

Pricing

Start free. Pay when it’s part of your week.

Priced by the page, so a 3-page abstract and a 40-page report cost what they should.

Free

Try PDFSifter on a handful of papers.

$0forever

No credit card needed


  • 200 pages a month
  • 1 seat
  • 2 templates
  • Train on PDFs of up to 20 pages
  • 1 GB of stored PDFs
  • Source PDFs kept 30 days

Starter

For one researcher working through a review.

$19per month

or $180 a year


  • 2,000 pages a month
  • 2 seats
  • 5 templates
  • Train on PDFs of up to 50 pages
  • 5 GB of stored PDFs
  • Source PDFs kept 60 days
  • Extra pages with credits, $0.012 each
Most popular

Lab

For a small team extracting every week.

$79per month

or $756 a year


  • 10,000 pages a month
  • 5 seats
  • Unlimited templates
  • Train on PDFs of up to 120 pages
  • API access and webhooks
  • 20 GB of stored PDFs
  • Source PDFs kept 90 days
  • Extra pages with credits, $0.009 each

Research

For labs running extraction at scale.

$300per month

or $2,880 a year


  • 50,000 pages a month
  • 10 seats
  • Unlimited templates
  • Train on PDFs of up to 300 pages
  • API access and webhooks
  • 50 GB of stored PDFs
  • Source PDFs kept 90 days
  • Extra pages with credits, $0.004 each

Prices in US dollars. A busy month? Top up with credits instead of upgrading. Need more seats or an invoice? Talk to us.

FAQ

Questions researchers ask us

Something else? The in-app Help page has step-by-step guides, or you can write to us.

Do I need to write prompts?
No. You write a short brief, list the fields and add a few examples. PDFSifter’s AI writes the extraction instructions itself and tests them against your examples. You can still edit its summary and rules if you want to.
How accurate is it?
Every template shows its measured accuracy on your own examples. Every value has a confidence score, and anything below your review threshold (80 by default) is flagged for a person to check.
What kinds of PDFs work?
Papers, reports and forms, with one record per PDF or many. Scanned pages are read with OCR. Password-protected files need to be unlocked first.
What happens to my files?
They’re stored in the EU and deleted after your retention period: 30 days on Free, up to 90 on Lab and Research. The values, quotes and text stay, so your results and search keep working.
What counts as a page?
Each page of each PDF you run. If a document fails and you retry it, only that document’s pages are counted again.

Turn your next stack of papers into a table

Your first 200 pages every month are free. Set up a template in about 15 minutes.