Introducing decision-machine-1: a language model for decisions, not prose
Extract

Extract every value from your documents.

With the proof attached. Send a file and the shape you want back. Every field returns cited to the page it came from, or declared absent. Never guessed.

How it works

Your schema shapes the output.

One call. A file, the shape you want, and a switch for proof. What comes back is JSON in exactly that shape.

  1. 01

    Send a file and a shape

    A URL, or the id of a file you already uploaded. Then a JSON Schema, or the name of a saved action so the request is two fields long. Up to 100 documents fit in one call.

    POST /v1/extractschema | actionmax 100 files
  2. 02

    We read it first

    Born-digital pages are read directly. Scans go through OCR. Recordings are transcribed. You never call parse yourself. Extract runs it on the way in.

    processing: autoocrtranscribe
  3. 03

    Read each field for what it is

    Mark a field as money, a date, a phone number, or an address. $1,234.56 comes back as 1234.56. A date comes back in ISO. A date nobody can settle, like 03/04/2026, comes back exactly as written instead of guessed into the wrong month.

    x-cr-kind: currencyformat: date13 field types
  4. 04

    Ground every field

    Turn citations on and each value points at its page and the box on that page, or at a timecode in a recording. A field with no evidence comes back null, with an entry that says so.

    citations: truebboxnotFound
  5. 05

    Wait, or don't

    The call holds for up to 120 seconds. A longer run comes back as a 202 with a run id, never as a timeout. Poll it, or let a webhook bring the result to you.

    ?wait=120202 + run idwebhook
The demo

Watch an invoice check its own math.

Fields come off any invoice, then the page checks the invoice against itself. Do the lines add up to the subtotal? Does subtotal plus tax make the total? Basic arithmetic still catches extraction errors.

The shape comes from a schema, a saved action, or a sentence of hints. The demo runs all three.Open in a new tab
What it reads

Clean data from any document.

You choose the fields. Extract fills them, whatever the file looked like when it arrived.

Invoices
PDF, Word, Excel, or a scan. Vendor, dates, line items, and totals. A list of ten line items comes back with the evidence for each one.
Contracts
Parties, renewal dates, notice periods. A term written across two clauses is cited at both, because that is where it lives.
Claim forms
Scanned, crooked, stapled to a cover letter. One sentence of instructions tells it to ignore the cover letter.
Client calls
A recording is transcribed, then cited by timecode. Seek to that second and hear the value being said.
Statements and ledgers
Ask for rows instead of one record and a twelve-page statement comes back as one row per transaction, each row cited.
Related files
An application and a bank statement. A contract and its amendment. Send them together and get one record back.
Grounding

Every value shows where it came from.

A citation carries the file, the page, and the box on that page. The value you doubt takes one click to check. Checked in one glance. Never invented.

invoice-4471.pdf1needs a lookSelect a field to see it in the document.
Page 1 of 1
Invoice INV-4471, Northwind Supply Co.

Northwind Supply Co.

Subtotal $13,518.52. Tax (8%) $1,081.48.
1

Extracted fields

TypeScript
const run = await client.extract({
  file: { url: "https://example.com/invoice-4471.pdf" },
  schema: invoiceSchema,
  citations: true, // off by default; ask for the evidence
});

// Every field is cited, or declared absent. Nothing in between.
for (const [field, cited] of Object.entries(run.output?.citations ?? {})) {
  for (const c of cited) {
    if (c.notFound) console.log(field, "not in the document");
    else console.log(field, c.page, c.bbox, c.text, c.confidence);
  }
}
Accuracy

It never guesses.

Most extraction fails quietly, with a plausible number in the wrong field. These are the places Extract is built to fail loudly instead.

  • Missing values
    What people assume The model fills the gap with something that looks right.
    What happens A value that is not in the document comes back null, with a notFound entry beside it. Your code can tell absent from empty.
  • Unsupported values
    What people assume If a field got filled, it must have come from somewhere.
    What happens A field filled without evidence gets a deterministic second pass. It either finds the citation or declares the value absent. Nothing is left in between.
  • Confidence
    What people assume One score for the whole document, and a shrug.
    What happens Every citation is scored on its own, against the exact quote it points at. Turn on the judge and a second pass re-scores each value against that evidence.
  • Ambiguous formats
    What people assume 03/04/2026 gets normalized to whichever month the model prefers.
    What happens When no reading can settle it, the value comes back exactly as written. A person decides, not a coin flip.
From first try to production

Start loose, then pin it down.

Three ways to say what you want back. They are one ramp, not three products.

Hints, to explore
Skip the schema and write one sentence about the document. We infer the fields and extract them in the same call. Two runs can pick different field names, so this is for finding the shape, not for shipping it.
A saved action, to scale
Save the schema and instructions once. Ops owns the shape, every service calls it by name, and the same action is editable in the workspace without a deploy.
Many documents, one record

Many documents become one record.

Send related files with unit across_documents. Each one is extracted on its own, then the results fold into one record by fixed rules. A value is taken whole, with its citation, from exactly one document.

  1. 01application.pdfsent first
  2. 02bank-statement.pdf
  3. 03pay-stub.pdf

Account number

account_number
application.pdfnot foundnot found
bank-statement.pdf•••• 4471citedkept
pay-stub.pdfnot foundnot found

Rule appliedA real value beats an empty one. The application left the box blank, so it had nothing to contribute and was not asked to guess.

Monthly income

monthly_income
application.pdf$8,200uncited
bank-statement.pdfnot foundnot found
pay-stub.pdf$7,940citedkept

Rule appliedThe strongest citation wins. The pay stub pointed at the line it came from. The application only asserted a number.

Home address

home_address
application.pdf88 Elm St, Apt 4citedkept
bank-statement.pdf88 Elm Streetcited
pay-stub.pdfnot foundnot found

Rule appliedBoth cited the address just as plainly, so the tie went to the earlier file in the request. Upload order decided this one.

Never blended, never averaged, and lists are never merged across files. Same input, same answer, every time. Because ties go to the earlier file, the order you send them in is a choice. Put the document that should win first.

For the engineers in the room

Built for the production path.

Nothing piles up. Nothing times out.

Sync, then async
The call holds up to 120 seconds, then hands back a run id. A large scan degrades into a handle, never into an error.
Webhooks
Name a webhook on the request and the finished run comes to you. No polling loop to own.
Idempotency keys
Send an idempotency key and a retry returns the original run. Nothing gets extracted twice.
Batch
One saved action over up to 100 files per call, one run per file. A bad URL in the list never sinks the rest.
Auto-expiry
Runs and their files are purged after 24 hours by default, seven days at most. Nothing accumulates in your organization.
A stable shape
One file or a hundred, the response carries one entry per document. Your parser never branches on how many you sent.
Part of the pipeline

One step in a bigger job.

The values Extract pulls out are what the next step needs. One CloudRaker pipeline, no tools stitched together.

Four ways in

Extract works wherever you do.

Same fields, same proof. Start in one, move to another.

Workspace
Your team, no code. Build the action in the app, edit it there, and run it without writing a line.
See the Workspace
CLI
paperwork extract extract, for scripts, CI, and backfills. JSON out, straight into jq.
See the CLI
Agents
One MCP connection for ChatGPT, Claude, or Codex. The agent extracts, reads the citations, and answers with the page instead of a paraphrase.
See Agents
Show me the code

One call returns cited JSON.

Send the file and the schema with citations on. JSON comes back in exactly that shape, with the evidence keyed by field.

import { CloudRakerClient } from "@cloudraker/api";

const client = new CloudRakerClient({ token: "YOUR_API_KEY" });

const run = await client.extract({
  file: { url: "https://example.com/invoice-4471.pdf" },
  citations: true, // off by default; ask for the evidence
  schema: {
    type: "object",
    properties: {
      vendor: { type: ["string", "null"] },
      due_date: { type: ["string", "null"], format: "date" },
      total: { type: ["number", "null"], "x-cr-kind": "currency" },
      po_number: { type: ["string", "null"] },
    },
  },
});

if (run.status === "processed") {
  console.log(run.output?.value);
  // { vendor: "Northwind Supply Co.", due_date: "2026-11-14",
  //   total: 14600, po_number: null }
  console.log(run.output?.citations?.po_number);
  // [{ fileId: "...", notFound: true }]
}
from cloudraker import CloudRaker

client = CloudRaker(token="YOUR_API_KEY")

run = client.extract(
    file={"url": "https://example.com/invoice-4471.pdf"},
    citations=True,  # off by default; ask for the evidence
    schema={
        "type": "object",
        "properties": {
            "vendor": {"type": ["string", "null"]},
            "due_date": {"type": ["string", "null"], "format": "date"},
            "total": {"type": ["number", "null"], "x-cr-kind": "currency"},
            "po_number": {"type": ["string", "null"]},
        },
    },
)

if run.status == "processed":
    print(run.output.value)
    # {"vendor": "Northwind Supply Co.", "due_date": "2026-11-14",
    #  "total": 14600.0, "po_number": None}
    print(run.output.citations["po_number"])
    # [{"fileId": "...", "notFound": True}]
Where to go next

Extract is one of many.

Same foundation under each one. Pick the door that fits the job.

Lease extraction
A lease, its abstract, and a fact sheet that disagree. See one cited record come back, and which document won each field.
See the use case
Get started

Data you can prove, from the first call.

Point Extract at your worst document with citations on. Read what comes back, then open the page behind any value you doubt.