Account number
account_numberRule appliedA real value beats an empty one. The application left the box blank, so it had nothing to contribute and was not asked to guess.
With the proof attached. Send a file and the shape you want back. Every field returns cited to the page it came from, or declared absent. Never guessed.
One call. A file, the shape you want, and a switch for proof. What comes back is JSON in exactly that shape.
A URL, or the id of a file you already uploaded. Then a JSON Schema, or the name of a saved action so the request is two fields long. Up to 100 documents fit in one call.
Born-digital pages are read directly. Scans go through OCR. Recordings are transcribed. You never call parse yourself. Extract runs it on the way in.
Mark a field as money, a date, a phone number, or an address. $1,234.56 comes back as 1234.56. A date comes back in ISO. A date nobody can settle, like 03/04/2026, comes back exactly as written instead of guessed into the wrong month.
Turn citations on and each value points at its page and the box on that page, or at a timecode in a recording. A field with no evidence comes back null, with an entry that says so.
The call holds for up to 120 seconds. A longer run comes back as a 202 with a run id, never as a timeout. Poll it, or let a webhook bring the result to you.
You choose the fields. Extract fills them, whatever the file looked like when it arrived.
A citation carries the file, the page, and the box on that page. The value you doubt takes one click to check. Checked in one glance. Never invented.
Extracted fields
const run = await client.extract({
file: { url: "https://example.com/invoice-4471.pdf" },
schema: invoiceSchema,
citations: true, // off by default; ask for the evidence
});
// Every field is cited, or declared absent. Nothing in between.
for (const [field, cited] of Object.entries(run.output?.citations ?? {})) {
for (const c of cited) {
if (c.notFound) console.log(field, "not in the document");
else console.log(field, c.page, c.bbox, c.text, c.confidence);
}
}Most extraction fails quietly, with a plausible number in the wrong field. These are the places Extract is built to fail loudly instead.
Three ways to say what you want back. They are one ramp, not three products.
Send related files with unit across_documents. Each one is extracted on its own, then the results fold into one record by fixed rules. A value is taken whole, with its citation, from exactly one document.
Rule appliedA real value beats an empty one. The application left the box blank, so it had nothing to contribute and was not asked to guess.
Rule appliedThe strongest citation wins. The pay stub pointed at the line it came from. The application only asserted a number.
Rule appliedBoth cited the address just as plainly, so the tie went to the earlier file in the request. Upload order decided this one.
Never blended, never averaged, and lists are never merged across files. Same input, same answer, every time. Because ties go to the earlier file, the order you send them in is a choice. Put the document that should win first.
Nothing piles up. Nothing times out.
The values Extract pulls out are what the next step needs. One CloudRaker pipeline, no tools stitched together.
Same fields, same proof. Start in one, move to another.
Send the file and the schema with citations on. JSON comes back in exactly that shape, with the evidence keyed by field.
import { CloudRakerClient } from "@cloudraker/api";
const client = new CloudRakerClient({ token: "YOUR_API_KEY" });
const run = await client.extract({
file: { url: "https://example.com/invoice-4471.pdf" },
citations: true, // off by default; ask for the evidence
schema: {
type: "object",
properties: {
vendor: { type: ["string", "null"] },
due_date: { type: ["string", "null"], format: "date" },
total: { type: ["number", "null"], "x-cr-kind": "currency" },
po_number: { type: ["string", "null"] },
},
},
});
if (run.status === "processed") {
console.log(run.output?.value);
// { vendor: "Northwind Supply Co.", due_date: "2026-11-14",
// total: 14600, po_number: null }
console.log(run.output?.citations?.po_number);
// [{ fileId: "...", notFound: true }]
}from cloudraker import CloudRaker
client = CloudRaker(token="YOUR_API_KEY")
run = client.extract(
file={"url": "https://example.com/invoice-4471.pdf"},
citations=True, # off by default; ask for the evidence
schema={
"type": "object",
"properties": {
"vendor": {"type": ["string", "null"]},
"due_date": {"type": ["string", "null"], "format": "date"},
"total": {"type": ["number", "null"], "x-cr-kind": "currency"},
"po_number": {"type": ["string", "null"]},
},
},
)
if run.status == "processed":
print(run.output.value)
# {"vendor": "Northwind Supply Co.", "due_date": "2026-11-14",
# "total": 14600.0, "po_number": None}
print(run.output.citations["po_number"])
# [{"fileId": "...", "notFound": True}]Same foundation under each one. Pick the door that fits the job.
Point Extract at your worst document with citations on. Read what comes back, then open the page behind any value you doubt.