Embedded Document Processing
Accept complex documents from your customers and return clean, structured data through one configurable service. Your product sends a document in any format or language, and Staple returns the fields, tables, and per-field confidence scores your workflow needs, with a verifiable trust record attached to the result. Your customers never leave your interface, and you never build extraction, language coverage, or a review workflow yourself.
The Reference Flow
A document enters through the API or a channel you control, then is pre-processed, classified, split, and extracted, and returned as structured data with a per-field confidence score. A webhook fires on completion so your integration knows the moment the data is final. The reading is context-based: Staple interprets what a field means rather than where it sits on the page, which is why it holds up on layouts it has never seen before instead of breaking the first time a supplier changes a template. See the underlying data extraction engine for how that works.
What It Reads: Channels, Formats, And Languages
• Channels: the API, plus email, SFTP, shared drives, portals, and file sync, so documents reach Staple however your customers already send them.
• Formats: PDF, images, spreadsheets, forms, scans, photographs, dot-matrix prints, and mixed multi-document packs where several documents arrive in a single file.
• 300+ languages with native Asian script support, plus handwriting and rubber stamps, so a branch in any market runs on the same engine.
Within a document, Staple captures header fields such as vendor, invoice number, and date, line items in tables, and details a template-based tool ignores, including QR codes, signatures, and stamps, along with tax-authority metadata where it is present.
Complex Tables, Handled Correctly
Line items are where most extraction tools fall down. Staple's intelligent tables handle irregular, hierarchical, multilingual, and merged tables, cleaning and standardizing hundreds of line items presented in inconsistent formats into one predictable structure. Mixed packs are separated first by classification and splitting, which recognizes an invoice, a delivery order, and a remittance in the same file and keeps the relationships between related documents intact, so what you receive is already sorted rather than a single blob you have to untangle.
Configure With Model Builder And Intelligent Feedback
A model in Staple is a natural-language, prompt-driven blueprint for how a document should be understood: it defines what data to capture, not merely where it appears on a page. Model Builder lets you define fields with drag and drop, editable or locked at the model level, with data types set per field and per line-item header. Enrichment and inference are part of the model, so a value can be derived from other captured data, for example resolving the currency by reading the supplier's country, or interpreting a company name from a logo on a receipt. Intelligent Feedback turns a manual correction into reusable logic through a natural-language instruction, so training happens on the go, without code, and the next similar document is handled correctly. You decide whether models are managed centrally on your customers' behalf or exposed to them through no-code configuration.
The Review Loop
Queue rules handle validation, set-values, and routing after scanning. Any field that lands below the configurable confidence threshold, 0.9 by default, is routed to review rather than passed downstream as if it were certain, so a low-confidence value never quietly becomes a posted number. You can surface Staple's review step to your own users or build your own review experience on top of the API, and the final webhook fires only once the document is resolved.
Structured Data That Carries Its Own Trust
What comes back is more than fields. Every result can carry a Metastructured Data envelope, so the trust travels with the data rather than sitting in a separate log your customers have to reconcile later. Each field is accompanied by its provenance, which model read it and when, its proof, the confidence and any anomaly flags, and its policy, the residency and jurisdiction markers that follow it. The envelope is cryptographically signed, so a downstream system or an auditor can confirm the data was not altered after Staple produced it, without contacting Staple. That is what turns extraction output into evidence your customers can act on and defend.
Proven In Production
A European KYC technology provider onboarded new document types across multiple languages with no code, adding coverage without an engineering project each time. A global FMCG brand processed supplier documents from 5,600 stores in four APAC languages, including dot-matrix prints a legacy OCR tool could not read, at 99.6% accuracy. Across deployments, Staple processes more than 10 million documents a year for Fortune 100 companies and global financial institutions in 60 countries.
Scope And Next Step
Packaging is usage-based and scoped per deployment, and you decide whether models are managed centrally or configured per tenant. Discuss your document types and tenant model with the platform team, or view the API documentation to see the request and response shapes.
Related Topics
See Staple process your documents
Book a 30-minute demo with a document processing specialist.
Not ready yet?
Take your time to decide.