Ingestion & Pre-processing
Accept documents from fragmented channels and turn them into ready inputs, automatically.
Documents arrive from everywhere, in every condition. Staple receives them from any channel, cleans and normalizes each one, and records where it came from, all before classification and extraction begin.

Before a document can be processed, someone has to receive it, open it, and clean it up. That someone is usually a person.
Invoices come by email. Statements land in a shared drive. A supplier sends a photo of a delivery note from their phone. Scans arrive skewed, dark, or half-legible.
In most operations, a person collects all of this, renames it, straightens it, and drops it into the right folder before any automation can touch it.
That manual front end is slow, inconsistent, and invisible to any audit.
Staple removes it. Documents enter through any channel and come out the other side normalized, readable, and ready, with a record of exactly what arrived and what was done to it.
From any channel to processing-ready, before classification and extraction begin.
Input Channels
Receives from every channel your senders actually use
Email, SFTP, shared drives, messaging, file sync, and API. Documents arrive the way your suppliers, customers, and partners already send them, no portal to force on anyone, no new habit to enforce.

Format Coverage
Any format, any condition
PDFs, images, spreadsheets, forms, scans, and mixed document packs. Digital-native files, photographed documents, and dot-matrix prints are all accepted in the same flow.

Image Preparation
Cleans the document so extraction doesn't inherit the mess
Staple corrects orientation, de-skews crooked scans, enhances quality, and removes noise. A dark photo or a rotated fax becomes a clean, machine-readable input, so accuracy downstream doesn't drop because of how a document was captured.

Control Point
Every arrival is recorded before anything else happens
Staple logs the source, timestamp, and every pre-processing action taken on each document. Auditability starts at the front door, not after extraction, so even how a document was received and prepared is part of the record.

Output
What comes out
Normalized, readable files, ready for classification, forensic checks, and extraction. The messy, multi-channel intake becomes a clean, consistent queue for the rest of the first mile.

See how Staple handles your messiest intake.
Book a 30-minute demo. Bring documents from your real channels and formats, and we'll run them through ingestion and pre-processing live.
FAQ
What channels can Staple ingest documents from?
Staple receives documents through email, SFTP, shared drives, messaging, file sync, and API. Senders keep using whatever channel they already use, and Staple collects and prepares the documents automatically, with no manual gathering step in front.
What does pre-processing actually do to a document?
Pre-processing prepares a raw document for reliable extraction. Staple corrects orientation, de-skews crooked scans, enhances image quality, and reduces noise, so a photographed, faxed, or poorly scanned document becomes clean and machine-readable before any data is pulled from it.
Can Staple handle mixed document packs in a single upload?
Yes. A batch containing different document types and formats is accepted as-is. Pre-processing normalizes every file, and the Document Layer then classifies and splits the pack automatically, so there is no need to separate documents before sending them.
Is the ingestion step auditable?
Yes. Staple records the source, timestamp, and every pre-processing action for each document. This means auditability begins at the moment a document arrives, not only once its data is extracted, so document handling itself is part of the evidence trail.
What format does the output take?
Pre-processing produces normalized, readable files ready for classification, forensic tampering checks, and extraction. The output is a clean, consistent input for the rest of the pipeline, regardless of how varied the original documents were.
