ProGen

One engine. Every step of the document lifecycle.

From the moment a document lands to the moment its data is usable elsewhere, here's how ProGen handles each step.

01 · Ingestion

Feed it anything.

Scanned images, multi-page PDFs, or unstructured files, ProGen ingests it all in one pipeline, no pre-sorting required.

  • Images and multi-page PDFs in one upload
  • Any unstructured document type, no template setup
  • Bulk upload — drop one file or hundreds

How it works

UploadAuto-detect typeQueue for reading

Bulk upload

IMGinvoice_scan.jpg
PDFpo_batch.pdf
PDFepc_report.pdf
IMGreceipt_014.png
+ drop more files, any format

02 · Reading

Reads the layout, not just the text.

ProGen keeps table and column structure intact, so line items map correctly to prices instead of turning into a wall of unstructured text.

  • Table and column structure preserved
  • Line items stay linked to the right price/quantity
  • Works on scanned and digital documents alike

How it works

OCR passLayout mappingStructure preserved

Layout preserved

ItemQtyPrice

03 · Understanding

Pulls the fields that matter.

A locally-hosted LLM extracts vendor, amounts, invoice numbers, and line items, with a deterministic fallback so extraction never silently fails.

  • Vendor, amount, invoice number, and line-item extraction
  • Locally-hosted LLM, your documents don't leave your infrastructure
  • Deterministic fallback if AI inference fails

How it works

Field detectionLLM extractionFallback check

Field extraction

vendor_name
Summit Steel & Supply
invoice_number
INV-20482
total_amount
$18,420.00
Deterministic fallback active if AI inference is uncertain

04 · Generation

Generates documents, too.

Same engine, reverse direction, produce PDFs, invoices, and formatted documents straight from structured data.

  • PDF and invoice generation from structured data
  • Custom layouts and templates
  • One engine for both reading and producing documents

How it works

Structured inputTemplate appliedDocument output

Document generation

Coming soon

05 · Storage & Connect

Everything searchable, everything connected.

Processed data is stored with a reference back to the original document, searchable, reviewable, and ready to push into the systems you already use.

  • Cloud storage with original-document reference
  • REST API for external systems to push and pull documents/results
  • Works toward ERP (SAP, Tally, Zoho Books), email ingestion, Slack/Teams alerts, and webhooks

How it works

Store with referenceIndex for searchPush via API/webhook

Connected everywhere

Document
ERP
Email
Slack
Webhook