Skip to content

Cascade Extraction: Deterministic Invoice Intake from First Scan to Field Guarantee

Today, Helpifyr's document pipeline crossed a threshold: supplier invoices can now be ingested, parsed, and enriched with deterministic, provenance-tracked field extraction-across diverse layouts and tax regimes. This is not just a smarter OCR, but a system-wide architectural shift that guarantees every extracted value is explainable, reproducible, and anchored to its source.

Jadda Helpifyr2 min read
Cascade Extraction: Deterministic Invoice Intake from First Scan to Field Guarantee

At a glance

9

merged changes

3

code projects involved

Most changes in

  • jhf-loom7
  • jhf-web1
  • jhf-spindle1

Underlined terms are explained: just hover or tap.

Picture a supplier invoice arriving by email: its layout is unfamiliar, its tax details buried among multilingual clutter, and the key fields-amount, invoice number, supplier address-are scattered and ambiguous. Historically, extracting this data required hand-tuned rules or brittle templates, with errors surfacing only after downstream processing. Today, that uncertainty ends. With a new cascade extraction engine, every field is captured through a deterministic, explainable chain-no matter the supplier, format, or country.

Why This Day Mattered

Operators and integrators can now onboard new suppliers and invoice formats without custom scripting or manual verification cycles. Developers gain a true contract: every extracted field is not just a best guess but comes with traceable, explainable provenance, making downstream automation and auditing reliable. For users, this means faster, error-free intake-no more waiting for a human to disambiguate fields or correct tax calculations.

The closed UTC day 2026-08-05 resolved into 9 merged PRs across 3 repos, led by jhf-loom (7), jhf-web (1), jhf-spindle (1).

What Actually Changed

The document pipeline now routes every invoice through a swappable OCR engine and a deterministic extraction cascade. This cascade applies a sequence of spatial, anchor-based, and exclusion-list extraction passes, layering heuristics and evidence trails for each field. Tax and net/gross relationships are derived with fail-closed logic, ensuring that missing or ambiguous data never silently propagates. The intake deduplication logic now scopes by document class, preventing cross-type collisions. Each extraction step is explainable and reproducible, with unified label assignment strategies and supplier-specific handling baked in.

Why It Holds Better Now

By grounding every field extraction in deterministic rules and explicit provenance, the system eliminates guesswork and silent failure modes. The multi-level cascade means new suppliers or layouts can be supported by adding or tuning extraction passes, not by rewriting brittle templates. Fail-closed tax derivation blocks partial or invalid data from entering the system, while the deduplication scoping prevents misclassification across document types. The unified evidence chain ensures that every value, from invoice number to tax rate, can be traced to its source-enabling confident automation and rapid onboarding.

Want to Know More?

How will this deterministic extraction pipeline enable real-time feedback and correction for users uploading new document types, and what new automation possibilities does it unlock for downstream finance workflows?

Terms in this post

fail-closed
Block when in doubt: if evidence is missing, the action does not run.
provenance
Proof of origin: where a piece of information or an artefact comes from.
PR
Pull request: a reviewed code change that gets merged into the project.
repo
Repository: a code project under version control.
operator
The person or team running the system.

What would this look like in your company?

A pilot shows it with a real process.

Request a pilot

More on Data and business logic

See all
Automatic Invoice Field Extraction: Supplier Intake Meets Real-Time EnrichmentData and business logic

2 min

Automatic Invoice Field Extraction: Supplier Intake Meets Real-Time Enrichment

Today, invoice intake on the Helpifyr / JaddaHelpifyr stack gained a new edge: invoices entering via supplier intake are now automatically enriched with structured data through live field extraction, closing the loop between document arrival and actionable, queryable records.

Read
No More Phantom Settlements: Reference Pinning Secures In-Flight DecisionsData and business logic

2 min

No More Phantom Settlements: Reference Pinning Secures In-Flight Decisions

Today's engineering work closes a subtle but critical loophole in the Helpifyr / JaddaHelpifyr stack: settlement decisions in progress can now survive service crashes and restarts without risk of double-processing or data drift. By introducing reference pinning for in-flight settlements, the platform now guarantees that every settlement either completes with the exact evidence it started with or is safely retried, never both.

Read