Cascade Extraction: Deterministic Invoice Intake from First Scan to Field Guarantee
Today, Helpifyr's document pipeline crossed a threshold: supplier invoices can now be ingested, parsed, and enriched with deterministic, provenance-tracked field extraction-across diverse layouts and tax regimes. This is not just a smarter OCR, but a system-wide architectural shift that guarantees every extracted value is explainable, reproducible, and anchored to its source.

At a glance
9
merged changes
3
code projects involved
Most changes in
- jhf-loom7
- jhf-web1
- jhf-spindle1
Underlined terms are explained: just hover or tap.
Picture a supplier invoice arriving by email: its layout is unfamiliar, its tax details buried among multilingual clutter, and the key fields-amount, invoice number, supplier address-are scattered and ambiguous. Historically, extracting this data required hand-tuned rules or brittle templates, with errors surfacing only after downstream processing. Today, that uncertainty ends. With a new cascade extraction engine, every field is captured through a deterministic, explainable chain-no matter the supplier, format, or country.
01Why it matters
Why This Day Mattered
Operators and integrators can now onboard new suppliers and invoice formats without custom scripting or manual verification cycles. Developers gain a true contract: every extracted field is not just a best guess but comes with traceable, explainable provenance, making downstream automation and auditing reliable. For users, this means faster, error-free intake-no more waiting for a human to disambiguate fields or correct tax calculations.
The closed UTC day 2026-08-05 resolved into 9 merged PRs across 3 repos, led by jhf-loom (7), jhf-web (1), jhf-spindle (1).
02What changed
What Actually Changed
The document pipeline now routes every invoice through a swappable OCR engine and a deterministic extraction cascade. This cascade applies a sequence of spatial, anchor-based, and exclusion-list extraction passes, layering heuristics and evidence trails for each field. Tax and net/gross relationships are derived with fail-closed logic, ensuring that missing or ambiguous data never silently propagates. The intake deduplication logic now scopes by document class, preventing cross-type collisions. Each extraction step is explainable and reproducible, with unified label assignment strategies and supplier-specific handling baked in.
03Why it holds better now
Why It Holds Better Now
By grounding every field extraction in deterministic rules and explicit provenance, the system eliminates guesswork and silent failure modes. The multi-level cascade means new suppliers or layouts can be supported by adding or tuning extraction passes, not by rewriting brittle templates. Fail-closed tax derivation blocks partial or invalid data from entering the system, while the deduplication scoping prevents misclassification across document types. The unified evidence chain ensures that every value, from invoice number to tax rate, can be traced to its source-enabling confident automation and rapid onboarding.
04Food for thought
Want to Know More?
How will this deterministic extraction pipeline enable real-time feedback and correction for users uploading new document types, and what new automation possibilities does it unlock for downstream finance workflows?
Terms in this post
- fail-closed
- Block when in doubt: if evidence is missing, the action does not run.
- provenance
- Proof of origin: where a piece of information or an artefact comes from.
- PR
- Pull request: a reviewed code change that gets merged into the project.
- repo
- Repository: a code project under version control.
- operator
- The person or team running the system.
What would this look like in your company?
A pilot shows it with a real process.
More on Data and business logic
See all
Data and business logic2 min
Automatic Invoice Field Extraction: Supplier Intake Meets Real-Time Enrichment
Today, invoice intake on the Helpifyr / JaddaHelpifyr stack gained a new edge: invoices entering via supplier intake are now automatically enriched with structured data through live field extraction, closing the loop between document arrival and actionable, queryable records.
Read
Data and business logic2 min
Settling for Certainty: Reference Pinning Ends Phantom State in Decision Flows
A new reference pinning mechanism now guarantees that settlement decisions in flight cannot be lost or misapplied, eliminating phantom or orphaned states and making every decision traceable and replay-safe, even across restarts.
Read
Data and business logic2 min
No More Phantom Settlements: Reference Pinning Secures In-Flight Decisions
Today's engineering work closes a subtle but critical loophole in the Helpifyr / JaddaHelpifyr stack: settlement decisions in progress can now survive service crashes and restarts without risk of double-processing or data drift. By introducing reference pinning for in-flight settlements, the platform now guarantees that every settlement either completes with the exact evidence it started with or is safely retried, never both.
Read