Document automation

AI document calculations: how to check the answers

When document AI generates calculations, check the source values, business rule and saved code. A practical review plan using Unsiloed’s beta announcement.

By Clairevue · · 5 min read

A magnifying glass checks a highlighted invoice row beside a blank-key calculator and pencil.
AI-generated illustration of document and calculation review; not actual invoice data or an Unsiloed screenshot.

Copying a total from an invoice and calculating a total from its line items leave different things to check. In the first case, you can compare the answer with a number on the page. In the second, you also need to know which rows the system included and what rule it applied.

An announcement for Unsiloed’s Higher-Order Extraction, described as a beta, proposes combining document extraction with generated calculation code. That could make derived answers easier to inspect than a number returned by a model alone. But a finance or operations team still needs to check the source values and the business rule before using the result.

Follow the calculation from the page

Unsiloed describes a request becoming a dependency graph: extracted values feed calculations, and later calculations wait for the inputs they need. A model writes a deterministic function for each calculation, a second model reviews it, and the system compiles the graph into a script.

A developer can run and test the arithmetic. The model interprets the request and document; code then performs the calculations it specified. Those are separate opportunities to make a mistake.

Suppose a fictional invoice lists ten units at $120 each, with a 10% discount on the subtotal and 10% tax on the discounted amount. Those are assumptions for this example, not tax guidance. The calculation should produce:

CalculationExpected answer
Quantity × unit price$1,200 subtotal
Subtotal less its 10% discount$1,080
Tax on the discounted amount$108
Discounted amount plus tax$1,188 total

If the system reads the discount as a fixed $10 rather than 10%, it can execute flawless arithmetic and return $1,309. Re-running that code won’t repair the interpretation. Someone needs to compare the discount with the source and the agreed rule.

Unsiloed’s extraction guide describes schema-defined fields with confidence scores. A field can contain a valid number while representing the wrong column or period. The schema needs to say which amount you want, including its currency and units.

Check the rule as well as the code

The announcement says a second model reviews each function against the calculation it should perform. That adds a code-review step, but both models could accept the same mistaken interpretation of the request. The announcement doesn’t establish that the reviewer checks against an independently approved business rule or known correct answers.

For a pilot, have the person responsible for the process specify the calculation before examining the model’s output. They should record expected answers for a set of documents and explain what should happen when an input is missing. A missing tax rate should trigger review rather than silently become zero.

In its separate bank-statement extraction guide, Unsiloed describes comparing running balances and transaction totals with the statement summary. The guide explicitly warns that errors can cancel out: consistent extracted figures may still fail to match the page.

Arithmetic checks test relationships among the inputs you have; they don’t prove you read all the correct inputs. A missing row needs a completeness check, and a wrong account number needs a source check even if every balance adds up.

Amazon Textract’s best-practice guidance also warns that merged or irregular table cells can yield inconsistent extraction results. A neat JSON response shouldn’t stop you checking a difficult table against the original.

Save the evidence that lets someone reproduce it

Before using a calculated answer, ask the supplier to show the record of one completed run. You should be able to follow the number back through:

  • The original document and the page location for each input, with the value and its units.
  • The agreed calculation rule, including how it treats discounts, missing fields and rounding.
  • The exact generated code and dependencies that ran, together with their version.
  • The reviewer’s findings, test results and any human corrections or approval.

Ask the supplier to demonstrate which of these records the beta retains. Unsiloed’s published extraction response reference describes job metadata, schema and field values with scores; that alone doesn’t establish what a Higher-Order calculation run retains.

The announcement also says a changed document format prompts new extraction and a regenerated graph. Keep that new run separate from the old one. Reproducing yesterday’s answer means executing the saved calculation with its saved inputs and relevant runtime settings, rather than asking a model to generate another solution today.

Generated code needs execution controls too. OWASP’s guidance on improper output handling warns about passing model-generated output into downstream systems without validation. Have the technical team restrict where calculation code runs and what it can access. Calculating an invoice total shouldn’t grant access to payment credentials.

Test one answer before automating the next action

Choose one derived value your team already checks manually. Keep the pilot read-only: the system can propose a figure, while a person decides whether to use it. Automatic ledger entries or payments can wait.

Use representative documents with expected answers recorded in advance. Include changed layouts and credit notes, plus at least one case with missing information. Test the approved calculation code with known inputs separately from the extraction, so a failed result doesn’t leave you guessing which part broke.

Measure how often the complete answer is correct, which exceptions the system catches, and how long reviewers spend resolving them. Also record wrong answers that pass without a warning; a confident mistake creates different work from a document routed to review. Count review time and integration work when judging whether the pilot saves effort.

Before putting the result into a spreadsheet or ledger, ask someone to reproduce it from the saved inputs and point to where those inputs appear in the original document.