Reading a receipt and computing a split are two different jobs. A language model can do the first well: it can turn a crumpled photo into a tidy table of items and prices. The research below is about why the second job is where it fails. Ask the same model to then work out who owes what, and the failures are characteristic: a fee that never shrinks after an item is removed, a tax recomputed on the wrong base, five shares that add up to less than the bill. The fix we recommend is architectural, not a better prompt: the model should read, and ordinary code should compute. Three quick checks tell you whether any scanner, chatbot, or app you hand a receipt to gets this right.

The clearest public demonstration is a 2024 test by the AI company Random Walk, which fed a five-person restaurant bill to ChatGPT 4.0. The model “handled the task flawlessly” when asked to turn the receipt into a table. Asked to recompute the bill without one dish, it scaled an alcohol-only tax down with the whole subtotal. Asked for each person’s share of the total, it returned percentages that, added together, account for about 87% of the money. The table was right. The arithmetic was not.

55% of three-digit by three-digit multiplications off-the-shelf ChatGPT got right; GPT-4 managed 59% (Dziri et al., NeurIPS 2023)
87.04% is what the five per-person percentages ChatGPT returned in Random Walk's bill test sum to, against a total that should reach 100%
97.9% of math problems on which Toolformer, a model trained to call tools, chose to hand the arithmetic to a calculator (Schick et al., 2023)

Sources: Dziri et al., “Faith and Fate: Limits of Transformers on Compositionality,” NeurIPS 36 (2023); Random Walk, “The Story of a Bill,” 2024 (the 87.04% figure is the sum of the five percentages the post prints); Schick et al., “Toolformer,” NeurIPS 36 (2023).

Why is reading a receipt a different job from computing the split?

Reading is pattern recognition: find the item names, find the prices, keep them on the right rows. Computing is executing a procedure exactly: multiply, carry, sum, and get the same answer every time. Language models are built for the first. They predict the next piece of text from patterns in their training data, and a long column of digits is just more text to them. That helps explain how the same system can transcribe a receipt perfectly and then stumble on the sum of a column it just transcribed.

The research teams that build these models say so in their own papers. The OpenAI group that built the GSM8K math benchmark wrote that “our models frequently fail to accurately perform calculations,” that larger models make fewer arithmetic mistakes but “this remains a common source of errors,” and that to mitigate it they trained every model to use a calculator that “will override sampling” whenever the model chooses to invoke it. A 2021 study by Nogueira, Jiang, and Lin found that a small pretrained transformer (T5-220M) trained on 1,000 addition examples “fails to learn addition of five-digit numbers when using subwords,” that is, when numbers are split into multi-digit tokens such as “32.” The authors add that “with enough training data, models can learn the addition task regardless of the representation.” The limit they found at every scale is a different one: “regardless of the number of parameters and training examples, models cannot seem to learn addition rules that are independent of the length of the numbers seen during training.”

Reading the receiptComputing the split
What the job is Recognize items, prices, quantities, and totals in a photoAssign items to people, allocate tax and tip, sum each share
What it needs Tolerance for blur, fonts, abbreviations, layoutExactness: the same inputs must always give the same cents
Language model fit Strong: this is pattern recognition over text and pixelsWeak: digits are tokens, and multi-step arithmetic compounds errors
How it fails A misread price or a dropped line, visible on inspectionA total that looks plausible and is wrong, invisible without a check
Right tool A vision or OCR modelDeterministic code operating on exact cents

The deeper reason is compositional. Dziri and colleagues formalized tasks like long multiplication as graphs of small steps and showed that transformer accuracy “decreases to near zero as task complexity increases,” in part “due to error propagation”: one wrong intermediate digit poisons everything downstream. A restaurant split is exactly this kind of task. Subtotal feeds tax, tax and tip feed the total, the total feeds every per-person share. Get the service charge wrong in step two and all five shares in step five are wrong, each by a plausible-looking amount.

Sources: Cobbe et al., “Training Verifiers to Solve Math Word Problems,” arXiv (2021); Nogueira, Jiang & Lin, “Investigating the Limitations of Transformers with Simple Arithmetic Tasks,” arXiv (2021); Dziri et al., NeurIPS 36 (2023).

What does it look like when a language model does the bill math?

It looks right until you check. Random Walk’s September 2024 test is worth walking through step by step, because its arithmetic failures are ones a reader can catch with the checks below. The bill was a five-person restaurant tab of ₹13,545.00 before charges, with a service charge, a 14.5% VAT that applied only to alcohol, and two 2.5% goods-and-services taxes that applied only to food and non-alcoholic drinks, for a printed total of ₹15,562.00.

Step 1 Convert the receipt to a table. The post reports ChatGPT 'handled the task flawlessly.' Reading: pass.
Step 2 Compute the total with and without taxes. Correct: 13,545.00 and 15,562.00, matching the receipt.
Step 3 Remove one dish (the 449.00 Adana Kebab) and recompute. The model cut the alcohol-only VAT in proportion to the whole subtotal (1,424.92 ÷ 13,545.00 × 13,096.00), as if VAT were a fixed share of the whole subtotal, although the kebab is food and carried no VAT. The post: it 'struggled to provide the correct total amount when taxes were included.'
Step 4 After correction, the model still 'did not proportionally reduce the service charge after removing the Adana Kebab.' A percentage fee stayed frozen while its base shrank.
Step 5 In a second, consolidated chat, compute each person's share of the total including taxes. The model divided each person's pre-tax spend by the post-tax total: 14.96%, 13.68%, 16.89%, 19.78%, 21.73%. Those five numbers sum to 87.04%.
Step 6 List the food and non-alcoholic items. The post: the model 'hallucinated and provided an inaccurate response, including items that were not on the bill.'

The post’s own verdict: the model “required multiple prompts and follow-up queries to grasp the logic of adding and subtracting taxes and service charges when removing an item and splitting each person’s bill by percentage.” Notice the shape of the errors in Steps 3 to 5. None of them is a misread digit. Each is a computation that used the wrong base, skipped a dependent line, or mixed a pre-tax numerator with a post-tax denominator. Those are arithmetic-procedure failures, and a clean table does not protect you from them. Not every failure in the post was arithmetic, though: in Step 6 the model listed items that were not on the bill, and elsewhere the post says “the LLM confused food and beverages” and that ChatGPT “mistakenly excluded the cost of the Adana Kebab.” Getting the table right once is not the same as keeping track of it across a long conversation.

The sum is the tell. Five shares of one bill must add to 100% of it. 87.04% means about 13% of the money was assigned to nobody, and that missing slice is almost exactly the tax-and-service overhead the model left out of the numerators. A split that fails this check is wrong before you ask who owes what.

Source: Random Walk, “The Story of a Bill: How Well Can AI Models Handle Real-World Math,” 2024-09-27. All quoted phrases and figures are from the post; the 87.04% sum is computed from the five percentages it prints.

Check 1: Do the shares add up to the bill total?

Add every person’s share. The sum must equal what the table actually pays, to the cent: the printed total, plus any tip added at the table. Any tool that fails this has either dropped money or invented it. One easy way to drop it is the one Random Walk’s model hit: dividing each person’s pre-tax items by the post-tax total. The shares then add to the subtotal’s share of the total, and the overhead is orphaned.

Here is the same mistake on an illustrative US check. The numbers are chosen for clean arithmetic, not taken from any receipt.

Illustrative two-person check
Food & drink subtotal (A ordered $50, B ordered $34)$84.00
Sales tax, 8% of the subtotal$6.72
Tip, 20% of the pre-tax subtotal$16.80
Total$107.52

Wrong direction: A’s share = $50 ÷ $107.52 = 46.5%; B’s share = $34 ÷ $107.52 = 31.6%.
Sum: 78.1%. The missing 21.9% is the tax and tip, assigned to nobody.

Right direction: multiplier = $107.52 ÷ $84.00 = 1.28.
A owes $50 × 1.28 = $64.00; B owes $34 × 1.28 = $43.52. Sum: $107.52.

The right direction is the proportional method: one multiplier, and each person’s items times that number. This example was built so the multiplier comes out to exactly 1.28 and every share lands on a whole cent. On a real bill the multiplier is usually a long decimal, so each share has to be rounded, and the rounded shares may not add back to the total exactly. A well-built tool has to place those leftover cents on purpose, and Check 1 is how you find out whether it did. The full method, including why it beats an even split of tax and tip, is in how to split tax and tip in proportion. One honest limit: a sum that reconciles proves the arithmetic closed, not that the right items landed on the right people. A shared plate assigned to one diner leaves the total untouched. The sum check is necessary, not sufficient.

Check 2: Does removing an item lower every line that depends on it?

Take one item off the bill and watch what moves. The subtotal must fall by the item’s price, and every percentage line whose base included that item must fall with it: sales tax on a taxable item, a percentage tip, an automatic service charge. A line whose base did not include the item, such as an alcohol-only tax when a food item comes off, should stay put. This check is about repricing the bill, as when a dish is voided or was never on your order; a splitter that keeps the printed charges while you reassign an item among diners is doing a different job, and Check 2 does not condemn it. Random Walk’s model got the subtotal right and the rest wrong in both directions: it cut the alcohol-only VAT when a food item was removed, and later left the service charge at its original amount after the base it was charged on had shrunk.

On the illustrative check above, suppose B drops a $10 dish. The subtotal becomes $74.00, tax becomes $5.92, the 20% tip becomes $14.80, and the total becomes $94.72. A tool that only subtracts the $10 and leaves tax and tip frozen reports $97.52, which is $2.80 too high: $0.80 of stale tax and $2.00 of stale tip. The error is small on one dish and compounds when several items come off the bill before it is settled.

Why models can miss this. One plausible reason: a frozen fee is a natural continuation of the text. The model has seen the service-charge figure once already, and nothing in next-token prediction forces a recomputation. The Random Walk post reports the frozen fee, not the reason for it. Deterministic code that defines the fee as a formula over the subtotal has no such option. Code that stores the fee as a fixed number does, which is why the check applies to every tool.

Check 3: Does each tax and fee line recompute from its own base?

Every percentage on a receipt has a base, and the bases differ. Sales tax applies to taxable items. A tip may be figured on the pre-tax subtotal. A delivery service fee may be a percentage of the food only. On the Indian bill in Random Walk’s test, one tax applied only to alcohol and two applied only to food and non-alcoholic drinks, and the model’s first recomputation scaled all of them with the whole subtotal. That is the base error: right percentage, wrong column.

The check is simple. For each percentage line, divide the printed amount by the base the tool used and confirm the rate matches the receipt. Check 3 confirms each line is consistent with a base and a rate. It cannot by itself show which items were in that base, since different item sets can share a subtotal, and it does not show the line was then allocated to the right people: an alcohol-only tax can be computed correctly and still be spread across the whole table, so the proportional split reconciles while quietly charging the water drinker for part of the bar tab. On a bill like that, compare a drinker’s share and a non-drinker’s share by hand. If a tool cannot tell you what base it used, you cannot tell which of those it did. The anatomy of each line, and which base it legally sits on, is in your restaurant receipt, explained.

1

Shares sum to the total

Add every person's share. It must equal what the table actually pays, the printed total plus any tip added at the table. Fail: money was dropped or invented, for example by dividing pre-tax items by a post-tax total.

2

Removing an item moves every dependent line

Delete one item and confirm the subtotal and every percentage line whose base included that item fall, while lines whose base excluded it stay put. Fail: a fee stayed frozen while its base shrank, or a tax moved that should not have.

3

Each percentage line recomputes from its own base

Divide each tax or fee by the base the tool used and confirm the rate. Fail: an alcohol tax or a pre-tax tip was computed on the whole total.

Why do the research labs hand the arithmetic to a calculator?

Because the model’s reasoning is often fine and its arithmetic is not, and the two can be separated. The PAL paper from Carnegie Mellon, published at ICML 2023, had a model write its reasoning as a short Python program and let the Python interpreter produce the answer. On the GSM8K word-problem benchmark the program-aided model solved 72.0% of problems against 65.6% for the same model reasoning in plain text. The more telling result came from GSM-Hard, a copy of the benchmark with the numbers replaced by large ones: the text-reasoning model’s accuracy fell from 65.6% to about 20% (the paper’s text says 20.1%; its Table 1 prints 23.1%), a relative drop the text puts at “almost 70%.” The program-aided version stayed at about 61%: the text says it “remains stable at 61.5%, dropping by only 14.3%,” and Table 1 prints 61.2%.

The authors then asked whether the big numbers confused the reasoning or just the arithmetic. In 16 of the 25 cases they examined, the model’s written-out reasoning was “nearly identical” with small and large numbers, “indicating that the primary failure mode is the inability to perform arithmetic accurately.” Their design conclusion applies directly to a receipt splitter: “since PAL offloads the computation to the Python interpreter, any complex computation can be performed accurately given the correctly generated program.”

65.6% → 20.1%accuracy of plain-text reasoning on GSM8K when the problems’ numbers were replaced with large ones (GSM-Hard), per the paper’s text; its Table 1 prints 23.1% for the large-number run. In most sampled cases the reasoning barely changed; the arithmetic broke. Gao et al., ICML 2023.

Meta’s Toolformer, published at NeurIPS 2023, reached the same conclusion from the other side. The team trained a 6.7-billion-parameter model to decide for itself when to call a calculator. On three arithmetic word-problem benchmarks, allowing those calls “more than doubles performance for all tasks,” lifting the model from 14.8%, 6.3%, and 15.0% to 40.4%, 29.4%, and 44.0%, and past the much larger GPT-3. The reason, in the authors’ words: “across all benchmarks, for 97.9% of all examples the model decides to ask the calculator tool for help.” That is a measurement under the paper’s setup, which used a decoding setting of k = 10 “to increase the disposition of our model to make use of” its tools, not a claim that it knew its own limits; the direction is still the point: the arithmetic went to the calculator.

Cobbe et al. (2021): models “frequently fail to accurately perform calculations,” so a calculator overrides the model whenever it is invoked→A splitter should never let the reading model emit a final share; code should compute it from the extracted items
Gao et al. (2023): in 16 of 25 sampled cases, reasoning stayed “nearly identical” while arithmetic failed as numbers grew→On word problems the failure was mostly in the digits, not the plan. A receipt adds a second failure, the wrong base, that exact arithmetic preserves; so a splitter needs both the right bases and a calculator
Schick et al. (2023): a tool-trained model asked the calculator on 97.9% of math examples→The labs’ pattern is a model that plans and a program that computes; the receipt-app analog is a model that reads and code that computes

Sources: Gao et al., “PAL: Program-aided Language Models,” ICML 2023 (PMLR 202); Schick et al., “Toolformer: Language Models Can Teach Themselves to Use Tools,” NeurIPS 36 (2023); Cobbe et al., arXiv (2021). The PAL paper’s prose and its Table 1 print different GSM-Hard figures (20.1% vs 23.1% for text reasoning, 61.5% vs 61.2% for PAL); both are shown above.

How can you tell which architecture a bill splitter uses?

Run the three checks, and if you can, read the plumbing. Some products publish it. A custom GPT in the ChatGPT store called “Split Bill Calculator, No Hallucinations” advertises that you can “simply upload your bill and Billy will split for you … accurately with no hallucinations.” Its page source shows how it earns the name: the GPT is wired to an external “Split Bill Action” at a separate domain, described by an OpenAPI 3.1.0 schema. The model’s job, per the schema, is to produce structured input, a list of people, the subtotal, the total, and a list of items with prices. The backend returns the shares. The schema describes a division of labor, with the chatbot handing the math to something else; it is not a published accuracy result, and it does not show that every answer goes through that path. We probed the endpoint: it answered a bare GET with HTTP 405 and an unauthenticated POST with HTTP 403, which shows a separate service exists and responds. It shows nothing about whether that service computes correctly. The three checks are the reader’s best test of that, though they cannot reveal architecture and cannot rule out a wrong assignment on their own.

The schema also shows what the reading layer can lose. Its item field is documented as “list of individual item and its total price. The total price is regardless of quantity,” so a line like “2 × margarita” is documented to arrive as one priced item; whether the count survives somewhere else, the schema does not say. If it does not, clean arithmetic downstream cannot recover a quantity the extraction flattened. That is the other half of the architecture: the deterministic layer is only as good as the structure it is handed.

Competitors are starting to sell the distinction outright. A Splitwise explainer published on split-the-bill.app in September 2026 names “deterministic calculation versus AI estimation” as “another real point of differentiation,” and describes its own flow as working “with no AI estimation—it’s deterministic math, the exact same calculation every time.” The same post says that some apps use AI OCR to “guess” receipt items, which “can make mistakes on crumpled receipts or small print,” and advises that “it’s always worth double-checking the total before finalizing the split, whichever app you choose.” That is a competitor’s pitch aimed at other apps, and its example is a reading error rather than a computation error, but the distinction it sells, reading a receipt versus computing from it, is the one this article is about.

Prone to fail the checks

Chatbot does everything

  • You paste or photograph the receipt into a chat
  • The model reads, assigns, and computes in one reply
  • Failures seen in the Random Walk test: frozen fees, wrong tax bases, shares under 100%
  • Every re-ask risks a new arithmetic error
Can pass the checks

Model reads, code computes

  • A vision or OCR step extracts items, prices, quantities
  • You confirm the extraction against the paper
  • Deterministic code assigns, allocates tax and tip, and sums
  • Still worth running all three checks: code can be built on the wrong base too
Can pass, slowly

You type, code computes

  • A calculator or spreadsheet with no reading step
  • Arithmetic is exact once the numbers are in
  • The transcription is the weak link, one chance of a typo per line item
  • Thirty line items is thirty chances to mistype

Sources: Split Bill Calculator, No Hallucinations, ChatGPT GPT Store (page metadata and embedded Action schema; endpoint probed 2026-10-05); Split The Bill, “Splitwise: What It Is, How It Works, and Pricing,” published 2026-09-12, updated 2026-09-18.

Does deterministic code guarantee the right answer?

No. Deterministic means repeatable, not correct, and money has a classic trap even for ordinary code. David Goldberg’s 1991 ACM survey explains that the decimal number 0.1, “although it has a finite decimal representation, in binary it has an infinite repeating representation,” so in standard binary floating point it “is exactly representable by neither” of its two nearest neighbors. Goldberg’s survey is about floating point in general, not receipts. Our own inference for bill splitting: code that stores dollars as binary fractions can pick up tiny rounding errors, and those can surface as a split that misses the total by a cent. A common engineering defense, which is general practice and not a claim from Goldberg’s paper, is to keep amounts in whole cents and round on purpose.

This is why Check 1 earns its place even against a tool that uses no language model at all. A sum that is off by a rounding slip is a different bug from a sum that misses the total by 13%, but both are caught by the same test. It is also why the mental math route fails for a different reason again: as that article’s sources argue, working memory rather than tokenization is the bottleneck when a human tries to carry twenty line items, a tax rate, and a tip percentage at once.

Source for the quoted passages on 0.1: Goldberg, “What every computer scientist should know about floating-point arithmetic,” ACM Computing Surveys 23(1), 1991. The application to receipts and the whole-cents practice are the author’s, not Goldberg’s.

Where does a receipt-splitting app sit in this?

The three checks apply to a receipt-splitting app like anything else, including splitty. splitty’s scan is the reading step: it scans an itemized receipt and reads the line items printed on it. Every item starts shared by everyone and you tap to remove whoever did not have it. Tax and tip are then allocated in proportion to each person’s items, and each person gets a pre-filled payment request for their share.

A proportional split of tax is the right method when one tax rate covers the whole bill. It is not the right method on a bill like Random Walk’s, where one tax applies only to alcohol: there, Check 3 is the one to run. If your receipt has a tax that applies to only some items, check those shares by hand, whichever tool you use.

Run the checks on any tool you use. Add the shares and compare them to the printed total. Remove an item and watch which lines move. The point of this article is not that one app is trustworthy; it is that trust should be earned by a sum you can check yourself, and that any tool which asks a language model to produce that sum has chosen the wrong worker for the job. If you want the arithmetic without the app, the bill split calculator on this site splits shared items among their sharers and allocates tax and tip by each person’s subtotal. For why the typing route fails in practice, see calculator vs bill splitting app; for what a scanner should be measured on, see which receipt OCR is better.

FAQ

AI bill splitting: quick answers

01 Can ChatGPT split a restaurant bill accurately?

It can turn a receipt into an accurate table; the arithmetic is the bigger risk. In a published 2024 test, ChatGPT converted a five-person bill to a table correctly, then wrongly cut an alcohol-only tax when a food item was removed, on a later attempt left the service charge unchanged, and returned per-person percentages that sum to about 87% of the total. If you use a chat model, check that the shares add up to the printed total and that removing an item moves every tax and fee line.

02 Why do language models make arithmetic mistakes?

They predict text rather than execute procedures. In the arithmetic research, how numbers are tokenized affects whether digit-by-digit rules transfer (the results depend on the representation and on training data), and multi-step calculations compound each error downstream. The teams behind GSM8K, PAL, and Toolformer all addressed this the same way: let the model reason, and hand the actual computation to a calculator or interpreter.

03 What is the difference between OCR and AI math on a receipt?

OCR, or any vision model, reads the receipt: it turns pixels into item names, prices, and quantities. That is pattern recognition, which these models do well. AI math is asking the same model to compute shares from what it read. That is procedure execution, which it does unreliably. A well-built splitter uses a model for the first step and deterministic code for the second.

04 If the shares add up to the total, is the split correct?

Not necessarily. A reconciled sum proves the net total matches; equal omissions and additions could cancel. It does not prove the right items landed on the right people; a shared appetizer assigned to one diner leaves the total unchanged. Treat the sum check as necessary, then confirm the assignments against the receipt, ideally with everyone able to see their own items.

05 What three checks catch a bad automated split?

One: the shares sum to what the table actually pays, the printed total plus any tip added at the table, to the cent. Two: removing an item lowers the subtotal and every percentage line whose base included that item, such as tax on a taxable item, a percentage tip, or a service charge, while a line whose base excluded the item stays put. Three: each tax or fee line, divided by the base the tool used, gives the rate printed on the receipt. A tool that fails any of them has a computation problem, whether a model or code did the math.