Picture the cheap version. At the table it takes a few seconds. “Hey, I think I got charged for the second round — can we redo that bit?” Someone looks at the receipt, someone says “oh, that’s mine,” and the whole thing is over before the card comes back.

Now the expensive version: the same sentence, sent from the car an hour later, becomes a thread that runs for days. The scenes here are illustrative rather than measured — but the difference between them is not. Nothing about the arithmetic changed. What changed is that the sentence is now travelling through a channel that gives you far less way to check whether it landed as you meant it.

There is a mechanism behind that, and it is more specific than “don’t argue over text.” The research does not say your message loses its tone in transit. It says something stranger and more useful: your confidence that the tone arrived barely shifts between channels, while your actual success at conveying it measurably does. You are about as sure over text and somewhat less clear, and nothing in the channel tells you which.

88.8% vs. 70.4% what senders predicted they would successfully communicate, versus what they managed — pooled across all three channels in Study 3 of Kruger et al. (2005), not a text-only penalty
Ability moved. Confidence barely did. in that study, how well people communicated varied significantly by medium, while their confidence in their own ability showed no detectable variation
$100–149 the band the median bill falls in across splitty's own scanned US restaurant receipts (sample of 5,645) — a self-selected sample of people who scan receipts, not a national average

Why do bill arguments get worse in the group chat?

Because the group chat thins and delays your feedback, without changing your meaning. When you say something awkward across a table, you find out quickly — a flicker, a pause, a laugh that arrives half a beat late — and you repair it in the same breath. In a thread the reply comes eventually, but not while you are still able to adjust. You spend the gap assuming your version was received, because nothing has told you otherwise.

Friedman and Currall built a model of this in Human Relations in 2003. Their argument is that e-mail is unusual in being simultaneously asynchronous, textual and electronic, and that the diminished feedback this produces should push disputes toward escalation rather than resolution. Separately, they argue that e-mail’s reviewability and revisability — the fact that you can re-read a message many times and draft a reply at length — invites what they call excess attention: rumination on the receiving side, and on the sending side a growing psychological investment in your own argument the more you redraft it.

Their own counterpoint, which is worth keeping. The same paper notes that slow feedback can prevent escalation: the added time may let people calm down and choose their words rather than fire off something rash. And it says plainly that e-mail “does not turn all communications into escalated conflicts.” The claim here is about susceptibility, not inevitability — most splits get settled in a thread without incident.

A note on what this source is. Friedman and Currall proposed a conceptual framework — the dispute-exacerbating model of e-mail — and a set of numbered propositions intended as a foundation for later empirical work. It is a model, not an experiment, and it reports no effect sizes. It is cited here for the shape of the mechanism, and the measured claims in this piece come from the experimental work below.

A restaurant table is the opposite environment on every one of those dimensions. It is synchronous, it is spoken, and the correction loop is close to immediate. That is why the same objection costs so little there — not because people are more generous at dinner, but because a misunderstanding gets caught before it becomes a position anyone has to defend.

How small are the sums actually in play?

It is worth being concrete about the stakes. Across splitty’s own scanned US restaurant receipts — a snapshot taken on 2026-08-21 — the median bill falls in the $100–149 band, and about 48% of receipts come in at $150 or more. Divided across a table of four, that band works out to something like $25 to $37 a head — plain arithmetic on the band rather than a measured figure, since the receipts carry no party-size dimension. The line item people actually re-open — a round someone didn’t drink, an appetizer that got split three ways instead of four — is some fraction of that again. How large that fraction typically is, splitty’s receipts cannot say: they record what was ordered, never what was contested.

$100–149

The familiar version of this is that nobody sends a message at 11pm because they need six dollars — they send it because the six dollars has stopped being six dollars. That is an observation about how these arguments feel, not a measured finding: splitty’s receipts carry no dispute or motive data, and no study cited here examined why people re-open a split. What the research does support is the mechanism underneath it, which is where the rest of this piece goes.

What splitty’s data can and cannot say here. The receipt sample supports the premise above — the size of a typical bill in this sample, and therefore of a typical share. It is not a national statistic: these are receipts from people who chose to scan them, so read it as splitty’s own window rather than the US average. It carries no party-size dimension either, so the per-head figure above is arithmetic on the band rather than an observed split. And it says nothing about how often splits get disputed, when disputes happen, or where they happen, because scanned receipts carry no dispute or channel dimension at all. Every claim in this piece about escalation rests on the published research, not on splitty’s receipts.

You are about as confident over text, and measurably less accurate

The central finding comes from Kruger, Epley, Parker and Ng, writing in the Journal of Personality and Social Psychology in 2005. Across five experiments they compared how well people thought they were communicating tone with how well they actually were.

The cleanest version is their third study, which ran the same task through three channels — e-mail, voice only, and face-to-face — and measured both halves of the exchange. Take the sending half first, because that is the half you control.

Pooled across all three channels, senders predicted they would successfully communicate 88.8% of the emotions and tones they were trying to convey, and managed 70.4%. That pooled gap is not the text penalty — it is how badly people read their own clarity in general. The channel-specific part is separate, and it points the way you would expect: participants were significantly more overconfident over e-mail than when using their voice.

The finding that matters most for a group chat is the one the authors put almost as an aside. Participants’ ability to communicate varied considerably depending on whether they used e-mail or their voice — while their confidence in that ability showed no detectable variation.

That sentence is the whole article. Your skill at getting a tone across genuinely changes when you switch from a table to a thread. Your sense of how well you are doing it, as far as this study could detect, does not. So the channel moves one variable and leaves your only instrument for noticing pointing at roughly the same number. Worth stating precisely: a non-significant result means no difference was detected, not that confidence is proven identical. The claim is that your confidence is a poor alarm, not that it is a constant.

Sender figures and both interactions from Study 3 of Kruger, Epley, Parker & Ng (2005), Journal of Personality and Social Psychology 89(6), 925–936: predicted vs. actual communication, F(1, 148) = 118.26, p < .001, η² = .44; greater overconfidence over e-mail than voice, F(2, 148) = 3.31, p = .039, η² = 0.04; ability varying by medium, F(2, 151) = 4.68, p = .011, η² = 0.06, while confidence in that ability did not, F < 1, ns.

The same study measured the receiving half too, and it fails in the matching way. After each message, receivers picked the intended emotion or tone from a list of four and ticked a box saying whether they were confident in that pick. They ticked yes about 89% of the time in every condition. How often they picked correctly moved a great deal:

ChannelConfident in their pickPicked correctlyGap

Figures are the reported condition means from the same study, F(2, 148) = 3.66, p = .028. This is the receiver’s measure, and it is a forced choice among four candidate tones plus a yes/no confidence tick made after each guess — not a prediction made in advance, and not a measure of general understanding. The gap column is simple arithmetic on the two columns, not a separately reported statistic.

So both ends of the thread have the same defect pointing in the same direction. You overestimate how well you sent it; they overestimate how well they read it — and over text that second gap is roughly ten points wider than it would be face-to-face. Neither side has much of an instrument for noticing that anything has gone wrong.

The honest distance between this study and your group chat. Kruger and colleagues ran a constrained laboratory task in 2005: pairs of people, one sentence at a time, emoticons prohibited, a fixed menu of tones, and no group audience, reactions, read receipts or live thread. A modern group chat is none of those things. What transfers is the asymmetry between confidence and accuracy across media; the specific percentages describe that task, not your Saturday-night thread.

The authors trace this to egocentrism. When you write “fine, I’ll just cover it,” you can hear your own delivery — the shrug, the lightness, the joke you intended. That private soundtrack is unavailable to everyone else, and it is very hard to set aside when estimating what they received.

The authors’ own caution, which matters. In their first study people expected 97% of their intended meanings to be decoded and 84% actually were. The researchers explicitly warn against reading this as evidence that people are bad at conveying tone in writing — 84%, as they note, is quite high. Their conclusion is narrower and is the one this piece is built on: however good people are, they are not as good as they believe. The danger is the unearned confidence, not incompetence.

Text isn’t as cue-free as the folk theory says

There is a tidy version of this argument that is worth resisting, because the evidence does not support it. The tidy version says text is a barren channel: no face, no voice, no cues, therefore no emotional information.

Blunden and Brodsky tested something close to that assumption in Personality and Social Psychology Bulletin in 2021 and found the opposite of a barren channel. Across six studies, communication mistakes — typos, errors — amplified readers’ perceptions of the sender’s emotion, both negative and positive. People treat a sloppily typed message as evidence of a state of mind. In one study, readers partially excused errors in emotional contexts, attributing them to how the sender was feeling rather than to how smart they are. The authors conclude that nonverbal behaviour in text-based and face-to-face communication may be more comparable than previously thought.

So the problem is not an absence of cues. It is that some of the cues that survive are accidental ones. Note the precise finding: errors amplify an emotion the reader is already inferring — they are a volume knob, not a signal generator. Which matters most in exactly our case, because a message re-opening a settled bill already carries a plausible reading of annoyance for the typos to turn up. You dashed it off at a red light; the sloppiness reads as feeling, and you have no way to know it did, because you know you were merely in traffic.

Whatever they already believed about you gets louder

The gap that opens up gets filled, and it does not get filled with neutral material. Epley and Kruger ran a set of experiments published in the Journal of Experimental Social Psychology in 2005 with an unusually clean design: they held the word-for-word content constant and varied only whether it arrived by e-mail or by voice.

Across three experiments, pre-existing expectancies about the person shaped impressions of them more strongly over e-mail than over voice, despite identical wording. Follow-up analysis attributed the effect, at least in part, to the greater ambiguity of e-mail. Ambiguity is not neutral. It is a space, and what fills it is whatever the reader already thought.

What was actually manipulated. Those experiments used racial stereotypes and bogus, experimenter-supplied expectancies as the prior beliefs. They were not about money, friends, or restaurant bills, and they did not study anyone’s reputation for being tight with a check. What transfers is the structural finding — identical words, read through a prior, land harder over text — not a claim that this specific research examined a bill dispute.

Structurally, though, it is the same shape as the thing every group has. If someone at that table has a reputation, earned or not, for going quiet when the check arrives, your message is not being read cold. It is being read through that. Say the identical sentence at the table and your face, your timing and your tone are all competing with the prior. Send it in text and the prior meets less competition. Epley and Kruger showed a relative shift, not an all-or-nothing one: expectancies mattered more over e-mail, partly through its greater ambiguity — not that they become the only thing operating.

Why a disagreement in text reads as a character flaw

The final step is the one that turns a bookkeeping question into a grievance. Schroeder, Kardas and Epley, writing in Psychological Science in 2017, found that hearing someone explain a position makes them seem more mentally capable — more in possession of a thinking, feeling mind — than reading the identical content. Their title is the summary: speech reveals, and text conceals, a more thoughtful mind in the midst of disagreement.

Crucially, the effect showed up precisely where you would least want it to: when the other person disagrees with you. That is the condition under which people are most inclined to treat an opponent as relatively mindless, and it is exactly the condition a bill dispute creates. The researchers’ own framing is that the medium through which people communicate may systematically influence the impressions they form of each other, and that the tendency to denigrate the minds of the opposition may be tempered by giving them, quite literally, a voice.

Scope, again. These experiments used polarising attitudinal and political topics — not dinner. The transferable claim is about disagreement and medium in general. Nobody has run this study on a restaurant check.

Line the findings up and you get a plausible account of the conversion. You send a message you are sure is reasonable. It arrives more ambiguous than you believe, carrying accidental cues you did not intend. The ambiguity is resolved using whatever the reader already suspected. And because you are disagreeing in text rather than in voice, you are likely to come across as less thoughtful than you would have sounded. Six dollars went in; a claim about your character came out.

This is a compounding account, not a tested pathway. Each link is separately evidenced, but no study here tested them as a chain, and none of them studied a bill dispute. Read the sequence as the most reasonable explanation the available research supports for a familiar experience — not as a measured causal mechanism.

The table versus the thread

The same dispute, run through two channels, is really two different events:

At the tableIn the group chat

The practical lesson is not that text is forbidden. It is that the two columns are not substitutes, and the cheap column has a closing time.

How to keep a bill disagreement cheap

1

Raise it while the table still exists

The quick version is only available for as long as everyone is sitting down. If something looks wrong on the receipt, say it before the group stands up — the fast correction loop that makes it cheap disappears when the group does. No study here ranked these remedies against each other; this one is first because it is the only one that avoids the cue-poor channel altogether rather than compensating for it.

2

Send a record, not a verdict

A bare total — 'you owe me forty-seven fifty' — asks to be accepted or rejected. An itemization can be checked line by line without anyone having to concede anything first. The reasoning is that a checkable record leaves less of the ambiguity the research says a reader fills in from their priors; none of the cited studies tested receipts or itemization, so this is an argument from the mechanism rather than a measured effect.

3

Assume your tone did not arrive

Not because you wrote it badly — because the research says your confidence about tone is roughly the same over text as face-to-face while your accuracy is not. Write the version that survives being read in the worst plausible voice, since you have no way to check which voice it got read in.

4

If it has already escalated, change the channel

Once a thread has turned into a character question, more text is the wrong instrument — it is the medium the trouble started in. A short call restores the paralinguistic cues the research says text lacks, and Schroeder and colleagues found voice narrows exactly the mindedness gap that opens during disagreement. Nobody has tested a phone call as a de-escalation remedy in a bill dispute, so treat this as reasoning from mechanism, not a measured cure.

5

Make the record the default, not the escalation

A receipt produced during an argument is evidence in a dispute. The same receipt shared when the check lands is just how the group sees the bill. If itemization only appears when someone is defending themselves, it reads as an accusation — so it has to be routine to be cheap.

Where splitty fits

The recurring failure here is structural rather than emotional: the itemization stays at the restaurant while the ask travels to the thread. What arrives in the group chat is a bare number, stripped of the only thing that would have made it checkable, in the exact channel where ambiguity gets filled in by whatever people already believed.

splitty is built so the record and the ask travel together. It scans the receipt and splits it line by line — everyone starts on every item, and you tap to remove whoever didn’t have the second round. Each person gets a request built from their own lines rather than a division of the total, so what lands in their app is an itemization they can read against their own memory of the meal. Share Split goes one step further and sends the record itself: a web link showing each recipient their own portion, the itemized breakdown behind it, and pay buttons, with no app install required. That does not make anyone more generous. It removes the ambiguity that the channel would otherwise fill for them.

It is worth being honest about the limit. No app can restore a paralinguistic cue to a text message, and none of this makes a conversation about money stop being a conversation about money. Nor has any of the research cited here tested receipts, itemization, or splitty — the case for a shared record is an argument from the mechanism these studies describe, not a measured result. The aim is narrow: give the argument a smaller surface, so it stays a question about a line on a receipt rather than a question about a person.

Common questions

FAQ

Questions & Answers

01 Why do money arguments feel worse over text than in person?

Because your confidence that you were clear does not adjust to the channel, but your accuracy does. In Study 3 of Kruger, Epley, Parker and Ng (2005), senders predicted they would successfully communicate 88.8% of the emotions and tones they intended and actually managed 70.4%, and they were significantly more overconfident over e-mail than by voice. The decisive detail is that their ability to communicate varied considerably by medium while their confidence in that ability showed no detectable variation — a null result, meaning no difference was found rather than equality proven. The receiving end fails the same way: people said they were confident in their reading about 89% of the time in every condition, but were right 62.8% of the time over e-mail against 73.9% face-to-face.

02 Is it true that text messages have no emotional cues?

No, and the tidy version of that claim is contradicted by the evidence. Blunden and Brodsky (2021) found across six studies that communication mistakes such as typos amplify readers' perceptions of a sender's emotion, both negative and positive, and concluded that nonverbal behaviour in text-based and face-to-face communication may be more comparable than previously thought. The issue is not that cues are absent — it is that the surviving cues are accidental ones the sender never intended and cannot see landing.

03 Why does a small disagreement about a bill turn into a fight about character?

Ambiguity gets resolved using prior beliefs. Epley and Kruger (2005) held wording word-for-word constant and varied only whether it arrived by e-mail or voice; pre-existing expectancies shaped impressions more strongly over e-mail, an effect they attributed at least partly to e-mail's greater ambiguity. Separately, Schroeder, Kardas and Epley (2017) found that hearing someone explain a position makes them seem more mentally capable than reading identical content, specifically when they disagree with you. Text widens the gap and then fills it with what the reader already suspected.

04 What is the best thing to do about it?

Raise it before the group stands up. The mechanisms that make a dispute expensive — delayed feedback, unverified tone, ambiguity resolved by priors, disagreement read as mindlessness — are all strongest in asynchronous text. They are attenuated at the table, not abolished: face-to-face comprehension in Kruger et al. was 73.9%, not perfect, and priors and denigration can operate in person too. The cited studies show voice and presence reduce these effects rather than eliminate them. If that window has closed, changing channel beats sending more text, though no study here compared those remedies head to head.

05 How should I word a message asking someone to pay me back?

Worry less about wording and more about attaching the itemization. A bare total is a claim someone has to accept or reject, which makes it about trust; an itemized record is a document they can check, which makes it about a line item. Then write assuming your tone did not survive — the finding from Kruger et al. (2005) is that people are not as clear as they believe, so the safe assumption is that the lightest reading you intended is not the one that arrived.

06 Does this mean I should never discuss a split in a group chat?

No — most splits are settled in a thread without incident, and Kruger and colleagues explicitly caution against concluding that people communicate tone badly in writing, noting that 84% accuracy in their first study is quite high. The asymmetry is about disagreement specifically. Coordinating a split over text is routine; re-opening a contested one over text is the case where the channel works against you, and that is the one worth moving.