Relationship Texting Benchmarks 2026
Aggregate benchmarks from 4,600,611 messages across 29 de-duplicated histories: 97.4% carried no emotional signal, plus the six measures we withheld.
What do long relationship message histories actually show?
Across 4,600,611 messages in 29 de-duplicated two-party histories from 20 accounts, about 97.4% of messages carry no detectable emotional signal, and expressions of affection outnumber explicit repair attempts by roughly 10.7 to 1. Rates are computed per account before being combined, because one uploader contributed nine of the qualifying threads. This is a cohort of people who chose to have a relationship analysed, so it is not a representative sample of relationships, and nothing here diagnoses abuse, deception or illness, or predicts an outcome.
- The person who asked for the analysis sent 53% of the messages. Computed per account, then taken across accounts. This describes who chooses to have a thread analysed, not relationships in general.
- Affection ran higher outbound in 14 of 20 accounts. Per 1,000 messages: 22.463 outbound against 20.819 inbound.
- 97.4% of messages carried no detectable emotional signal. Long histories are mostly logistics, timing and small talk. The affective classes are the exception, which is why single screenshots mislead.
- Expressions of affection outnumbered repair attempts by about 10.7 to 1. Warmth is common; explicitly trying to fix a rupture is rare. The gap between them is the most stable ratio in the cohort.
- Hostile language was near-absent: the most common such class ran 0.032 per 1,000 messages. Distress, Threat did not reach a measurable rate. This is a property of who uploads, not evidence that hostility is rare in relationships.
- Not reported -- Criticism: Direction split is a measurement artifact, not a behavioural finding. Withheld pending a direction-blind re-scoring.
- Not reported -- Manipulation: Direction split is a measurement artifact, not a behavioural finding. Withheld pending a direction-blind re-scoring.
- Not reported -- Repair: The account majority (12 outbound vs 6 inbound) points the opposite way to the combined rate (2.108 vs 2.481). No direction claim is made.
- Not reported -- Blame: Rate rounds to zero across the cohort; the class is too rare here to report.
- Not reported -- Reply latency: Threads split 9 to 13 on which side replies more slowly -- no consistent direction, so no reply-speed claim is made.
- Not reported -- Thread duration: Span is clustered on a few exact values from repeat uploads of the same relationship, so a median span would describe the upload pattern rather than the relationships.
- Limitation: The cohort is people who chose to have a relationship analysed. It is not a random or representative sample of relationships.
- Limitation: Direction is defined by who uploaded. Any statistic split by direction inherits that selection.
- Limitation: Classifier labels are model output, not clinical judgements, and carry the error rate of the model that produced them.
- Limitation: The cohort is small. Figures show shape and order of magnitude, not precision, and no significance test is claimed.
- Limitation: Grey Mirror measures patterns in text. It does not diagnose abuse, deception, illness or predict outcomes.
Why a screenshot is the weakest evidence about a relationship
The single most consequential number in this cohort is how little of a relationship is emotionally legible at all. About 97.4% of messages carried no detectable emotional signal. Long histories are overwhelmingly logistics, timing and small talk; the affective classes are the exception rather than the texture.
That is what makes a screenshot structurally misleading. A screenshot is not a random sample of a relationship - it is the fraction of a percent that somebody chose to keep, usually while upset. Reading a relationship from saved screenshots means reading the least representative sample available, selected precisely because it was unrepresentative enough to be worth saving.
Measured per 1,000 messages, affection ran 22.463 outbound against 20.819 inbound, while explicit repair ran 2.108 outbound against 2.481 inbound. Neutral or uncertain messages ran 973.105 per 1,000. Those proportions are invisible in any excerpt.
Methodology
Every figure comes from one reproducible read-only aggregation over the production analysis store. No message text, sender name, email address or per-thread row enters the published artifact. Account identity exists in the aggregation only as a salted hash, and participant names are replaced before any text is scored.
A thread qualifies when it carries at least 500 messages, has two participating sides, has usable timestamps, and spans at least 30 days. The median span in the published cohort is 264 days. A re-upload of the same relationship counts once, keyed on the owning account and the thread endpoints rather than the message count, because a partial re-upload has a different count and the same conversation.
The account, not the thread, is the unit of analysis. One account contributed nine of the qualifying threads; averaging over threads would have given that account nine correlated votes. Every rate is therefore computed per account first and only then combined, and no account may supply more than a quarter of the cohort. Class breakdowns are suppressed entirely below 10 accounts, and nothing is published below 15.
- Eligible after all rules: 89 threads, from which the published cohort of 29 de-duplicated histories was drawn.
- Excluded, below the 500-message floor: 119 threads.
- Excluded, implausible timestamps that fell back to a parser sentinel: 46 threads.
- Excluded, span shorter than 30 days: 37 threads.
- Excluded, one-sided with no genuine second participant: 19 threads.
- Excluded, account cap so no single uploader dominates: 8 threads.
The six measures that were computed and withheld
A benchmark that reports only what worked is not a benchmark. Six measures were computed and then deliberately left unpublished. Each is named with the reason it was withheld, so the decision can be argued with rather than taken on trust.
- Criticism: Direction split is a measurement artifact, not a behavioural finding. Withheld pending a direction-blind re-scoring.
- Manipulation: Direction split is a measurement artifact, not a behavioural finding. Withheld pending a direction-blind re-scoring.
- Repair: The account majority (12 outbound vs 6 inbound) points the opposite way to the combined rate (2.108 vs 2.481). No direction claim is made.
- Blame: Rate rounds to zero across the cohort; the class is too rare here to report.
- Reply latency: Threads split 9 to 13 on which side replies more slowly -- no consistent direction, so no reply-speed claim is made.
- Thread duration: Span is clustered on a few exact values from repeat uploads of the same relationship, so a median span would describe the upload pattern rather than the relationships.
How each term is defined
Two of these terms are measured directly in this cohort. The rest are Grey Mirror report vocabulary, included here so a writer quoting the product is not guessing at what the words mean. The distinction is stated on each so the two are not confused.
- Emotional signal (measured here): a message the classifier assigns to one of the affective classes - affection, repair, blame, distress, harassment or threat - rather than to Neutral/Uncertain. The 97.4% headline is the share that fell to Neutral/Uncertain.
- Repair (measured here): a message the classifier scores as an explicit attempt to de-escalate or reconnect after friction. It is the rarest common class in this cohort.
- Reciprocity (report vocabulary): how initiation, follow-up and response are distributed between the two sides, reported as initiation balance - the share of conversation starts attributed to each participant.
- Evidence window (report vocabulary): the span of surrounding messages a claim is grounded in. A metric drawn from many well-distributed windows across months carries a narrower confidence band than one drawn from a few clustered windows in a single week.
- Timing drift (report vocabulary): a sustained change in response timing across comparable periods, as distinct from one isolated slow stretch. Not measured as a published figure in this cohort - see the withheld reply-latency measure below.
- Conflict loop (report vocabulary): the same disagreement recurring in different wording across a history, identified by recurrence rather than by any single exchange.
Limitations, stated plainly
This is a self-selected cohort of people who chose to have a relationship analysed, drawn from a consumer product. It is not a random sample, not a representative sample of relationships, and not a clinical population. Every figure should be read as shape and order of magnitude, not precision. No significance test is claimed.
- The cohort is people who chose to have a relationship analysed. It is not a random or representative sample of relationships.
- Direction is defined by who uploaded. Any statistic split by direction inherits that selection.
- Classifier labels are model output, not clinical judgements, and carry the error rate of the model that produced them.
- The cohort is small. Figures show shape and order of magnitude, not precision, and no significance test is claimed.
- Grey Mirror measures patterns in text. It does not diagnose abuse, deception, illness or predict outcomes.
What this does not prove
The most common way a finding like this gets misreported is by turning a property of the uploaders into a claim about relationships in general. These are the specific inferences the data cannot support.
- It does not show that hostility is rare in relationships. Blame, distress and threat did not reach a measurable rate here, and harassment ran 0.032 per 1,000 messages. That is a property of who uploads a history to a consumer tool, not evidence about the population.
- It does not show that one side texts more than the other in general. The uploader sent about 53% of messages, with a range across accounts of 46.5 to 59.3%. Direction is defined by who uploaded, so any direction-split statistic inherits that selection.
- It does not show that either side replies faster. Threads split 9 to 13 across 22 decided threads, which is why no reply-speed claim is made.
- It does not diagnose abuse, deception, manipulation or mental illness, and it does not establish intent. Classifier labels are model output carrying that model error rate, not clinical or legal judgements.
- It does not predict whether a relationship continues or ends. Nothing in this artifact is an outcome model.
Why this matters
For anyone writing about relationship analysis, the practical implication is that the evidence base almost everyone reasons from is the wrong sample. If roughly 97.4% of a real thread carries no emotional signal, then advice built on saved screenshots is built on the 2.6% that someone found notable enough to keep.
For AI text interpretation, it sets a floor on what a single-message tool can honestly claim. A model shown one exchange has no access to the base rate it came from, so it cannot tell a characteristic pattern from an outlier. That is a limitation of the input, not of the model.
For privacy, it is the reason this artifact is aggregate-only. A message history is among the most sensitive things a person owns, and publishing per-thread rows would expose contributors to re-identification. Anonymisation happens before scoring, and small cohorts are suppressed rather than shown.
For screenshot culture more broadly, the finding is uncomfortable in a useful way: the artifact people screenshot and circulate is selected for being atypical, and then read as if it were typical.
Cite this
Suggested citation: Weant, L. (2026). Relationship Texting Benchmarks 2026: What Long-Term Message Histories Reveal About Reciprocity, Conflict and Repair. Wentropy Labs / Grey Mirror. https://justlay.me/research/relationship-texting-benchmarks-2026
Published 2026-08-22. Dataset: 4,600,611 messages across 29 de-duplicated two-party histories from 20 accounts. Data version relationship-texting-benchmarks-v2, code version 0bc2ae04. The full aggregate artifact, including every withheld figure and its reason, is published as JSON at https://justlay.me/benchmarks/relationship-texting-benchmarks.json under CC BY 4.0.
When quoting any figure from this study, the correct basis line is: 29 de-duplicated histories totalling 4,600,611 messages from 20 accounts. Press and research enquiries go to press@justlay.me.
Related reading
- Full-thread relationship text analyzer - the method these benchmarks come from
- Screenshot analysis versus full-thread relationship text analysis
- Grey Mirror methodology - the full pipeline
- Evidence standards - what a claim must carry before it is reported
- Metrics library - every measure defined
- Privacy and deletion - how uploaded histories are handled
- Red flag analysis across a full thread
- Response time patterns and timing drift
- Interactive sample report
- The Signal Is Rare - the companion conflict-cohort study
Relationship Texting Benchmarks 2026. Aggregate benchmarks from 4,600,611 messages across 29 de-duplicated histories: 97.4% carried no emotional signal, plus the six measures we withheld.
The page explains its subject, evidence quality, and the appropriate next step.
Data, methodology and charts
Every figure on this page is reproducible from the files below. They are published under the same open terms as the study itself.
Start here
- Analyze your messages — upload a full exported history on the web and watch the first scenes free.
- Grey Mirror for iPhone — the free native iOS app, App Store id 6799236359, iOS 18 or later. Same exports, same report as the web.
- Pricing — the complete report is a one-time $25 unlock, no subscription.
- The Signal Is Rare — conflict outran repair 36 to 1, 97.4% of messages carried no detectable signal, and one claim was withheld.
More from Grey Mirror
View the canonical Relationship Texting Benchmarks 2026 page