# The Signal Is Rare: methodology and evidence boundaries

Version: August 25, 2026

## Scope

This package combines findings from two observed Grey Mirror research cohorts and one synthetic turning-point benchmark. It does not merge those evidence bases into one sample, one message total, or one claim of real-world accuracy.

The 29-history signal cohort contains 4,600,611 messages from 20 accounts. It supports findings about detectable emotional signal, affection, repair, direction splits, and measures that were deliberately withheld.

The 44-history conflict cohort contains 2,994,932 messages. It supports descriptive findings about escalation, repair attempts, apology attempts, and pursue-withdraw balance.

The turning-point benchmark contains 40 steady generated threads and 40 generated threads with a known planted change. It supports false-positive, detection, direction, and localization measurements only. It contains no uploaded user conversations.

The observed cohorts may overlap. They have different inclusion rules and different reported metrics. Their message totals must not be added.

## Data provenance audit

The August 25 Desktop export adds provenance information without expanding any research cohort.

Its analytics metadata records 347 rollup conversations totaling 20,794,571 messages across 88 accounts. The export also records minimum eligibility settings of 500 messages and a 30-day span, 126 eligible conversations, and the final 29-history signal cohort. The final cohort is therefore presented as a curated subset of a larger operational rollup, not as all messages Grey Mirror has processed.

The upload manifest contains 661 stored-file rows and 470 unique content hashes. The other 191 rows repeat an existing hash. These are storage records, not relationship histories, but they demonstrate why counting files is not the same as counting independent evidence.

The export also contains 33 quality-notification records. Overlapping reason codes include 24 missing-timestamp occurrences, 11 unsupported-or-empty-export occurrences, 8 missing-sender-direction occurrences, and 1 low-confidence-import occurrence.

Seventy-three cached import-quality objects collapse to 45 distinct conversation IDs. Thirty-three were marked safe to analyze and 12 were marked unsafe. Repeated cache windows were not counted as additional conversations.

The 89-row operational run table was excluded from research findings because only 2 rows carried core metric values. The 325-conversation reply-latency rollup was also excluded because it is not joined to the published study's eligibility and de-duplication decisions.

The provenance audit is not a fourth cohort. None of its totals are added to the signal cohort, conflict cohort, or synthetic benchmark. Full notes are provided in `DATA-EXPORT-AUDIT.md` and `data-export-audit.csv`.

## Observed signal cohort

The published benchmark includes 29 de-duplicated two-party histories totaling 4,600,611 messages from 20 accounts. One account contributed nine qualifying histories, so rates were computed per account before being combined. This prevents that uploader from dominating the combined rate.

Published findings used in this package:

- About 97.4% of messages carried no detectable emotional signal.
- Affection outnumbered explicit repair attempts by about 10.7 to 1.
- The person who requested the analysis sent 53% of messages when computed per account and combined.
- Affection was higher outbound in 14 of 20 accounts.
- Affection rates were 22.463 outbound and 20.819 inbound per 1,000 messages.
- Reply-latency direction split 9 to 13 on which side replied more slowly. No universal direction claim was made.

The 2.6% detectable-signal figure in the package is the arithmetic complement of the published 97.4% no-detectable-signal figure. It is labeled as derived in the data file.

Six measures were published as not reported:

1. Criticism direction, because the split was treated as a measurement artifact pending direction-blind rescoring.
2. Manipulation direction, for the same reason.
3. Repair direction, because the account majority and the combined rate pointed in opposite directions.
4. Blame, because its rate rounded to zero and was too rare in this cohort to report.
5. Reply latency, because the directional split was 9 to 13 and did not establish a consistent direction.
6. Thread duration, because repeat uploads clustered on a few exact spans and would make the median describe the upload pattern rather than the relationships.

## Observed conflict cohort

The published study contains 44 complete conversation archives totaling 2,994,932 messages. The median thread contains 1,570 messages and the largest contains 163,636 messages.

Every event rate is expressed per 1,000 messages so histories with different lengths can be compared. Medians and interquartile ranges are used instead of means because thread length is heavily skewed.

Published descriptive values:

| Measure | Median per 1,000 | Interquartile range |
| --- | ---: | ---: |
| Escalation | 195.3 | 52.6 to 226.2 |
| Repair attempt | 5.43 | 3.80 to 7.64 |
| Apology attempt | 4.31 | 0.91 to 4.59 |

The reported 36 to 1 figure is an approximate ratio of the published escalation and repair medians.

Only 5 of the 44 histories showed repair occurring more often than escalation. The other 39 count is a direct arithmetic complement and is labeled as such in the data file.

Pursue-withdraw distribution:

- 28 histories skewed toward pursuit, published as 64%.
- 10 skewed toward withdrawal, published as 23%.
- 6 were balanced, published as 14%.

The published rounded shares sum to 101%. The package preserves the published counts and calls out the rounding instead of silently changing a percentage.

The conflict study describes escalation as emotional-momentum spikes across the thread rather than a simple keyword count. Repair follows behavior after conflict rather than counting the word sorry. Apology attempts are reported separately.

## Synthetic turning-point benchmark

The benchmark uses 40 trials per arm. One arm contains generated threads with no real change. The other contains generated threads with a planted change of known date and direction.

Published results:

- 0% false positives across 40 steady synthetic threads.
- 100% detection across 40 synthetic threads containing a known change.
- 100% correct direction on the planted-change arm.
- Median localization error of 0 days.

The detector uses ten behavioral signals per time bucket: warmth, enthusiasm, curiosity, reciprocity, responsiveness, repair, vulnerability, conflict, investment, and initiation. Signals are standardized with a robust median and median absolute deviation. Candidate changepoints come from PELT over a summed multivariate cost. Candidates must survive a circular block permutation significance test. Threads shorter than 14 days are refused.

These results demonstrate behavior on generated data where the answer is known. They do not promise the same rates on real conversations. Detecting a measurable change does not explain why it happened.

## Separate engineering retrieval benchmark

The export includes a promoted evidence-retrieval pass on one full-fidelity 163,636-message artifact. It evaluated 30 metric outputs, 11 evidence-backed metric families, and 40 known evidence message IDs.

Mean recall was 0.8788 at 5, 8, and 16 retrieved results and 1.0000 at 32. Full evidence coverage was 0.7273 at 5, 8, and 16 and 1.0000 at 32. Recorded throughput was 328.6001 messages per second in that benchmark environment.

This test evaluates whether the retrieval layer can recover message IDs already linked to metric evidence. It does not evaluate classifier accuracy, relationship truth, clinical validity, causality, or relationship outcome. It is not added to any research cohort and it is not the synthetic turning-point benchmark.

The exported taxonomy artifact contains 142 base communication labels. The label count describes system scope, not accuracy. Multiple exported model checkpoints have different evaluation profiles, so this package makes no blanket model-accuracy claim.

## What these findings cannot establish

- The observed histories are self-selected and are not a random sample of relationships.
- The cohorts are small and support descriptive shape and order of magnitude, not population precision.
- Text does not capture calls, in-person interactions, deleted messages, or conversations on other apps.
- The data does not contain relationship outcome labels and does not predict whether a relationship continued, improved, or ended.
- Observable behavior does not establish private intent.
- Classifier labels are model output, not clinical judgments and not validated psychometric instruments.
- Synthetic accuracy is not real-relationship accuracy.
- Evidence-retrieval recall is not classifier accuracy or relationship validity.
- Operational archive counts are not independent relationship counts.

## Privacy and release format

This package publishes aggregate findings only. It contains no raw private messages, names, contact details, individual timelines, or individually identifying thread records. Every graphic includes its own cohort label and claim boundary so it can be reused without losing essential context.

## Canonical sources

- Signal benchmark: https://justlay.me/research/relationship-texting-benchmarks-2026
- Conflict and repair study: https://justlay.me/grey-mirror/relationship-conflict-repair-study
- Turning-point benchmark: https://justlay.me/grey-mirror/turning-point-accuracy-benchmark
- General methodology: https://justlay.me/grey-mirror/methodology
- Metrics library: https://justlay.me/grey-mirror/relationship-text-analysis-metrics
- Evidence standards: https://justlay.me/grey-mirror/evidence-standards
