Every metric should answer a real relationship question.
Grey Mirror metrics are meant to turn a full message history into inspectable signals. A metric is useful only when it has enough evidence, a clear definition, and confidence labels that prevent over-reading.
What metrics does Grey Mirror use?
Grey Mirror uses relationship signal families such as repair rate, rupture versus repair, response time, timing drift, turn-taking, effort imbalance, positivity reciprocity, power dynamics, future planning, communication shifts, emotional momentum, recurring loops, and evidence confidence.
How should users read the metrics?
Metrics are not verdicts. They are structured signals that become more useful when they agree with each other across the timeline and less useful when the evidence is thin.
- Prefer trend direction over one isolated count.
- Prefer participant-aware metrics over report-wide averages.
- Prefer supported evidence over confident-sounding summaries.
- Treat low-confidence metrics as missing evidence, not bad news.
Metric definitions
The public metric library defines what each metric means, what question it answers, how it is roughly measured, and what evidence context it needs.
Relationship health score
A composite report-level read of stability, repair, reciprocity, timing, and risk signals when enough data exists.
Is the thread moving toward steadier connection or toward more strain?
Combines normalized metric families rather than treating one dramatic message as the whole answer.
Best read as a multi-signal score with supporting evidence windows and confidence labels.
Needs enough timeline coverage, both participant labels, and stable parsed timestamps.
Repair rate
How often conflict or tension is followed by useful de-escalation, accountability, or reconnection.
Do conflicts actually recover, or do they restart with the same unresolved pattern?
Looks for rupture windows, follow-up windows, apology or accountability language, and return-to-baseline signals.
Best with complete exchanges around disagreement and aftermath; offline repair lowers evidence density.
Best with complete exchanges around disagreement and aftermath.
Rupture vs repair
The balance between moments that damage trust and moments that restore it.
Is the relationship absorbing stress or accumulating unresolved damage?
Contrasts tension markers, escalation, blame, invalidation, repair bids, and follow-through windows.
Best when rupture and repair both appear in the text record with clear before, during, and after context.
Needs enough before, during, and after context around conflict.
Response time
Reply latency patterns by participant, time period, and context.
Who tends to wait, who tends to chase, and whether the pace has shifted.
Uses timestamp gaps after normalization and excludes gaps that are not countable.
Best with reliable timestamps, intact sequence, and enough repeated windows to separate ordinary schedule from drift.
Requires reliable timestamps and correct participant mapping.
Timing drift
Whether reply patterns get faster, slower, more selective, or more asymmetric over time.
Did the relationship cool gradually, or did one person suddenly change cadence?
Compares timeline windows and participant-level latency distributions.
Best with long histories and consistent timestamp coverage so slower pacing can be compared across phases.
Best with long histories and consistent timestamp coverage.
Turn-taking
How evenly the conversation passes between participants instead of collapsing into monologues or pursuit.
Does one person carry the thread, explain more, or repeatedly reopen the same issue alone?
Compares message runs, response density, unanswered streaks, and follow-up burden.
Best when topic type, response quality, and repeated back-and-forth are visible.
Needs accurate speaker attribution and enough back-and-forth.
Effort imbalance
Differences in initiation, follow-up, emotional labor, planning, and repair attempts.
Is the thread reciprocal, or is one person maintaining the relationship by default?
Combines initiation, unanswered bids, follow-through, planning, repair, and engagement markers.
Best when ordinary days, conflict days, planning days, and follow-through all appear in the upload.
Best when the uploaded history includes ordinary days, conflict days, and planning days.
Positivity reciprocity
Whether warmth, affirmation, humor, appreciation, and bids for connection are returned or ignored.
Does positive energy circulate, or does it mostly travel one way?
Counts and compares positive bids, reciprocal responses, and ignored warmth across windows.
Best with surrounding message context around affectionate, humorous, sarcastic, or supportive exchanges.
Needs message-level context around affectionate or supportive exchanges.
Power dynamics
Signals of pressure, control, invalidation, apology avoidance, boundary testing, or one-sided decision control.
Does one person repeatedly shape the terms of the relationship while the other adapts?
Looks for repeated pressure patterns, dismissal, coercive wording, threats, blame loops, and boundary response.
Best with repeated pattern evidence, recurrence, intensity, and trust-page resources for safety-related concerns.
Requires repeated pattern evidence, not a single phrase.
Future planning
How often concrete future plans appear, who initiates them, and whether they become action.
Is future talk specific and mutual, or vague and noncommittal?
Separates concrete logistics, time-bound plans, vague someday language, and follow-through.
Best with enough planning conversations, follow-through windows, and platform coverage to compare.
Needs timestamps and enough planning conversations to compare.
Communication shifts
Changes in length, tone, response cadence, conflict style, affection, or planning across the timeline.
What changed, when did it change, and which signals changed together?
Compares windowed metric changes and participant-level trend direction.
Best with multiple months or distinct relationship phases so timeline shifts can be compared.
Best with multiple months or distinct relationship phases.
Emotional momentum
The direction of emotional movement across windows when sentiment, repair, warmth, and tension align.
Is the relationship gaining connection, flattening, or moving toward shutdown?
Uses multiple signal families together instead of one sentiment counter.
Best with enough messages per window and surrounding context for emotional language.
Needs enough messages per window for confidence.
Recurring loops
Repeated relationship episodes that follow similar sequences, such as tension, apology, temporary warmth, and relapse.
Are you solving a problem or replaying the same episode with new words?
Compares episode structure, triggers, sequence order, and aftermath when sequence intelligence is available.
Best with repeated episodes that have clear starts, follow-up, and aftermath windows.
Needs repeated episodes with clear starts and aftermath windows.
Evidence confidence
A label that explains whether a metric is strongly supported, partially supported, or too thin to trust.
How much should the user rely on this particular claim?
Considers message count, window coverage, parser quality, participant confidence, and supporting evidence density.
Low confidence is not a negative finding; it means the evidence needs more support.
Uses parser diagnostics, timeline coverage, evidence references, and metric support.
Which metrics need the most evidence?
Power dynamics, recurring loops, emotional momentum, and relationship-health style scores should be treated carefully because they combine multiple weaker signals into a stronger claim only when the evidence is dense.
How do the metrics become a report?
The report should not dump a wall of analyzer names. It should translate the metric stack into a small number of useful reads: what is stable, what is drifting, what repeats, and which signals deserve caution.
Scorecard layer
A compact read of overall health, repair, affection, timing, and balance. It is useful for orientation, not final judgment.
Dashboard layer
Expandable metric families for users who want to inspect why a score moved instead of trusting a headline number.
Evidence layer
Windows, counts, confidence, and report notes that explain where the metric came from and when confidence should be lower.
What should a strong metric page answer?
A useful metrics library should own the exact questions people ask before they upload: what is measured, how to read it, what can go wrong, and how the metric connects to the report.
- Repair answers whether conflict is followed by accountability, de-escalation, or the same loop again.
- Timing answers whether silence and response gaps are random life friction or a repeated relationship pattern.
- Effort answers whether one participant carries initiation, follow-up, planning, and repair more often.
- Trajectory answers whether several signals move together across the timeline instead of changing in isolation.
- Confidence answers whether the evidence is strong enough to name the pattern plainly.
Frequently Asked Questions
Are counts enough?
No. Counts help, but Grey Mirror should also consider timing, sequence, participant role, surrounding messages, and confidence.
Why does a metric sometimes say insufficient evidence?
Because a trustworthy metric needs enough parseable messages, timeline coverage, and participant confidence to support the claim.
Can users compare metrics across runs?
When later runs are linked to the same report lineage and contain comparable data, longitudinal surfaces can show change over time.
How do you measure reciprocity in text messages?
Reciprocity is measured as a set of paired ratios across the whole thread rather than a single number: who opens conversations, who sends the second message after a reply, who asks questions back, and who closes threads. Each ratio is compared against that pair’s own historical baseline, because a 60/40 split is normal for some couples and a sharp change for others. The direction and stability of the ratio over time matters more than its value in any single week.
How do you measure effort balance in texts?
Effort balance combines four counts per participant: conversation initiations, follow-up messages after no reply, questions asked, and concrete plans proposed. Message volume alone is a poor proxy, because one person can send many short reactive messages while the other does the structural work of starting and planning. Grey Mirror reports each component separately so a single lopsided component does not read as global imbalance.
How do you measure emotional reciprocity in texts?
Emotional reciprocity looks at what happens after one person raises something vulnerable. The measurable unit is the response to a disclosure: whether it was acknowledged, matched with a disclosure of similar depth, met with a follow-up question, deflected, or left unanswered. Tracking that response distribution over time shows whether emotional openness is met symmetrically, and whether the discloser reduces openness in response to how it lands.
How do you measure positivity reciprocity in texts?
Positivity reciprocity measures whether warmth is returned in kind and how quickly. It counts appreciation, affection, humour, and encouragement per participant, then examines whether a warm message is typically answered with warmth, with neutral logistics, or with nothing. The ratio matters less than its trend: a stable imbalance is a different signal from one that widened over the last several months.
How do you measure repair attempts in texts?
A repair attempt is a message that tries to de-escalate or reconnect after a rupture: an apology, a check-in, an acknowledgement, or humour used to break tension. Measurement requires two parts — how often repair is attempted, and whether it was received, meaning the other person engaged rather than continued the conflict or went silent. Repair success rate is the more predictive of the two, because frequent unreceived repair attempts indicate a stuck cycle rather than a healthy one.
How do you measure blame shifting in texts?
Blame shifting is measured sequentially, not by keyword. The pattern is a specific move: a complaint is raised, and the response reframes the complaint as a fault of the person raising it rather than engaging with it. Counting how often that move follows a complaint, and whether the original topic ever returns, distinguishes a defensive moment from a recurring structure. Grey Mirror reports the sequence and its recurrence rather than labelling intent.
References and methodology
Related Grey Mirror guides
- Relationship text analyzer
- Methodology
- Public white paper
- Metrics library
- Evidence standards
- Privacy and deletion
- AI sycophancy vs measured analysis
- Interactive sample report
- Relationship text analysis glossary
- Long-term pattern analysis
- Love language in texting
- iMessage analysis
- WhatsApp chat analysis
- Instagram DM analysis
- ChatGPT vs Grey Mirror
- Screenshots vs full thread
- Repair attempts in texting
- Conflict escalation patterns
- Emotion word frequency
- Texting anxiety signs
- Friendship text analysis
- Telegram text analysis
- Improve text communication
- Apology insufficiency case study
- Full thread vs screenshot case study
- Couples text message analyzer
- Analyze chat history for patterns
- SMS and Android text analysis
- Pricing and free preview
- Private relationship text analyzer
- Best relationship text analyzer
- Best text message analyzers 2026
- Chat analyzer comparison
- Red flag text analyzer
- Situationship text analyzer
- Analyze relationship texts
- Relationship pattern analysis case study
- Criticism in texts case study
- Dismissiveness case study
- Emotional availability case study
- Emotional labor case study
- Emotional tone drift case study
- Talking about problems case study
- Mixed signals case study
- Post-conflict patterns case study
- Power dynamics case study
- Reading subtext case study
- Validation in texts case study
View the canonical Every metric should answer a real relationship question. page