02

Methods

Prediction and replication in marital conflict research

A set of laboratory studies reported that coded video of a fifteen-minute disagreement could forecast separation years later, at accuracies in the high nineties. The figure travelled widely. The methodological reply travelled less.

Published 2 July 2026 · Sources: 5 · Record 02

The research design is straightforward to describe. Couples are recorded in a laboratory while discussing an area of ongoing disagreement. The video is coded second by second for facial expression, tone and speech content using an observational coding scheme, and physiological measures such as heart rate and skin conductance are collected at the same time. The couples are then followed for years, and the coded behaviour is compared against who is still together at follow-up.

John Gottman and Robert Levenson ran several studies on this design, reporting in 1992 and again in 2000 that combinations of coded variables separated the couples who later divorced from those who did not with very high accuracy. The work also produced a set of descriptive categories for conflict behaviour that entered common use, among them criticism, contempt, defensiveness and withdrawal from the interaction.

The objection

In 2001 Richard Heyman and Amy Slep published a paper in the Journal of Marriage and Family whose title states the problem plainly: the hazards of predicting divorce without crossvalidation. Their argument is statistical rather than substantive.

In the studies at issue, the model was constructed after the follow-up outcomes were known. Variables and cut-off points were selected because they separated the groups well in that particular sample. A model built that way will always describe its own data closely, because the peculiarities of the sample have been absorbed into the model. The test of a prediction is whether it holds on data it has never seen, and that step, crossvalidation, had not been performed.

Heyman and Slep performed it. Working with longitudinal datasets, they built models in the manner described, then applied them to independent samples. Accuracy fell substantially. Some models performed close to what the base rate alone would deliver, which is to say that on a fresh sample they added little beyond knowing how many couples separate in general.

Sample size compounds the issue. Several of the influential studies followed a few dozen couples, largely white, largely middle-class, recruited in one metropolitan area of the United States. With samples that small, the number of parameters available for selection is high relative to the number of cases, which is precisely the condition under which a fitted model absorbs noise.

What survives

The distinction that matters is between description and forecast, and the two parts of this literature have fared differently.

The observational coding systems remain in use and remain useful. Trained coders agree with one another at acceptable rates, which means the behaviours can be measured rather than merely asserted. The associated findings, that negative affect during conflict discussions tends to be reciprocated, and that patterns of that kind correlate with self-reported satisfaction measured at the same time, are consistent across a range of samples and laboratories.

The forecasting claim is a different matter. Accuracy figures generated without crossvalidation describe how well a model fits the sample it was built from, and the popular reporting of that period rarely made the distinction. Where the claim has been tested prospectively, it has not held at the level first reported.

None of this establishes that laboratory observation of conflict is uninformative. It establishes that a figure quoted in a magazine may rest on a step that was never taken, and that the reply to it sits in a journal that fewer people read.

Sources

  1. Gottman, J. M., & Levenson, R. W. (1992). Marital processes predictive of later dissolution: Behavior, physiology, and health. Journal of Personality and Social Psychology, 63(2), 221–233.
  2. Gottman, J. M., & Levenson, R. W. (2000). The timing of divorce: Predicting when a couple will divorce over a 14-year period. Journal of Marriage and Family, 62(3), 737–745.
  3. Heyman, R. E., & Slep, A. M. S. (2001). The hazards of predicting divorce without crossvalidation. Journal of Marriage and Family, 63(2), 473–479.
  4. Heyman, R. E. (2001). Observation of couple conflicts: Clinical assessment applications, stubborn truths, and shaky foundations. Psychological Assessment, 13(1), 5–35.
  5. Stanley, S. M., Bradbury, T. N., & Markman, H. J. (2000). Structural flaws in the bridge from basic research on marriage to interventions for couples. Journal of Marriage and Family, 62(1), 256–264.

Back to the record