Skip to main content
EN
Back to Insights
Markets

Measuring Payments Cohort Cohesion: A Market-Adjusted Methodology

The quantitative methodology behind the cohesion read the Pulse reports each week: market-adjusted residual correlation, a leave-one-out reclassification test, and a full accounting of where the method fails.

FDP
Franco Di PietroThe Payments Corner
August 3, 202612 min readLinkedIn
Audio Briefing
0:00/12:25
Infographic showing payments and fintech stocks moving together before diverging, with panels illustrating correlation, regression, time-varying R² and network structure.
Cohesion or divergence? A quantitative framework for identifying when payments and fintech stocks move as one category—and when they begin to separate into individual company stories.

The Payments Corner — methodology paper

The Payments Corner groups its payments-equity universe into six cohorts and, for each, reports whether the cohort is trading as a single bet or diverging name by name. This paper documents exactly how that judgment is computed: the statistic, the two adjustments that make it honest, the reclassification test built on top of it, the guardrails, and the limits. The design goal throughout is that every step map to a standard, named technique in empirical finance, so the measure is defensible rather than bespoke.

Provenance and contribution

The methods below are standard, and deliberately so. None of the individual steps is proprietary. Working in returns rather than price levels, to avoid manufacturing correlation from shared trend, is the spurious-regression result of Granger and Newbold (1974). Regressing a name's return on the market's and keeping the residual is the market model that underlies the Capital Asset Pricing Model (Sharpe, 1964) and the abnormal-return residual of the classic event study (Fama, Fisher, Jensen and Roll, 1969). Correlating those residuals to isolate co-movement beyond the market is ordinary factor-model risk decomposition — the systematic-versus-specific split that every commercial risk model performs. Excluding a name from the factor it is scored against is leave-one-out, a generic bias correction. A factor-risk desk does all of this as a matter of routine. The durability of the measure rests on that fact, not on any secret construction: every step maps to a named, defensible technique, and a reader who wants to audit it can.

What The Payments Corner contributes is the construction and the application, not the mathematics:

  • The construction — a leave-one-out, equal-weighted cohort residual factor used as a taxonomy-validation test. Each name is scored by how well its residual fits its assigned cohort versus every other cohort's, and a persistent better fit elsewhere is a reclassification signal. It is applied to a curated payments universe, with stated thresholds, under the framing that a cohort is a testable hypothesis rather than a fixed label.
  • The application — turning a risk-desk technique into a reader-facing read: whether a repricing across payments names is structural (the cohort moving as one bet) or name by name (the label doing less work than the individual stories inside it).

The rest of this paper documents each step, and where each one fails.

1. The object: daily returns, never price levels

The measure is built on daily returns — r_t = (P_t − P_{t−1}) / P_{t−1} — not prices.

This is the single most important choice, and it is the difference between a real statistic and a classic trap. Correlating two price series almost always yields a correlation near +1 — even for unrelated companies — because both are non-stationary: they trend, they carry a unit root. This is the Granger–Newbold spurious-regression result: trending series manufacture correlation out of nothing. Two random walks drifting upward will read as roughly 95% correlated and mean nothing.

Returns resolve this. Returns are approximately stationary — roughly constant mean and variance over time, with no trend to manufacture agreement. A correlation of return series is a real statement about co-movement, not an artifact of two series both rising over a year. Every serious market-structure and risk model works in return space for exactly this reason.

Working in returns also raises the stakes on clean data. A stock split, left uncorrected, injects a single fabricated return — a large "move" on the split date that never happened — into an otherwise clean series. In price terms that is a visible cliff; in return terms it is one catastrophic outlier that contaminates every correlation, beta, and fit the name touches, and drags its cohort's factor along with it. Split-adjusting the price history is therefore not cosmetic housekeeping; it is a precondition for the cohesion statistics to mean anything.

2. The core problem: raw correlation overstates cohesion

Correlating cohort members' raw returns would yield high numbers — but misleading ones. On any risk-on or risk-off day, everything moves together: a network, a Latin American fintech, and a crypto rail all rise three percent because the market rose three percent. Raw return correlation conflates two distinct facts:

  1. Names move together because they are the same kind of business (the quantity of interest); and
  2. Names move together because the market moved (noise to be removed).

Any equity-only universe shows inflated raw cohesion simply because it inherits the market's common factor. Reporting that as "the cohort is moving as one trade" would be a category error.

3. The fix: residualization (the market adjustment)

The market is removed from each name using a single-factor market model — the empirical regression that underlies both the Capital Asset Pricing Model and the base case of every factor model in empirical finance. (The market model is the estimated regression, carrying an intercept α; the CAPM is its equilibrium counterpart, which predicts α = 0. The two are related but distinct, and it is the market model — the regression — that is used here.)

r_i = α_i + β_i · r_market + ε_i        (market proxy = S&P 500)

The name's beta β_i is estimated by ordinary least squares of its returns on the market's returns; the residual is then formed:

ε_i = r_i − β_i · r_market

The residual is the part of the name's daily move the market does not explain — its idiosyncratic, cohort-specific component. From there:

  • Cohesion = the average pairwise Pearson correlation of residuals across the cohort's members.
  • Market beta (the reported cohort statistic) = the mean of the members' betas — how much of the cohort's motion is simply market exposure.

The measure now reports what its label claims: co-movement beyond the market tide, i.e. payments-specific co-movement. Because correlation mean-centers each series, the intercept α is immaterial to cohesion — only the β · r_market term matters, and that is exactly what is removed.

The effect is not marginal. In a controlled test, a synthetic name constructed to be pure market beta and nothing else shows a raw return correlation to a cohort of about 0.94 and a residual correlation of about 0.06. That gap is the entire argument: 0.94 is the tide, 0.06 is the truth. A cohesion measure that cannot separate the two is measuring the weather, not the cohort.

4. Per-name fit and reclassification: leave-one-out

Cohesion describes the group; the reclassification question concerns the individual. For each member, its residual series is regressed on the cohort's equal-weighted residual factor, and the fit — the R² — is read as the fraction of the name's idiosyncratic variance explained by the cohort's idiosyncratic factor. A high R² means the name's payments-specific moves are the cohort's payments-specific moves — it belongs. A low R² (below 0.40) marks a diverger.

One failure mode must be designed out. If the cohort factor includes the name being scored, the name is correlated with itself by construction and its fit is inflated — every name then "belongs" to whatever bucket it was placed in, and the audit can never flag anything. The correction is leave-one-out: when a name is scored, it is excluded from the factor it is measured against. This is standard cross-validation hygiene against target leakage, and it is what makes a "fits another cohort better" flag trustworthy — a name earns a reclassification candidacy only by fitting a different cohort's leave-one-out factor better than its own.

5. Equal weighting

The cohort factor is the equal-weighted mean of member residuals, not capitalization-weighted. Under cap weighting the largest members would dominate — "Card Networks" would collapse into "does this name track Visa" — and the factor would cease to represent the cohort as a set of peers. Equal weighting represents the cohort as peers. The trade-off is more noise, since a small, volatile name receives an equal vote, which is precisely why the sample gate below matters.

6. The honesty guards

  • Sample gate (≥15 trading days). Below roughly fifteen observations, the sampling error on a correlation or a beta is large enough that the estimate is closer to noise than to signal; the sampling distribution of a correlation coefficient is wide at small samples. Windows that cannot clear the gate are shown but flagged, with the reader directed to the shape of the chart rather than the printed number.
  • Fixed, stated thresholds. Cohesion at or above 0.70 reads as moving together; 0.40 to 0.70 as mixed; below 0.40 as diverging. A member below a 0.40 fit is an outlier. Note that these are two distinct tests that share a cutoff by convention, not by construction: the 0.40 floor on cohort cohesion (the average pairwise residual correlation of §3) measures whether the category as a whole is moving together, while the 0.40 floor on an individual name's R² fit (the leave-one-out regression of §4) measures whether a single name belongs to it. Fixed thresholds make the classification reproducible rather than eyeballed, and let a reader disagree with a specific cutoff rather than with an opaque verdict.
  • Dispersion. A second, independent read on spread — the cross-sectional standard deviation of window returns across full-window members — that does not depend on the correlation math at all.

7. Why the method is durable

The method is durable because every step is the standard tool for its job rather than a bespoke heuristic:

  • Return space yields stationarity, hence non-spurious correlation (Granger–Newbold).
  • Single-factor residualization is how a risk model decomposes specific risk; stripping common factors and analyzing residual co-movement is textbook.
  • OLS beta is the most-studied, most-robust estimator in finance.
  • Leave-one-out is generic anti-leakage cross-validation.
  • Sample gates and fixed thresholds make the output reproducible and self-documenting.

Nothing here is exotic. It is the machinery a factor-risk desk uses, applied to a taxonomy question. That is the source of the durability: an audit of the method maps every choice to a named, defensible technique.

8. Limits and failure modes

Durability includes naming what the method cannot do.

  1. Single-factor only. The market (S&P 500) is stripped; style, size, and sector factors are not. Residual co-movement can therefore still reflect a shared style — "unprofitable, high-beta fintech" is itself a factor, and names can cohere on that rather than on being genuinely the same payments business. A multi-factor residualization — adding a technology factor, or size and growth factors — would isolate the payments-specific signal more cleanly. It is the natural next refinement.
  2. Beta is assumed constant across the window. Real beta drifts, particularly over a year and particularly for names that repriced hard through a rate cycle. A rolling beta would track that drift rather than average over it.
  3. A single window is a candidate, not a verdict. One window's "fits another cohort better" flag can be noise; the signal is persistence — a name that diverges for six consecutive months, not one. The reclassification surface accordingly states "confirm it holds" rather than moving anything automatically, and persistence tracking is the intended companion to the single-window read.
  4. Pearson correlation is sensitive to fat tails. A single earnings-gap day can swing a correlation. A rank (Spearman) correlation would serve as a robustness check.
  5. History depth bounds the long windows. With roughly one year of clean history, the longest offered window is one year; multi-year windows are deferred (see §9).
  6. Cohesion is descriptive, not causal. It reports that names moved together beyond the market; it does not report why.

The two refinements that would make the method more surgical — multi-factor residualization, and rolling beta with persistence — are already identified as future work. Neither changes the foundation; both refine it.

9. Window length: is five years better than one?

Not automatically. Window length is a bias–variance and stationarity trade-off, not a "larger sample wins" question.

The case for a longer window. Correlation and beta estimated on roughly 1,250 days carry far tighter standard errors than on roughly 250 — less sampling noise, more stable numbers. A multi-year window also spans several rate cycles and risk regimes, so the cohesion is less a function of one year's tape.

The case against, for this application. The cohort question is which names move together now. Relationships drift: business models change, beta drifts, and cohort membership itself shifts. A company at an IPO-era, zero-rate valuation and the same company several years later are, for these purposes, different instruments. A five-year correlation averages across those structural breaks and returns a blend of dead regimes; for a "current membership" answer, that is bias, and it can exceed the noise it was meant to reduce. Composition compounds the problem: over five years a universe accumulates splits, rebrands, spinoffs, and acquisitions, and a large share of the most relevant payments names — the recent listings — do not have five years of history at all. A five-year window would therefore either drop the newest, most-in-flux names or run on ragged, uneven samples, which are exactly the names most in need of classification.

The sweet spot. Medium windows — six months to one year — are long enough to clear the sample gate and average out day-to-day noise, yet short enough to reflect the current regime. One week is too noisy; five years is too stale and carries the composition gap.

Better than any single length. The superior answer is not a longer static window but a rolling correlation through time with a persistence requirement: observing how cohesion evolves and acting only on relationships that hold across windows. That yields the stability of more data without the staleness of collapsing it into one number — and it is what makes the reclassification flags trustworthy.

FDP

Franco Di Pietro

The Payments Corner

30+ years across payments, fintech, banking, and financial infrastructure. Operator-level perspectives on the systems that move money.

Share:

Related Insights