Strategy Similarity and Decision Diversification
Measuring whether multiple strategies actually make different portfolio decisions
QM011 · Portfolio Methods · Intermediate
Core idea. Return correlation describes payoff co-movement; decision similarity describes whether strategies make the same underlying portfolio choices.
Use it for. Diagnosing strategy redundancy through separate what, when, and how-much dimensions before or alongside portfolio combination.
It does not establish. A universal scalar “decision-diversification score,” causal independence, or a guaranteed reduction in portfolio risk.
The Question
If a portfolio contains five named strategies, does it contain five genuinely different sources of decision-making?
Not necessarily. Two strategies can have different labels yet hold almost the same assets, change positions on the same dates, and take nearly identical exposure. Conversely, strategies with similar realized return correlation can still reach those returns through different holdings or different timing.
The methodological question is therefore:
How similar are the decisions themselves, rather than only the realized returns?
Why It Matters
Traditional diversification analysis starts from the joint distribution of returns. That is essential: Markowitz’s portfolio framework formalized why covariance matters for portfolio risk (Markowitz 1952). But return correlation is an outcome statistic. It does not by itself identify how two systematic rules generated those outcomes.
For strategy research, that distinction matters because apparent diversification can be redundant at the decision layer. Several sleeves may react to the same price trend, move risk-off together, or maintain nearly identical exposures. Counting strategies then overstates the breadth of the decision architecture.
QM011 treats decision similarity as a diagnostic layer that complements, rather than replaces, return-based diversification.
Intuition
A strategy decision can be decomposed into three questions:
- What? Which assets, directions, or exposure buckets does it choose?
- When? When does it materially change that choice?
- How much? How large is the chosen exposure or adjustment?
These dimensions should usually be reported separately. Collapsing them immediately into one score hides which kind of similarity is driving redundancy.
The what/when/how-much decomposition is a reusable diagnostic framework, not a universally standardized statistic with one canonical formula. The formulas below are transparent choices that can be adapted to the portfolio representation, provided the convention is declared before interpretation.
The Method
Return Similarity Is Not Decision Similarity
For two strategy returns \(r_{A,t}\) and \(r_{B,t}\), the familiar return correlation is
\[ \rho_{AB}=\operatorname{Corr}(r_A,r_B). \]
This tells us how the realized payoffs co-moved. It does not tell us whether the strategies held the same assets or changed positions together.
Decision diagnostics operate on portfolio states or actions. Let \(\mathbf w_{A,t}\) and \(\mathbf w_{B,t}\) denote target-weight vectors at decision date \(t\).
The What Dimension: Holdings or Exposure Similarity
For long-only, fully invested portfolios, a particularly interpretable overlap measure is
\[ O_t=\sum_{i=1}^{N}\min(w_{A,i,t},w_{B,i,t}). \]
Because both weight vectors sum to one,
\[ O_t=1-\frac{1}{2}\|\mathbf w_{A,t}-\mathbf w_{B,t}\|_1. \]
The value ranges from 0 to 1:
- \(O_t=1\) means identical holdings weights;
- \(O_t=0\) means no overlap under the long-only fully invested convention.
The half-\(L_1\) distance resembles the logic used by Active Share to quantify portfolio-weight differences relative to a benchmark (Cremers and Petajisto 2009), although here it is used symmetrically between two strategies rather than as the Active Share statistic itself.
For long/short, leveraged, or factor-exposure vectors, the long-only overlap formula may no longer be appropriate. Cosine similarity, normalized \(L_1\) distance, or category-specific exposure comparisons may be better choices.
The When Dimension: Decision-Timing Disagreement
A strategy should not be classified as “acting” merely because its weights drifted with market returns. The timing diagnostic should be based on intended target changes or another explicit action variable.
Define
\[ I_{k,t}=\mathbf 1\{\|\mathbf w_{k,t}-\mathbf w_{k,t-1}\|_1>\tau\}, \]
where \(k\in\{A,B\}\) and \(\tau\) is a pre-declared materiality threshold. A simple timing disagreement rate is
\[ D_{when}=\frac{1}{T}\sum_{t=1}^{T}\mathbf 1\{I_{A,t}\ne I_{B,t}\}. \]
This statistic asks how often one strategy makes a material target change while the other does not.
It can be refined. For example, when both strategies act, a researcher may separately test whether the direction of risk change agrees. The key is to avoid mixing action timing with weight drift, which belongs to portfolio accounting rather than the decision rule itself.
The How-Much Dimension: Magnitude Disagreement
Suppose \(x_{k,t}\) is a scalar risk-exposure measure, such as equity weight, risky-asset share, duration target, or gross exposure. A simple average absolute disagreement is
\[ D_{mag}=\frac{1}{T}\sum_{t=1}^{T}|x_{A,t}-x_{B,t}|. \]
For multi-asset target weights, an analogous measure is the average half-\(L_1\) distance:
\[ D_{w}=\frac{1}{T}\sum_{t=1}^{T}\frac{1}{2}\|\mathbf w_{A,t}-\mathbf w_{B,t}\|_1. \]
The normalization should match the portfolio domain. A 20-percentage-point exposure difference is interpretable only if the exposure scale itself is meaningful.
Keep a Decision Fingerprint, Not a Forced Composite
A useful summary is a vector such as
\[ \mathcal F_{AB}= (\overline O, D_{when}, D_{mag}), \]
where \(\overline O\) is average holdings overlap.
A high-overlap, low-timing-disagreement, low-magnitude-disagreement pair is strongly redundant at the decision layer. A low-overlap pair with frequent timing disagreement is making visibly different choices.
There is no need to combine the three numbers into a single index unless the research question supplies defensible weights. A scalar composite can create an arbitrary trade-off—for example, treating 10 percentage points of holdings difference as equivalent to one additional timing disagreement.
How to Interpret the Result
Decision diversification is best treated as diagnostic evidence, not as a substitute for portfolio risk analysis.
- High decision similarity can reveal that a multi-strategy architecture is less diverse than its labels imply.
- Low decision similarity can reveal genuinely different portfolio mechanisms.
- Neither result guarantees low or high return correlation in a finite sample.
- Greater decision diversity does not automatically improve Sharpe ratio, drawdown, or tail risk.
The portfolio still needs conventional outcome evaluation. QM011 answers a different question: whether the strategy set is relying on distinct decisions.
Financial / Economic Example
Consider two long-only strategies over four decision dates with weights in Equity and Bonds:
| Date | Strategy A | Strategy B | Overlap |
|---|---|---|---|
| 1 | (0.80, 0.20) | (0.75, 0.25) | 0.95 |
| 2 | (0.80, 0.20) | (0.40, 0.60) | 0.60 |
| 3 | (0.30, 0.70) | (0.40, 0.60) | 0.90 |
| 4 | (0.30, 0.70) | (0.70, 0.30) | 0.60 |
With a material-action threshold of 20 percentage points in full-\(L_1\) target change, Strategy A acts only between dates 2 and 3, while Strategy B acts between dates 1 and 2 and again between dates 3 and 4. Their timing disagreement is therefore high even though some dates have substantial holdings overlap.
The example shows why “same holdings today” and “same decision process over time” are different claims.
The weights are synthetic and demonstrate the diagnostics only.
Implementation
For long-only, fully invested weights:
import numpy as np
def overlap(w_a, w_b):
a = np.asarray(w_a, float)
b = np.asarray(w_b, float)
return np.minimum(a, b).sum()
def half_l1_distance(w_a, w_b):
return 0.5 * np.abs(np.asarray(w_a) - np.asarray(w_b)).sum()For valid long-only fully invested vectors, overlap(a, b) equals 1 - half_l1_distance(a, b). The included validator checks this identity and the timing-disagreement example.
Common Mistakes
Using return correlation as a proxy for identical decisions. It is an outcome statistic, not a holdings/action diagnostic.
Counting strategy labels. Five strategy names do not imply five independent decision rules.
Measuring action from drifted realized weights. Market movement can change weights even if the strategy did nothing. Use target changes or an explicit trade instruction.
Forcing a composite score. What, when, and how much have different units and meanings.
Calling decision diversity “risk reduction.” It may create robustness to rule uncertainty, but the realized portfolio risk still depends on exposures and return covariance.
When Not to Use It
Do not use long-only holdings overlap unchanged for short, leveraged, derivative, or nonlinear exposures. In those settings, choose a representation that captures the economically relevant risk—such as factor beta, delta-equivalent exposure, duration, or scenario P&L sensitivity.
Do not use QM011 as the sole portfolio selection objective. A strategy can be wonderfully different and economically poor. Decision diagnostics should sit beside performance, risk, cost, and robustness analysis.
Used in SlackQuant Research
- Diversify the Decisions, Not Just the Assets: applied strategy-level analysis distinguishes return similarity from disagreement in what sleeves hold, when they change, and how much risk they take.
QM011 generalizes that reusable diagnostic logic. It is not a summary of the ADAA paper and should be applicable to other multi-strategy architectures.
Reproducibility
The deterministic Python example verifies the long-only overlap identity and computes separate timing and magnitude diagnostics from target weights. No empirical claim depends on the synthetic numbers.