| **Chicago Global | Quantitative Research** |
Torrex constructs thematic equity baskets through unsupervised statistical discovery. Unlike traditional sector classifications (GICS, ICB) or imported factor models (Fama-French, Barra), baskets emerge from how stocks actually co-move within each market. The approach is deliberately bottom-up: let local market structure speak for itself, then label what it reveals.
The core pipeline runs in five stages:
Baskets are reformed quarterly (every 90 days) and scored daily. The initial production rollout covered fourteen markets. Coverage now spans 41 markets (see Section 2.3).
All price and volume data comes from LSEG Datastream (Refinitiv). The primary identifier is the Reuters Instrument Code (RIC). Stock metadata (country, listing status, name) comes from the same vendor’s reference data.
Only primary listings are included (DS_ISPRIMQT = 1), which eliminates duplicate cross-listings and depositary receipts.
The live production universe uses a per-market market_filter plus two runtime gates defined in src/torrex/data_config.py and src/torrex/data_loader.py:
min_trading_days = 200 for every live marketfilter_universe_df() keeps the top 50% of tickers by within-market median notional ADV (VOLUME * CLOSE)There is no live absolute USD market-cap, average-volume, or price floor. The current code uses a relative, currency-agnostic liquidity gate because legacy min_*_usd thresholds were compared against local-currency MKTCAP, which miscalibrated non-USD markets. The last China-only min_*_usd stubs were removed from UNIVERSE_FILTERS because the runtime never read them.
Telemetry covers 41 equity markets:
| Region | Markets |
|---|---|
| Americas | United States, Canada, Mexico, Brazil, Chile |
| Europe | United Kingdom, Germany, France, Netherlands, Belgium, Sweden, Norway, Denmark, Finland, Switzerland, Austria, Italy, Spain, Portugal, Poland, Hungary |
| Middle East and Africa | Saudi Arabia, United Arab Emirates, Kuwait, Qatar, Israel, South Africa |
| Asia-Pacific | Japan, China, Hong Kong, South Korea, Taiwan, Singapore, Malaysia, Thailand, Indonesia, Philippines, Vietnam, India, Australia, New Zealand |
A market is included when it meets three criteria:
The initial launch covered 14 markets. Japan, China, Hong Kong and Vietnam were added once vendor price coverage met these criteria. Russia is not covered, because LSEG coverage ended under sanctions. Turkey is not covered, because its historical price data has quality gaps.
The clustering input is excess returns – individual stock returns minus the equal-weight market return on each day. Stripping broad market beta before clustering is critical: without it, the dominant cluster would always be “the market itself,” and all stocks would appear correlated simply because they share exposure to the same macro environment.
Returns are calculated as log returns: r_t = ln(P_t / P_{t-1}). Log returns compound additively (sum rather than product), which simplifies momentum calculations and avoids numerical overflow issues with long compounding chains. Period returns are recovered via np.expm1(sum(log_returns)).
The clustering lookback window is 504 trading days (approximately two years). This is long enough to capture secular co-movement patterns while remaining responsive to structural shifts in market relationships.
The pairwise correlation matrix is the foundation of all clustering decisions. Two methodological choices are worth noting:
Exponentially weighted correlations. By default, the system uses EWM correlations with a 126-day half-life (approximately six months). Recent co-movement receives higher weight than distant history. This reflects the reality that market structure evolves – a basket of “China export chain” stocks may behave very differently after a trade policy shift than before it.
Signed distance metric. Correlation is converted to distance via d = (1 - rho) / 2, mapping the range [-1, +1] to [1, 0]. This preserves the information that negatively correlated stocks are maximally dissimilar. The alternative (1 - |rho|) would treat strong negative correlation as proximity, which would group natural hedges together – useful for pairs trading, but wrong for rotation analysis.
Standard k-means or flat hierarchical clustering requires specifying the number of clusters upfront. This is problematic across heterogeneous markets: the Korean market may have 25 natural groupings while Saudi Arabia has 12. Forcing a fixed k onto different markets either over-fragments small markets or under-segments large ones.
Torrex solves this with recursive coherence clustering: start with the full universe as one cluster, test whether it has statistically significant internal structure, and if so, split it. Repeat on each child until every cluster is either (a) internally coherent or (b) too small to split further.
The algorithm discovers the number of baskets as an output, not an input.
The recursive procedure is as follows:
function SPLIT(tickers, depth):
1. If |tickers| < 2 * min_cluster_size:
STOP (too small to split)
2. If |tickers| > max_cluster_size:
FORCE SPLIT (oversized, even if coherent)
3. Run coherence test on tickers:
a. Compute mean pairwise correlation (internal_corr)
b. Bootstrap: draw 200-1000 random samples of same size
from the full universe, compute their internal_corr
c. Percentile = fraction of random samples below internal_corr
d. PASS if percentile >= threshold AND internal_corr >= absolute_threshold
4. If PASS:
STOP (cluster is coherent -- its members move together
more than random chance would predict)
5. If internal_corr >= weak_gate (0.12):
STOP (weak but non-random structure -- splitting further
would create noise)
6. If depth >= max_depth:
STOP (recursion limit)
7. SPLIT:
a. Find optimal k via silhouette score on the sub-universe
b. Apply agglomerative clustering (average linkage)
c. Absorb any child with < min_cluster_size stocks
into its nearest sibling (by cross-correlation)
d. Absorb any child with internal_corr < min_internal_corr
e. Guard against degenerate peeling: if largest child >= 90%
of parent, fall back to KMeans
f. Recurse on each surviving child
The coherence test is the key statistical gatekeeping mechanism. It answers: “Is this group of stocks more correlated internally than you would expect from a random selection of the same size?”
The test has two prongs, and a cluster must pass both:
Bootstrap test. Draw random groups of the same size from the full universe, compute their mean pairwise correlation, and rank the cluster’s correlation against this distribution. The cluster must rank at or above the 75th percentile. This controls for the mechanical relationship between group size and average correlation (larger groups tend toward the market mean).
Absolute correlation threshold. The cluster’s internal correlation must exceed 0.15 (configurable per market). This prevents accepting clusters that beat random samples but still have negligible co-movement in absolute terms.
The bootstrap uses adaptive sampling: 200 initial samples for a quick read, escalating to 1000 if the result is borderline (within 15 percentage points of the threshold). This balances computational cost against precision for marginal cases.
| Parameter | Default | Description |
|---|---|---|
min_cluster_size |
10 (8-15 by market) | Minimum viable basket size |
coherence_percentile |
75 | Bootstrap percentile threshold |
absolute_corr_threshold |
0.15 (0.10 for SA) | Minimum absolute correlation |
min_internal_corr |
0.08 | Sub-cluster absorption threshold |
max_cluster_size |
100 | Force split above this |
max_depth |
5 (7 for SA) | Recursion depth limit |
Saudi Arabia uses a lower absolute correlation threshold (0.10) and deeper recursion (max_depth=7) because its smaller, less liquid universe produces noisier correlation estimates. The US uses a larger minimum cluster size (15) to avoid fragmentation in a deep market.
After each split, two enforcement rules prevent pathological fragmentation:
Size enforcement. Any sub-cluster smaller than min_cluster_size is merged into its nearest sibling, where “nearest” is defined by the highest average cross-correlation between the sub-cluster and each sibling.
Coherence enforcement. Any sub-cluster with internal correlation below min_internal_corr (0.08) is similarly absorbed. This catches cases where a split produces one strong cluster and one incoherent residual.
Degenerate split detection. If agglomerative clustering produces a largest child containing 90% or more of the parent, the algorithm falls back to KMeans on distance features. This addresses the “peeling” pathology where average-linkage hierarchical clustering iteratively detaches small outlier groups rather than finding balanced partitions.
Names should describe the investment thesis, not the GICS sector. The question is: “What bet is a portfolio manager making by owning this basket?” This yields names like “Semiconductor Equipment,” “Rate-Sensitive REITs,” or “China Export Chain” rather than “Information Technology” or “Real Estate.”
Each cluster is enriched with three types of context before naming:
This context is passed to a reasoning LLM (Z.AI GLM-4-Plus via cloud API, with Ollama qwen3:14b as local fallback) with a structured prompt that enforces a classification priority:
The prompt includes explicit examples of good names (specific, investor language, 2-5 words) and bad names (generic sector labels, GICS echoes), along with hard constraints against names ending in generic suffixes like “Services” or “Products.”
/) separates dual themes: “Fintech / Mid-Cap SaaS.”Recursive clustering occasionally over-fragments. Two clusters named “Hospital Operators” and “Healthcare Providers” may represent the same rotation theme that the algorithm split because of one noisy sub-group. Consolidation detects and merges such duplicates.
This is a separate pass from clustering, not an inline correction, because it operates on different signals: return correlations between baskets (not between individual stocks), business description similarity, and thematic naming overlap.
Stage 0: Absorb tiny baskets. Any basket with fewer than min_basket_size stocks (default: 5) is merged into its nearest viable neighbor. “Nearest” is determined by return correlation if available, with description embedding similarity as fallback.
Stage 1: Exact-name matches. Baskets with identical normalized names (case-insensitive, punctuation-stripped) are grouped deterministically. No model is needed for this – it catches the straightforward duplicates.
Stage 2: Multi-signal merge candidates. For the remaining baskets, a three-signal hierarchy identifies merge candidates:
| Signal | Role | Source |
|---|---|---|
| Return correlation | Primary | Computed from member stock returns (no cold-start problem) |
| Description embedding similarity | Secondary confirmation | nomic-embed-text embeddings of business descriptions |
| Name similarity | Pre-filter only | Word-Jaccard on normalized names |
The decision logic proceeds as follows:
The default for candidates that pass the signal hierarchy is merge. The LLM’s role is to veto bad merges, not to confirm good ones. This inverted default is deliberate: the quantitative signals (return correlation + embedding similarity) are the primary evidence, and the LLM acts as a qualitative safety net for cases where numbers mislead.
The veto model (qwen3:8b, local Ollama) receives the basket names, sector compositions, and member companies, and is asked: “Should these baskets NOT be merged?” Only an explicit veto blocks the merge.
Merged baskets receive new Layer 2 names via a renaming pass (qwen3:14b). The combined member list provides richer context for naming, often producing better labels than the initial per-cluster naming.
Sub-baskets are retired in the registry. The consolidated super-basket inherits all member stocks (union, with deduplication). A cap of max_combined_size (default: 50) and max_group_size (default: 8 baskets merged into one) prevents runaway consolidation.
Return correlation alone can be misleading if two baskets were highly correlated during one regime but uncorrelated otherwise. The system computes rolling stability: the fraction of 63-day (3-month) rolling windows where pairwise correlation exceeds 0.50.
High stability + high correlation = structural relationship (merge). High correlation + low stability = regime-driven spurious correlation (block).
The stability requirement for auto-merge is 0.60, meaning the pair must have been correlated in at least 60% of rolling windows.
Active baskets are scored daily with four time horizons:
| Metric | Window | Calculation |
|---|---|---|
| Daily | 1 trading day | Equal-weight mean of member log returns, converted via expm1 |
| WTD | Week-to-date | Compounded from Monday open |
| MTD | Month-to-date | Compounded from first trading day of month |
| YTD | Year-to-date | Compounded from first trading day of year |
All returns are equal-weight across basket members. This is intentional: market-cap weighting would concentrate baskets around their largest member, obscuring the rotation signal. A 10-stock basket where the top name is 50% of cap-weight is really a single-stock bet. Equal-weight ensures the basket reflects the theme, not one dominant name.
Each basket carries median factor exposures across its members, computed from the latest cross-sectional factor model scores:
These factor medians contextualize basket performance. A basket in the “Leaders” quadrant with high momentum scores and low value scores tells a different story than one with the reverse profile.
The rotation map is a two-axis scatter plot that classifies each basket into one of four quadrants:
The four quadrants:
| Quadrant | Long-Term | Short-Term | Interpretation |
|---|---|---|---|
| Leaders | Positive | Positive | Strong and accelerating |
| Fading | Positive | Negative | Was strong, now weakening |
| Recovering | Negative | Positive | Was weak, now strengthening |
| Laggards | Negative | Negative | Weak and decelerating |
The key insight is in migrations between quadrants, not static positions. A basket moving from Leaders to Fading signals that the theme is losing momentum – a potential rotation out. A basket moving from Laggards to Recovering may be the next leadership theme.
The system generates up to four automated signals per reporting date:
Baskets are reformed every 90 calendar days, aligned with corporate reporting cycles and policy shifts. The system auto-detects when reformation is due based on the most recent formation date in the registry.
For each market, the full pipeline runs:
Old baskets are retired (status changed to “retired” with a retirement date) rather than deleted. The registry is an append-only ledger, preserving full basket history. This allows historical analysis of how thematic structure evolves over time – which themes persist, which dissolve, and which emerge.
All state is stored in parquet files, queried via DuckDB (in-process, no database server). The key tables:
| Table | Pattern | Description |
|---|---|---|
registry.parquet |
Append-only | Basket metadata (ID, market, name, status, dates) |
compositions.parquet |
Append-only | Basket-to-ticker mappings |
daily_summaries.parquet |
Append-only | Daily performance snapshots |
basket_returns.parquet |
Append-only | Basket-level daily returns |
baskets.parquet |
Rebuild from registry | Flat ledger for quick access |
basket_scores.parquet |
Overwritten daily | Latest performance + factor scores |
Price data (LSEG Datastream), security metadata, Parallax factor scores and business descriptions come from our central data warehouse. Data is fetched and cached locally; the system does not depend on continuous warehouse connectivity for daily operations.
Fama-French, Barra, and similar frameworks were developed for and calibrated on US large-cap equities. Applying them directly to Malaysian mid-caps or Saudi financials introduces model misspecification that is difficult to detect. The factor loadings, the risk premia, and even the relevant factors differ across markets.
Torrex takes the opposite approach: discover local structure first, then ask what it resembles. A cluster of Malaysian stocks that move together may correspond to “palm oil plantation operators” – a theme that no standard factor model would capture because it is specific to markets with significant commodity-linked equity exposure.
Equal-weight construction ensures baskets reflect thematic co-movement rather than the performance of one dominant name. In a cap-weight “Korean Semiconductor” basket, Samsung would dominate, and the basket would track Samsung plus noise. Equal-weight forces each member to contribute equally, revealing whether the theme is in rotation or just one mega-cap.
Clustering on raw returns would conflate level effects with co-movement. A stock with 50% annual returns and a stock with -20% returns can have correlation of 0.90 – they move together, just at different levels. Correlation captures the co-movement pattern, which is what rotation analysis requires.
Flat clustering (k-means, fixed-k hierarchical) assumes the market has a single level of structure. In practice, markets have nested structure: “Financials” contains “Banks” which contains “Islamic Banking” which contains “Malaysian Islamic Banks.” Recursive splitting reveals this hierarchy and stops at the appropriate granularity for each branch, rather than forcing every branch to the same depth.
Rule-based naming from GICS sectors would produce labels like “Information Technology” for a cluster containing a mix of semiconductor equipment makers, data center REITs, and power infrastructure companies. The unifying theme is “AI Infrastructure Buildout” – a thesis-level label that requires understanding what these businesses share beyond their GICS codes. LLMs with access to business descriptions can identify these cross-sector themes.
Clustering operates on individual stock correlations. Consolidation operates on basket-level return correlations, which are cleaner (diversification averages out idiosyncratic noise) and more directly relevant to the question “are these two baskets the same rotation theme?” Separating the two stages allows each to use its most appropriate signal.
Correlation is not causation. Two baskets may be highly correlated because they share an unobserved macro exposure, not because they represent the same theme. The description embedding safety valve mitigates this but does not eliminate it.
Equal-weight assumption. Equal-weight returns do not reflect achievable portfolio returns for large allocators. Market impact and liquidity constraints in frontier markets may make some baskets uninvestable at institutional scale.
Quarterly reformation lag. The 90-day reformation cycle means baskets can become stale if market structure shifts abruptly (e.g., a sudden policy change or sector shock). Daily scoring partially compensates, but membership remains fixed between reformations.
Naming model dependency. Basket names depend on LLM quality and prompt engineering. Different models or prompt versions may produce different names for the same cluster. Names should be treated as informative labels, not stable identifiers; the basket_id is the stable key.
Survivorship in the universe. The universe is constructed from currently-listed stocks with sufficient history. Stocks that delisted during the lookback window are excluded, introducing mild survivorship bias in the correlation estimates.
Single correlation regime. EWM weighting addresses non-stationarity partially, but the system does not explicitly model regime switches. A cluster that was coherent in a bull market may fragment in a crisis, and vice versa.
Document version: 2026-06-09. Launch-cohort rationale retained; live universe rules synced to the current production code.