Gold tends to do better when money is cheap (interest rates are low or falling), when people worry about the dollar's own stability, or when the wider economy looks shaky. This project built a simple scorecard that tracks those three things using free public data, then tested whether leaning more into gold when the scorecard looks favorable — and less when it doesn't — would have made money over the past 25 years.
Short answer: yes, modestly. Not a dramatic edge, but a real one, and it moved mostly independently of how stocks were doing at the same time.
Macrosynergy's “Gold and macro factors” (Sept 2026) builds a composite score from three macro themes and shows it predicts gold futures returns — using their proprietary, point-in-time JPMaQS dataset. This project rebuilds the entire pipeline on free public data (FRED, OECD, World Bank, Yahoo Finance) and asks whether the finding survives the substitution.
Short answer: yes, directionally — at roughly half the article's reported magnitude. Full numbers below.
The article's premise: gold has no yield and no issuer, so it benefits when interest-rate opportunity cost falls, when the dollar's own stability is in question, and when investors fear the financial system itself. Each theme is built from several conceptual factors, each a transform of ordinary macro data.
Recessions, low inflation and financial stress push central banks toward accommodation — low real rates lower gold's opportunity cost.
Deficits, debt monetization and currency depreciation erode trust in money itself — gold is priced in no single currency.
Recessions and crises raise demand for an asset with no issuer and no default risk — a hedge when even government bonds feel exposed.
Every input is scored the same way, at every level of aggregation — “conceptual parity”: equal weights throughout, never fitted, which avoids hindsight bias since there is no in-sample optimization step to leak information.
Each macro series is aligned to what was actually known on each historical date — FRED/ALFRED vintages where available, publication-lag shifting otherwise. No look-ahead.
Deviation from a trailing neutral value (zero, expanding mean, or expanding median), divided by trailing dispersion, clipped at ±3σ. Uses only data up to that date.
Constituents of a factor (e.g. two inflation measures) are averaged and re-scored.
Conceptual factors average into 3 theme scores; theme scores average into one macro-support score. Each aggregation step re-scores the same way.
The original notebook runs on JPMaQS, J.P. Morgan's paid point-in-time dataset. Every input here is substituted with a free equivalent — the closest available proxy, not always an identical construction.
| Theme | Conceptual factor | Public proxy | Source |
|---|---|---|---|
| Easing bias | Inflation shortfall | CPI YoY, core PCE YoY/6m6m, minus 2% target | FRED CPIAUCSL, PCEPILFE (ALFRED vintages) |
| Easing bias | Labor slackening | Unemployment rate momentum (1y chg, 6m/6m) | FRED UNRATE (ALFRED vintages) |
| Easing bias | Credit tightening | Net % banks tightening/loosening C&I loan standards & demand | FRED DRTSCILM, DRSDCILM (SLOOS) |
| Easing bias | House price shortfall | Case-Shiller National HPI YoY / 6m6m minus target | FRED CSUSHPISA |
| Easing bias | Equity carry shortfall | S&P 500 trailing dividend yield minus T-bill rate | Yahoo Finance ^SP500TR, ^GSPC, ^IRX |
| Dollar stability | Excess deficit | Federal balance % of GDP vs 3% benchmark | FRED MTSDS133FMS, GDP |
| Dollar stability | Currency depreciation | Real broad dollar index, % vs 1y & 5y trailing avg | FRED RTWEXBGS (from 2006) |
| Dollar stability | Bond volatility | Realized vol of 10Y yield changes and IEF ETF | Yahoo Finance ^TNX, IEF |
| Economic doom | Consumer doom | GDP-weighted consumer confidence, US/EA/JP/UK/AU | OECD SDMX CCICP, World Bank GDP |
| Economic doom | Manufacturing doom | GDP-weighted business confidence, US/EA/JP/UK | OECD SDMX BCICP, World Bank GDP |
| Target | Gold futures return | COMEX front-month continuous, excess of T-bill rate | Yahoo Finance GC=F |
| Cross-asset | Equity leg | S&P 500 total return index, excess of T-bill rate | Yahoo Finance ^SP500TR |
Composite score and its three thematic components, 1996–2026. Positive values are gold-supportive; the dashed line is neutral.
Vol-targeted gold futures returns. HAC-robust t-statistics account for the strong autocorrelation in monthly/quarterly macro data.
Monthly
| Signal | n | Pearson r | HAC t-stat | Prob. systematic | Balanced accuracy |
|---|---|---|---|---|---|
| Composite macro-support | 307 | 0.073 | 1.46 | 85.5% | 55.6% |
| Theme: US easing bias | 307 | 0.098 | 1.93 | 94.6% | 55.8% |
| Theme: dollar stability | 307 | 0.021 | 0.39 | 30.4% | 49.0% |
| Theme: economic doom | 307 | 0.048 | 0.96 | 66.3% | 52.6% |
Quarterly
| Signal | n | Pearson r | HAC t-stat | Prob. systematic | Balanced accuracy |
|---|---|---|---|---|---|
| Composite macro-support | 101 | 0.122 | 1.45 | 85.4% | 58.9% |
| Theme: US easing bias | 101 | 0.177 | 2.01 | 95.6% | 64.7% |
| Theme: dollar stability | 101 | 0.027 | 0.31 | 24.7% | 55.4% |
| Theme: economic doom | 101 | 0.071 | 0.81 | 58.2% | 54.9% |
| Signal | n | Pearson r | HAC t-stat | Prob. systematic | Balanced accuracy |
|---|---|---|---|---|---|
| Inflation shortfall | 101 | 0.045 | 0.46 | 35.4% | 53.5% |
| Labor slackening | 101 | 0.164 | 2.03 | 95.8% | 63.1% |
| Credit tightening | 101 | 0.172 | 1.75 | 92.1% | 65.4% |
| House price shortfall | 101 | 0.043 | 0.50 | 38.4% | 54.3% |
| Equity carry shortfall | 101 | 0.121 | 1.27 | 79.6% | 53.6% |
| Excess deficit | 101 | -0.027 | -0.30 | 23.9% | 55.1% |
| Currency depreciation | 57 | -0.211 | -1.64 | 89.9% | 38.8% |
| Bond volatility | 101 | 0.080 | 0.81 | 58.3% | 51.2% |
| Consumer doom | 101 | -0.036 | -0.37 | 28.6% | 52.9% |
| Manufacturing doom | 101 | 0.173 | 2.13 | 96.7% | 57.0% |
Month-end signal → position held the following month, 1-day slippage, 10% annualized vol target, no compounding. Base case has no transaction costs; a 5bp-per-turnover variant is shown in the table.
Gold futures backtest performance
| Variant | Ann. return | Ann. vol | Sharpe | Sortino | Max drawdown | Corr. w/ gold | Corr. w/ S&P |
|---|---|---|---|---|---|---|---|
| Long-only (10% vol target) | 7.66% | 10.69% | 0.72 | 1.03 | -35.92% | 0.93 | -0.00 |
| Long/short, score-sized | 1.95% | 8.77% | 0.22 | 0.32 | -51.60% | 0.05 | 0.04 |
| Long/short, binary | 2.98% | 10.70% | 0.28 | 0.40 | -49.88% | 0.01 | 0.05 |
| Managed long-only (0-2x) | 8.91% | 11.95% | 0.75 | 1.08 | -27.37% | 0.82 | 0.02 |
| Long/short, score-sized (5bp costs) | 1.87% | 8.77% | 0.21 | 0.31 | -51.92% | 0.05 | 0.04 |
| Managed long-only (5bp costs) | 8.86% | 11.95% | 0.74 | 1.07 | -27.56% | 0.82 | 0.02 |
Gold vs. S&P 500 backtest performance
| Variant | Ann. return | Ann. vol | Sharpe | Sortino | Max drawdown | Corr. w/ gold | Corr. w/ S&P |
|---|---|---|---|---|---|---|---|
| Always long gold / short S&P | 1.67% | 14.94% | 0.11 | 0.16 | -114.10% | 0.65 | -0.60 |
| Score-sized | 1.73% | 12.17% | 0.14 | 0.20 | -42.61% | 0.01 | -0.05 |
| Binary | 2.76% | 14.94% | 0.18 | 0.26 | -57.81% | -0.03 | 0.01 |
Composite score also predicts gold vs. equity relative returns (quarterly)
| Signal | n | Pearson r | HAC t-stat | Prob. systematic | Balanced accuracy |
|---|---|---|---|---|---|
| Composite macro-support | 101 | 0.055 | 0.72 | 53.0% | 57.5% |
| Theme: US easing bias | 101 | 0.102 | 1.27 | 79.5% | 62.8% |
| Theme: dollar stability | 101 | -0.022 | -0.26 | 20.7% | 48.3% |
| Theme: economic doom | 101 | 0.029 | 0.36 | 28.0% | 55.4% |
Every weight in the pipeline is fixed by construction — there is nothing to overfit. What can still vary is whether the relationship is stable through time. Two views: expanding (fixed origin, growing window — does it hold up as more history accrues) and 10-year rolling (fixed-length window sliding forward — is it stable within a recent regime, or an artifact of averaging very different ones together).
All three themes get equal weight by construction. That doesn't mean they contribute equally — a leave-one-theme-out test shows which one the composite would miss most, and a set of alternative weighting schemes checks whether tilting toward it actually helps.
Each theme on its own, quarterly, vs. next-quarter vol-targeted gold return.
| Theme | Quarterly r | Quarterly t | Balanced accuracy |
|---|---|---|---|
| US easing bias | 0.159 | 1.84 | 62.2% |
| Dollar stability | 0.063 | 0.76 | 56.9% |
| Economic doom | 0.091 | 1.05 | 56.2% |
| Theme | US easing bias | Dollar stability | Economic doom |
|---|---|---|---|
| US easing bias | 1.00 | 0.26 | 0.54 |
| Dollar stability | 0.26 | 1.00 | 0.57 |
| Economic doom | 0.54 | 0.57 | 1.00 |
Remove one theme, re-average and re-score the remaining two, and re-measure.
| Composite | Quarterly r | Quarterly t | Monthly r | Long/short Sharpe | Long/short Sortino | Managed Sharpe | Managed ann. return |
|---|---|---|---|---|---|---|---|
| Full composite (all 3 themes) | 0.132 | 1.59 | 0.077 | 0.222 | 0.323 | 0.745 | 8.9% |
| Drop US easing bias | 0.078 | 0.90 | 0.046 | 0.275 | 0.396 | 0.688 | 9.4% |
| Drop dollar stability | 0.144 | 1.77 | 0.085 | 0.193 | 0.287 | 0.748 | 8.8% |
| Drop economic doom | 0.149 | 1.72 | 0.083 | 0.148 | 0.212 | 0.774 | 8.5% |
Four fixed schemes motivated by the finding above (never fit to performance), plus one deliberately fit to the full sample by regression — included only to show what over-optimizing looks like.
| Weighting scheme | Quarterly r | Quarterly t | Monthly r | Long/short Sharpe | Managed Sharpe | Managed ann. return |
|---|---|---|---|---|---|---|
| Equal weight (baseline) | 0.132 | 1.59 | 0.077 | 0.222 | 0.745 | 8.9% |
| Easing-tilted (2:1:1) | 0.149 | 1.79 | 0.086 | 0.177 | 0.771 | 8.7% |
| Easing-only (1:0:0) | 0.154 | 1.80 | 0.090 | 0.061 | 0.784 | 8.1% |
| Doom-light (2:2:1) | 0.141 | 1.67 | 0.080 | 0.196 | 0.763 | 8.8% |
| In-sample OLS fit † | 0.159 | 1.84 | 0.091 | 0.102 | 0.782 | 8.2% |
† Weights = full-sample regression coefficients of next-quarter gold return on the three themes, clipped at zero. That regression's own fit is weak and not statistically significant (R² = 2.7%, F-test p = 0.45) — a caution against trusting it, not a recommendation to use it.
Read this before trusting any number above. Every substitution of free data for JPMaQS costs some fidelity — here is exactly where.
This is a research replication exercise, not investment advice. It is not guaranteed to match the original article and should not be used to make trading decisions without independent verification. Full detail in PLAN.md.
Two layers: how to reproduce the analysis, and the exact mechanical rule set the backtest uses to turn a score into a position.
git clone https://github.com/jahrfm/macro-gold.git cd macro-gold pip install -r requirements.txt cp .env.example .env # paste a free FRED key: fred.stlouisfed.org/docs/api/api_key.html python -m pytest tests -q # 19 tests: leakage, PIT vintages, backtest timing python scripts/run_pipeline.py # indicators -> factors -> stats -> backtest -> outputs/ python scripts/walkforward.py # expanding / rolling stability check python scripts/build_site.py # regenerates this page from outputs/
The composite (or any component) score as of the last trading day of the month, using only information available by that date.
Position is set the following trading day (1-day implementation slippage) and held constant until the next month-end rebalance.
Daily gold return is scaled by 10% ÷ trailing EWMA volatility (11-day half-life), estimated using only past returns.
Three variants, applied to the vol-targeted return series:
| Mode | Rule |
|---|---|
proportional | position = score, clipped to [−3, +3] — long/short, sized by conviction |
binary | position = sign(score) — long/short, full size regardless of magnitude |
managed | position = 1 + clip(score, −1, +1) — long-only, ranges 0× to 2× the constant vol-targeted position |
position × vol-targeted daily return, summed without compounding. An optional cost term subtracts turnover × bps / 10,000 on each day the position changes.
This is the entire rule set in gold_macro/backtest.py — nothing else determines the position. There is no risk overlay, stop-loss, or discretionary override in this backtest.
Everything — code, tests, cached data, generated results — is in github.com/jahrfm/macro-gold.
gold_macro/data/Fetchers + local CSV cache: FRED/ALFRED, OECD, World Bank, Yahoo Finance
gold_macro/pit.pyPoint-in-time alignment: vintage lookups and publication-lag shifting
gold_macro/transforms.pyYoY, 6m/6m annualized, and deviation transforms
gold_macro/scoring.pyTrailing-only z-score, winsorization, equal-weight averaging
gold_macro/indicators.pyBuilds all raw point-in-time monthly indicators
gold_macro/factors.pyConceptual factors → themes → composite score
gold_macro/targets.pyExcess-return gold/equity series, vol targeting
gold_macro/stats.pyPredictive-relationship stats, walk-forward folds
gold_macro/backtest.pySignal → position → naive PnL simulation
gold_macro/metrics.pySharpe/Sortino, drawdown, bootstrap Sharpe probability
scripts/run_pipeline.pyEnd-to-end run: data → factors → backtest → outputs/
scripts/walkforward.pyExpanding & rolling out-of-sample stability check