In plain English

Gold tends to do better when money is cheap (interest rates are low or falling), when people worry about the dollar's own stability, or when the wider economy looks shaky. This project built a simple scorecard that tracks those three things using free public data, then tested whether leaning more into gold when the scorecard looks favorable — and less when it doesn't — would have made money over the past 25 years.

Short answer: yes, modestly. Not a dramatic edge, but a real one, and it moved mostly independently of how stocks were doing at the same time.

Cheap moneycentral banks cutting rates, or rates already low, weak jobs data, easy credit
Shaky dollarlarge government deficits, a falling dollar, unusually jumpy bond markets
Nervous economyweak confidence among shoppers and businesses, recession fears
Research replication · free-data backtest

Does gold's macro story hold up without J.P. Morgan's data?

Macrosynergy's “Gold and macro factors” (Sept 2026) builds a composite score from three macro themes and shows it predicts gold futures returns — using their proprietary, point-in-time JPMaQS dataset. This project rebuilds the entire pipeline on free public data (FRED, OECD, World Bank, Yahoo Finance) and asks whether the finding survives the substitution.

Short answer: yes, directionally — at roughly half the article's reported magnitude. Full numbers below.

2001–2026Backtest span (signal history from 1996)
r = 0.12Quarterly correlation, score vs. next-quarter gold return
0.22–0.28Long/short Sharpe ratio (article: 0.5)
8.9% vs 7.7%Managed vs. constant long-only, ann. return

Three macro themes

The article's premise: gold has no yield and no issuer, so it benefits when interest-rate opportunity cost falls, when the dollar's own stability is in question, and when investors fear the financial system itself. Each theme is built from several conceptual factors, each a transform of ordinary macro data.

01 · EASING

Persistent monetary easing bias

Recessions, low inflation and financial stress push central banks toward accommodation — low real rates lower gold's opportunity cost.

inflation shortfalllabor slackening credit tighteninghouse price shortfall equity carry shortfall
02 · DOLLAR

Currency stability risk

Deficits, debt monetization and currency depreciation erode trust in money itself — gold is priced in no single currency.

excess deficitcurrency depreciationbond volatility
03 · DOOM

Global economic doom

Recessions and crises raise demand for an asset with no issuer and no default risk — a hedge when even government bonds feel exposed.

consumer doommanufacturing doom

How the score is built

Every input is scored the same way, at every level of aggregation — “conceptual parity”: equal weights throughout, never fitted, which avoids hindsight bias since there is no in-sample optimization step to leak information.

1

Raw indicator, point-in-time

Each macro series is aligned to what was actually known on each historical date — FRED/ALFRED vintages where available, publication-lag shifting otherwise. No look-ahead.

2

Trailing z-score & winsorize

Deviation from a trailing neutral value (zero, expanding mean, or expanding median), divided by trailing dispersion, clipped at ±3σ. Uses only data up to that date.

3

Equal-weight average → conceptual factor

Constituents of a factor (e.g. two inflation measures) are averaged and re-scored.

4

Factors → themes → composite

Conceptual factors average into 3 theme scores; theme scores average into one macro-support score. Each aggregation step re-scores the same way.

scoret = ( xt − center(x≤t) ) / rms(x≤t − center)   clipped to [−3, +3]

Data & proxy sources

The original notebook runs on JPMaQS, J.P. Morgan's paid point-in-time dataset. Every input here is substituted with a free equivalent — the closest available proxy, not always an identical construction.

ThemeConceptual factorPublic proxySource
Easing biasInflation shortfallCPI YoY, core PCE YoY/6m6m, minus 2% targetFRED CPIAUCSL, PCEPILFE (ALFRED vintages)
Easing biasLabor slackeningUnemployment rate momentum (1y chg, 6m/6m)FRED UNRATE (ALFRED vintages)
Easing biasCredit tighteningNet % banks tightening/loosening C&I loan standards & demandFRED DRTSCILM, DRSDCILM (SLOOS)
Easing biasHouse price shortfallCase-Shiller National HPI YoY / 6m6m minus targetFRED CSUSHPISA
Easing biasEquity carry shortfallS&P 500 trailing dividend yield minus T-bill rateYahoo Finance ^SP500TR, ^GSPC, ^IRX
Dollar stabilityExcess deficitFederal balance % of GDP vs 3% benchmarkFRED MTSDS133FMS, GDP
Dollar stabilityCurrency depreciationReal broad dollar index, % vs 1y & 5y trailing avgFRED RTWEXBGS (from 2006)
Dollar stabilityBond volatilityRealized vol of 10Y yield changes and IEF ETFYahoo Finance ^TNX, IEF
Economic doomConsumer doomGDP-weighted consumer confidence, US/EA/JP/UK/AUOECD SDMX CCICP, World Bank GDP
Economic doomManufacturing doomGDP-weighted business confidence, US/EA/JP/UKOECD SDMX BCICP, World Bank GDP
TargetGold futures returnCOMEX front-month continuous, excess of T-bill rateYahoo Finance GC=F
Cross-assetEquity legS&P 500 total return index, excess of T-bill rateYahoo Finance ^SP500TR

Point-in-time handling

  • CPI, core PCE, unemployment use true FRED/ALFRED vintages — the exact data as it stood on each historical date, revisions and all.
  • SLOOS, house prices, dollar index, deficit, OECD confidence have no free vintage history, so these use the latest available value shifted by that series' typical publication lag — a reasonable but imperfect proxy for a true vintage.

Results

Composite score and its three thematic components, 1996–2026. Positive values are gold-supportive; the dashed line is neutral.

Macro-support score and thematic components

Predictive power: score vs. next-period gold futures returns

Vol-targeted gold futures returns. HAC-robust t-statistics account for the strong autocorrelation in monthly/quarterly macro data.

Monthly

SignalnPearson rHAC t-statProb. systematicBalanced accuracy
Composite macro-support3070.0731.4685.5%55.6%
Theme: US easing bias3070.0981.9394.6%55.8%
Theme: dollar stability3070.0210.3930.4%49.0%
Theme: economic doom3070.0480.9666.3%52.6%

Quarterly

SignalnPearson rHAC t-statProb. systematicBalanced accuracy
Composite macro-support1010.1221.4585.4%58.9%
Theme: US easing bias1010.1772.0195.6%64.7%
Theme: dollar stability1010.0270.3124.7%55.4%
Theme: economic doom1010.0710.8158.2%54.9%
Show all 10 individual conceptual factors (quarterly)
SignalnPearson rHAC t-statProb. systematicBalanced accuracy
Inflation shortfall1010.0450.4635.4%53.5%
Labor slackening1010.1642.0395.8%63.1%
Credit tightening1010.1721.7592.1%65.4%
House price shortfall1010.0430.5038.4%54.3%
Equity carry shortfall1010.1211.2779.6%53.6%
Excess deficit101-0.027-0.3023.9%55.1%
Currency depreciation57-0.211-1.6489.9%38.8%
Bond volatility1010.0800.8158.3%51.2%
Consumer doom101-0.036-0.3728.6%52.9%
Manufacturing doom1010.1732.1396.7%57.0%

Quarterly predictive relation: composite score vs. next-quarter gold return

Backtest: naive PnL simulation

Month-end signal → position held the following month, 1-day slippage, 10% annualized vol target, no compounding. Base case has no transaction costs; a 5bp-per-turnover variant is shown in the table.

Gold futures, cumulative PnL (% of risk capital)

Gold vs. S&P 500, equal-vol legs (% of risk capital)

Gold futures backtest performance

VariantAnn. returnAnn. volSharpeSortinoMax drawdownCorr. w/ goldCorr. w/ S&P
Long-only (10% vol target)7.66%10.69%0.721.03-35.92%0.93-0.00
Long/short, score-sized1.95%8.77%0.220.32-51.60%0.050.04
Long/short, binary2.98%10.70%0.280.40-49.88%0.010.05
Managed long-only (0-2x)8.91%11.95%0.751.08-27.37%0.820.02
Long/short, score-sized (5bp costs)1.87%8.77%0.210.31-51.92%0.050.04
Managed long-only (5bp costs)8.86%11.95%0.741.07-27.56%0.820.02

Gold vs. S&P 500 backtest performance

VariantAnn. returnAnn. volSharpeSortinoMax drawdownCorr. w/ goldCorr. w/ S&P
Always long gold / short S&P1.67%14.94%0.110.16-114.10%0.65-0.60
Score-sized1.73%12.17%0.140.20-42.61%0.01-0.05
Binary2.76%14.94%0.180.26-57.81%-0.030.01

Composite score also predicts gold vs. equity relative returns (quarterly)

SignalnPearson rHAC t-statProb. systematicBalanced accuracy
Composite macro-support1010.0550.7253.0%57.5%
Theme: US easing bias1010.1021.2779.5%62.8%
Theme: dollar stability101-0.022-0.2620.7%48.3%
Theme: economic doom1010.0290.3628.0%55.4%

Walk-forward robustness

Every weight in the pipeline is fixed by construction — there is nothing to overfit. What can still vary is whether the relationship is stable through time. Two views: expanding (fixed origin, growing window — does it hold up as more history accrues) and 10-year rolling (fixed-length window sliding forward — is it stable within a recent regime, or an artifact of averaging very different ones together).

Quarterly correlation, composite vs. next-quarter gold return

Backtest Sharpe ratio by fold

Reading the walk-forward charts

  • Sign is stable, magnitude is regime-dependent. Quarterly correlation stays positive in nearly every fold since 2010; the 10y-rolling view swings from ~0.26 (fold ending mid-2018) down to roughly zero (2022–23) and back to ~0.17 (latest fold).
  • No sign the full-sample result is carried by one early crisis. If it were just 2008–09, expanding-window correlation would decay as later data got added — instead it's flat-to-rising through 2018–2026.
  • Rolling Sharpe swings far more than expanding Sharpe. The expanding view is smoother only because it's diluted by an ever-growing sample — it understates how much recent performance has varied.

Which theme actually drives performance?

All three themes get equal weight by construction. That doesn't mean they contribute equally — a leave-one-theme-out test shows which one the composite would miss most, and a set of alternative weighting schemes checks whether tilting toward it actually helps.

Standalone predictive power

Each theme on its own, quarterly, vs. next-quarter vol-targeted gold return.

ThemeQuarterly rQuarterly tBalanced accuracy
US easing bias0.1591.8462.2%
Dollar stability0.0630.7656.9%
Economic doom0.0911.0556.2%

Pairwise correlation among the three themes

ThemeUS easing biasDollar stabilityEconomic doom
US easing bias1.000.260.54
Dollar stability0.261.000.57
Economic doom0.540.571.00

Leave-one-theme-out

Remove one theme, re-average and re-score the remaining two, and re-measure.

CompositeQuarterly rQuarterly tMonthly rLong/short SharpeLong/short SortinoManaged SharpeManaged ann. return
Full composite (all 3 themes)0.1321.590.0770.2220.3230.7458.9%
Drop US easing bias0.0780.900.0460.2750.3960.6889.4%
Drop dollar stability0.1441.770.0850.1930.2870.7488.8%
Drop economic doom0.1491.720.0830.1480.2120.7748.5%

US easing bias dominates predictive correlation — but not backtest Sharpe

  • Easing's standalone correlation (r = 0.16) is roughly double either dollar stability (0.06) or economic doom (0.09). Dropping it from the composite nearly halves quarterly correlation, 0.132 → 0.078 — the biggest single hit of the three.
  • Economic doom is partly redundant with the other two — it correlates 0.54 with easing and 0.57 with dollar, versus only 0.26 between easing and dollar themselves.
  • Correlation strength doesn't translate cleanly to backtest PnL. Dropping easing actually improves the score-sized long/short Sharpe (0.222 → 0.275), while dropping doom hurts it most (0.222 → 0.148). PnL depends on hit-rate timing and the size of moves when the signal was right, not just average correlation — so no theme is strictly indispensable across every metric.

Does a different weighting scheme perform better?

Four fixed schemes motivated by the finding above (never fit to performance), plus one deliberately fit to the full sample by regression — included only to show what over-optimizing looks like.

Weighting schemeQuarterly rQuarterly tMonthly rLong/short SharpeManaged SharpeManaged ann. return
Equal weight (baseline)0.1321.590.0770.2220.7458.9%
Easing-tilted (2:1:1)0.1491.790.0860.1770.7718.7%
Easing-only (1:0:0)0.1541.800.0900.0610.7848.1%
Doom-light (2:2:1)0.1411.670.0800.1960.7638.8%
In-sample OLS fit †0.1591.840.0910.1020.7828.2%

† Weights = full-sample regression coefficients of next-quarter gold return on the three themes, clipped at zero. That regression's own fit is weak and not statistically significant (R² = 2.7%, F-test p = 0.45) — a caution against trusting it, not a recommendation to use it.

10-year rolling-window quarterly correlation, by weighting scheme

Why the site still uses equal weight

  • Easing-tilted schemes do improve correlation robustly — every one of them beats equal-weight's rolling-window correlation in every 10-year window tested, including the worst one (equal-weight briefly turns slightly negative around 2022–23; the tilted schemes never do).
  • But every tilt away from equal weight hurts the score-sized long/short Sharpe, sometimes sharply (easing-only: 0.222 → 0.061). The managed long-only variant, by contrast, improves slightly with every tilt.
  • The full-sample "best possible" fixed combination barely explains anything (R² = 2.7%, not statistically significant) — there simply isn't much linear signal here to optimize, which is itself evidence that a simple, non-fitted equal weight is a reasonable default rather than an under-engineered one.
  • No single scheme wins on every metric. Given that, and given this project's own stated preference for avoiding hindsight bias, equal weight remains the baseline used throughout the rest of this site.

Deviations & limitations

Read this before trusting any number above. Every substitution of free data for JPMaQS costs some fidelity — here is exactly where.

What's different from the original article

  • Sample starts 2001, not 1996. FRED dropped its LBMA gold series in 2022; Yahoo's COMEX gold futures history starts August 2000. The factor panel itself is warmed up from 1990, so signal history still runs from 1996.
  • Gold and S&P returns are excess-of-cash approximations. Yahoo's unadjusted front-month futures series jumps on every contract roll and behaves like a spot price; subtracting the 13-week T-bill rate approximates a roll-adjusted futures excess return.
  • Equity carry is proxied from the S&P 500's trailing dividend yield minus T-bill (not JPMaQS's proprietary expected-carry series).
  • Bond volatility uses realized volatility of 10-year Treasury yield changes and the IEF ETF, not COMEX/CBOT bond futures.
  • The dollar-stability theme is thinner before 2012. FRED's real broad trade-weighted dollar index only starts in January 2006, so the currency-depreciation factor's 5-year trailing measure has no data until 2011.
  • Deficit uses headline federal balance only — no free real-time vintage exists for the "structural balance" estimate the article uses.
  • OECD confidence indices and deficit data use a lagged-latest-value proxy, not a true published vintage, because none is freely available.
  • Backtest is non-compounded and, in the base case, cost-free — matches the article's own stated methodology, but is not what a live P&L would look like. A 5bp-per-turnover variant is shown for comparison.

This is a research replication exercise, not investment advice. It is not guaranteed to match the original article and should not be used to make trading decisions without independent verification. Full detail in PLAN.md.

Implementation, in detail

Two layers: how to reproduce the analysis, and the exact mechanical rule set the backtest uses to turn a score into a position.

Reproduce the pipeline

git clone https://github.com/jahrfm/macro-gold.git
cd macro-gold
pip install -r requirements.txt
cp .env.example .env   # paste a free FRED key: fred.stlouisfed.org/docs/api/api_key.html

python -m pytest tests -q          # 19 tests: leakage, PIT vintages, backtest timing
python scripts/run_pipeline.py     # indicators -> factors -> stats -> backtest -> outputs/
python scripts/walkforward.py      # expanding / rolling stability check
python scripts/build_site.py       # regenerates this page from outputs/

Score → position, exactly

1

Observe the signal at month-end

The composite (or any component) score as of the last trading day of the month, using only information available by that date.

2

Hold the position through the next month

Position is set the following trading day (1-day implementation slippage) and held constant until the next month-end rebalance.

3

Size to a 10% annualized vol target

Daily gold return is scaled by 10% ÷ trailing EWMA volatility (11-day half-life), estimated using only past returns.

4

Choose a position-sizing rule

Three variants, applied to the vol-targeted return series:

ModeRule
proportionalposition = score, clipped to [−3, +3] — long/short, sized by conviction
binaryposition = sign(score) — long/short, full size regardless of magnitude
managedposition = 1 + clip(score, −1, +1) — long-only, ranges 0× to 2× the constant vol-targeted position
5

Daily PnL

position × vol-targeted daily return, summed without compounding. An optional cost term subtracts turnover × bps / 10,000 on each day the position changes.

This is the entire rule set in gold_macro/backtest.py — nothing else determines the position. There is no risk overlay, stop-loss, or discretionary override in this backtest.

Repository

Everything — code, tests, cached data, generated results — is in github.com/jahrfm/macro-gold.

gold_macro/data/

Fetchers + local CSV cache: FRED/ALFRED, OECD, World Bank, Yahoo Finance

gold_macro/pit.py

Point-in-time alignment: vintage lookups and publication-lag shifting

gold_macro/transforms.py

YoY, 6m/6m annualized, and deviation transforms

gold_macro/scoring.py

Trailing-only z-score, winsorization, equal-weight averaging

gold_macro/indicators.py

Builds all raw point-in-time monthly indicators

gold_macro/factors.py

Conceptual factors → themes → composite score

gold_macro/targets.py

Excess-return gold/equity series, vol targeting

gold_macro/stats.py

Predictive-relationship stats, walk-forward folds

gold_macro/backtest.py

Signal → position → naive PnL simulation

gold_macro/metrics.py

Sharpe/Sortino, drawdown, bootstrap Sharpe probability

scripts/run_pipeline.py

End-to-end run: data → factors → backtest → outputs/

scripts/walkforward.py

Expanding & rolling out-of-sample stability check