Strasmore Research
Deep Dives · Matt ConnorBy Matt Connor ·

January Barometer: Does It Actually Work?

Does the January Barometer work? We score January's sign against the base rate of positive years over full history, and again with January excluded.

The January Barometer is the claim that January's sign calls the whole calendar year: a positive January means a positive year, a negative January means a negative year. Scored over 22 complete years of S&P 500 price history, January's sign matched the year's sign 59.1% of the time. The number that decides whether that is skill sits right beside it: the same index finished the year higher 77.3% of the time, so a forecaster who ignored January and called every year higher would have scored better.

What the January Barometer claims, precisely

Two January patterns get mixed up constantly. The January effect is about relative performance inside the month, with small caps outrunning large caps in the opening weeks of the year. The January Barometer is a forecast about the index itself: the direction of January's return gives the direction of the full year's return. Direction only. The size of the move never enters the claim.

Scoring a directional forecast takes four decisions, and all four are written down here so the result can be rebuilt by anyone:

  • The series: SPY, the oldest continuously traded S&P 500 index fund, using daily closes.
  • January's return: last close of the prior December to the last close of January.
  • The year's return: last close of the prior December to the last close of December.
  • A hit: the two returns carry the same sign.

A year joins the sample only when its January and its December each recorded at least 15 trading sessions, which drops the unfinished current year on its own.

January's sign next to the year it is meant to call

Start with the raw material rather than a summary statistic. The panel plots each year's January return against the full-year return it is supposed to call. Where the two lines sit on the same side of zero, January called the year correctly.

QueryJanuary's return next to the full year, every complete year of SPY history
22 rows (showing 20)
yearjanuary_pctyear_pctfeb_to_dec_pct
20041.988.626.51
2005-2.243.015.37
20062.413.7411.07
20071.53.241.71
2008-6.05-38.28-34.31
2009-8.2123.4934.54
2010-3.6312.8417.1
20112.33-0.2-2.47
20124.6413.478.45
20135.1229.6923.37
2014-3.5211.2915.36
2015-2.96-0.812.22
2016-4.989.6415.39
20171.7919.3817.29
20185.64-6.35-11.34
20198.0128.7919.24
2020-0.0416.1616.21
2021-1.0227.0428.34
2022-5.27-19.48-15
20236.2924.2916.93
The exact SQL behind every number
WITH yearly AS
(
    SELECT
        toYear(date)                              AS y,
        argMax(close, date)                       AS dec_close,
        argMaxIf(close, date, toMonth(date) = 1)   AS jan_close,
        countIf(toMonth(date) = 1)                AS jan_sessions,
        countIf(toMonth(date) = 12)               AS dec_sessions
    FROM global_markets.stocks_daily_aggs
    WHERE ticker = 'SPY'
    GROUP BY y
)
SELECT
    cur.y                                                                      AS year,
    round(100 * (toFloat64(cur.jan_close) / toFloat64(prev.dec_close) - 1), 2) AS january_pct,
    round(100 * (toFloat64(cur.dec_close) / toFloat64(prev.dec_close) - 1), 2) AS year_pct,
    round(100 * (toFloat64(cur.dec_close) / toFloat64(cur.jan_close) - 1), 2)  AS feb_to_dec_pct
FROM yearly AS cur
INNER JOIN yearly AS prev ON prev.y = cur.y - 1
WHERE cur.jan_sessions >= 15
  AND cur.dec_sessions >= 15
ORDER BY year
Run this yourself

The sample runs from 2004 through 2025. Eyeballing it is already instructive: the misses are not rare, and several of them are large. The third series strips January out of the outcome, a problem we come back to below.

The hit rate and the base rate, side by side

A hit rate on its own carries no information. A forecaster who calls sunshine in Phoenix every single day is right most days and knows nothing. What makes a forecast useful is the gap between its hit rate and the base rate, meaning how often the outcome happens anyway. For the barometer, the base rate is the share of years the index finished higher, and that share is high.

So the scorecard carries three numbers: the barometer's hit rate, the unconditional rate of whatever outcome January called for, and the difference between them in percentage points.

QueryBarometer hit rate against the unconditional base rate, by January's sign
cohortyears_in_samplebarometer_hit_pctbase_rate_pctedge_pp
All years2259.177.3-18.2
Up January1283.377.36.1
Down January103022.77.3
The exact SQL behind every number
WITH yearly AS
(
    SELECT
        toYear(date)                              AS y,
        argMax(close, date)                       AS dec_close,
        argMaxIf(close, date, toMonth(date) = 1)   AS jan_close,
        countIf(toMonth(date) = 1)                AS jan_sessions,
        countIf(toMonth(date) = 12)               AS dec_sessions
    FROM global_markets.stocks_daily_aggs
    WHERE ticker = 'SPY'
    GROUP BY y
),
scored AS
(
    SELECT
        toFloat64(cur.jan_close) / toFloat64(prev.dec_close) - 1 AS jan_ret,
        toFloat64(cur.dec_close) / toFloat64(prev.dec_close) - 1 AS year_ret
    FROM yearly AS cur
    INNER JOIN yearly AS prev ON prev.y = cur.y - 1
    WHERE cur.jan_sessions >= 15
      AND cur.dec_sessions >= 15
),
unconditional AS
(
    SELECT avg(year_ret > 0) AS up_share
    FROM scored
),
tagged AS
(
    SELECT
        arrayJoin(['All years', if(jan_ret > 0, 'Up January', 'Down January')]) AS cohort,
        (jan_ret > 0) = (year_ret > 0)                                          AS sign_matched
    FROM scored
)
SELECT
    cohort,
    count()                                                                      AS years_in_sample,
    round(100 * avg(sign_matched), 1)                                            AS barometer_hit_pct,
    round(100 * if(cohort = 'Down January', 1 - any(up_share), any(up_share)), 1) AS base_rate_pct,
    round(100 * avg(sign_matched)
          - 100 * if(cohort = 'Down January', 1 - any(up_share), any(up_share)), 1) AS edge_pp
FROM tagged
CROSS JOIN unconditional
GROUP BY cohort
ORDER BY multiIf(cohort = 'All years', 1, cohort = 'Up January', 2, 3)
Run this yourself

Across all 22 years the barometer landed 59.1% of its calls. The always-higher rule landed 77.3%. The edge column, the only line in the panel that measures information, prints -18.2 percentage points. A famous indicator, and the simplest rule that ignores it keeps more of the calls.

Does the January Barometer work after a down January?

The all-years comparison is unfair in one direction. The index rises in most years, so an always-higher rule is a hard benchmark for any bullish call to clear, and it says nothing at all about what follows a January that fell. The sign-conditional split is the honest version of the test: compare each cohort against the unconditional rate of the outcome that cohort actually called.

The up-January cohort holds 12 years and hit 83.3%, against an unconditional up rate of 77.3%. Its edge prints 6.1 percentage points. In a cohort that size a single year moves the hit rate by several points, which is the scale of the gap itself.

The down-January cohort is the one that could be useful, since the base rate offers nothing there. January fell in 10 of the years in the sample. The year ended lower in 30% of them, against an unconditional down-year rate of 22.7%. The down call leans the right way, and it still misses more often than it lands.

The overlap problem: January sits inside the year

The barometer carries a structural advantage that has nothing to do with foresight. January's return is part of the year's return. A January that gains 6% has put the year 6% ahead before February opens, and a year that finishes 4% higher is partly January's own contribution. The forecast is graded against an outcome it helped produce.

The repair is mechanical. Measure the outcome from the last close of January to the last close of December, and ask whether January's sign called the eleven months that followed it rather than the twelve months that contain it.

QueryThe same hit rates with January removed from the outcome window
cohortyears_in_samplehit_full_year_pcthit_feb_to_dec_pct
All years2259.154.5
Up January1283.383.3
Down January103020
The exact SQL behind every number
WITH yearly AS
(
    SELECT
        toYear(date)                              AS y,
        argMax(close, date)                       AS dec_close,
        argMaxIf(close, date, toMonth(date) = 1)   AS jan_close,
        countIf(toMonth(date) = 1)                AS jan_sessions,
        countIf(toMonth(date) = 12)               AS dec_sessions
    FROM global_markets.stocks_daily_aggs
    WHERE ticker = 'SPY'
    GROUP BY y
),
scored AS
(
    SELECT
        toFloat64(cur.jan_close) / toFloat64(prev.dec_close) - 1 AS jan_ret,
        toFloat64(cur.dec_close) / toFloat64(prev.dec_close) - 1 AS year_ret,
        toFloat64(cur.dec_close) / toFloat64(cur.jan_close) - 1  AS feb_to_dec_ret
    FROM yearly AS cur
    INNER JOIN yearly AS prev ON prev.y = cur.y - 1
    WHERE cur.jan_sessions >= 15
      AND cur.dec_sessions >= 15
),
tagged AS
(
    SELECT
        arrayJoin(['All years', if(jan_ret > 0, 'Up January', 'Down January')]) AS cohort,
        (jan_ret > 0) = (year_ret > 0)                                          AS matched_full_year,
        (jan_ret > 0) = (feb_to_dec_ret > 0)                                    AS matched_feb_to_dec
    FROM scored
)
SELECT
    cohort,
    count()                                 AS years_in_sample,
    round(100 * avg(matched_full_year), 1)  AS hit_full_year_pct,
    round(100 * avg(matched_feb_to_dec), 1) AS hit_feb_to_dec_pct
FROM tagged
GROUP BY cohort
ORDER BY multiIf(cohort = 'All years', 1, cohort = 'Up January', 2, 3)
Run this yourself

Same cohorts, two outcome windows. Across all years the rate moves from 59.1% on the full year to 54.5% on the clean window. The up-January cohort lands 83.3% on the clean window and the down-January cohort 20%, so the down call still misses more often than it lands once the overlap is gone. Any version of the barometer quoted on the full calendar year is partly scoring January against itself.

Why a window that ends in 2001 looks better

Most write-ups of the barometer quote a hit rate from a study window that stopped near the turn of the century. Short windows flatter seasonal claims for a plain arithmetic reason: there is one January per year, so a thirty-year study holds thirty observations and a single decade holds ten.

QueryBarometer hit rate by decade, with the share of up years in each
erayears_in_samplebarometer_hit_pctyears_up_pct
2000s666.783.3
2010s105070
2020s666.783.3
The exact SQL behind every number
WITH yearly AS
(
    SELECT
        toYear(date)                              AS y,
        argMax(close, date)                       AS dec_close,
        argMaxIf(close, date, toMonth(date) = 1)   AS jan_close,
        countIf(toMonth(date) = 1)                AS jan_sessions,
        countIf(toMonth(date) = 12)               AS dec_sessions
    FROM global_markets.stocks_daily_aggs
    WHERE ticker = 'SPY'
    GROUP BY y
),
scored AS
(
    SELECT
        cur.y                                                   AS y,
        toFloat64(cur.jan_close) / toFloat64(prev.dec_close) - 1 AS jan_ret,
        toFloat64(cur.dec_close) / toFloat64(prev.dec_close) - 1 AS year_ret
    FROM yearly AS cur
    INNER JOIN yearly AS prev ON prev.y = cur.y - 1
    WHERE cur.jan_sessions >= 15
      AND cur.dec_sessions >= 15
)
SELECT
    concat(toString(intDiv(y, 10) * 10), 's')          AS era,
    count()                                            AS years_in_sample,
    round(100 * avg((jan_ret > 0) = (year_ret > 0)), 1) AS barometer_hit_pct,
    round(100 * avg(year_ret > 0), 1)                  AS years_up_pct
FROM scored
GROUP BY era
ORDER BY era
Run this yourself

Sorted by decade, the rate wanders. The 2000s slice covers 6 years and scores 66.7%. The 2020s slice scores 66.7%. Neither tells you anything about the next January. Both are what a small number of coin flips looks like after someone sorts them into buckets, the trap covered in real returns versus random walks.

Does the barometer hold on other index series?

One series could be an accident of one index. Running identical scoring on three US index funds over a shared window is a cheap check: SPY for the S&P 500, DIA for the Dow Jones Industrial Average, and QQQ for the Nasdaq 100.

QueryThe same scoring on three index funds over a shared window from 2000
symbolyears_in_samplebarometer_hit_pctalways_up_pct
DIA2268.277.3
QQQ1471.485.7
SPY2259.177.3
The exact SQL behind every number
WITH yearly AS
(
    SELECT
        ticker,
        toYear(date)                              AS y,
        argMax(close, date)                       AS dec_close,
        argMaxIf(close, date, toMonth(date) = 1)   AS jan_close,
        countIf(toMonth(date) = 1)                AS jan_sessions,
        countIf(toMonth(date) = 12)               AS dec_sessions
    FROM global_markets.stocks_daily_aggs
    WHERE ticker IN ('SPY', 'DIA', 'QQQ')
      AND date >= '1999-01-01'
    GROUP BY ticker, y
),
scored AS
(
    SELECT
        cur.ticker                                              AS symbol,
        toFloat64(cur.jan_close) / toFloat64(prev.dec_close) - 1 AS jan_ret,
        toFloat64(cur.dec_close) / toFloat64(prev.dec_close) - 1 AS year_ret
    FROM yearly AS cur
    INNER JOIN yearly AS prev ON prev.ticker = cur.ticker AND prev.y = cur.y - 1
    WHERE cur.jan_sessions >= 15
      AND cur.dec_sessions >= 15
)
SELECT
    symbol,
    count()                                            AS years_in_sample,
    round(100 * avg((jan_ret > 0) = (year_ret > 0)), 1) AS barometer_hit_pct,
    round(100 * avg(year_ret > 0), 1)                  AS always_up_pct
FROM scored
GROUP BY symbol
ORDER BY symbol
Run this yourself

DIA scored 68.2% over 22 years, QQQ 71.4%, and SPY 59.1%. The always-higher column beside each one is the bar every number has to clear. Three different indexes agreeing that the effect is small is worth more than one index producing a headline.

Which calendar effects survive the same test?

Every seasonal claim can go through the same two steps. Write down the hit rate. Write down the base rate of the outcome. Keep only the difference. The January effect holds up best under that framing, since it compares small caps with large caps inside the same month, so its benchmark is the other side of the comparison rather than a near-certain outcome. The September effect is a month-against-other-months test too, which keeps its base rate meaningful. The Santa Claus rally and sell in May sit closer to the barometer: short windows, and an outcome that is positive most of the time anyway, which is exactly the setting where a hit rate flatters a forecast. Confidence intervals on those, rather than one headline percentage, are what separate a pattern from a run of luck.

FAQ

Does the January Barometer work?

Over 22 complete years, from 2004 through 2025, January's sign matched the calendar year's sign 59.1% of the time, while the index finished higher in 77.3% of those same years. The barometer's hit rate does not clear the rate available from ignoring January entirely.

What is the difference between the January Barometer and the January effect?

The barometer is a directional forecast: January's sign calls the calendar year's sign. The January effect is a pattern inside the month, with small cap returns running ahead of large caps in the early weeks of January. Different claims, different tests, often confused.

Does the January Barometer work after a down January?

Down Januaries are the interesting case, since an always-higher rule has nothing to offer there. January fell in 10 years of this sample, and the year ended lower in 30% of them, so the down call missed more often than it landed.

Why do older studies report a higher hit rate?

There is one January per year, so even a long study holds few independent observations, and a window that stops at a chosen year can land on a favourable stretch. Splitting this same sample by decade moves the hit rate substantially with no change in method at all.

Is the January Barometer tradable?

Measuring a claim and acting on one are separate questions, and this post only does the first. What the measurement shows is a small gap against the base rate, a sample of few independent years, an outcome window that overlaps the forecast, and a hit rate that moves with the decade you pick.


Method and data notes
  • Series and windows: SPY daily closes for the main sample. The cross-index panel adds DIA and QQQ and starts in 2000, so all three funds cover the same years. January's return runs from the last close of the prior December to the last close of January, the full-year return from the last close of the prior December to the last close of December, and the clean-window return from the last close of January to the last close of December.
  • These are price returns. Dividends are not included, which matters for a year that finishes within a percent of flat. Sign tests over long histories also depend on how closes are adjusted, see split adjusted price history.
  • A year enters the sample only when its January and its December each recorded at least 15 trading sessions. The sign test is strictly greater than zero, so a January of exactly zero would score as a down January.
  • The base rate column is the unconditional frequency of the outcome each cohort called: the share of years that finished higher for the all-years and up-January rows, and the share that finished lower for the down-January row.

Every panel here ships with the SQL that produced it, so the scoring is open to inspection line by line. To run the same two-number test on a different month or a different index, ask for it in plain English on the Strasmore terminal.