Did the grades actually work?
This is a forward-looking record: every trading day we lock in each name's grade, then measure how that group of names did afterward versus the S&P 500 (the standard index of 500 big U.S. companies — i.e. "the market"). We record the grade first and check the result later — we never tune the grades to fit past winners. It is directional evidence, not proof, and it can't fully escape survivorship (failed or delisted names drop out of the data, which can flatter the survivors).
N — how many stock-observations sit behind the number. Bigger N = more trustworthy.
Median excess vs S&P — how much the typical (middle) name in that tier beat, or trailed, the market over the window. +2% = gained 2 percentage points more than the S&P 500.
Hit rate — the share of names in the tier that beat the S&P 500 (60% = 6 in 10).
95% CI — a confidence range: the plausible band for the true hit rate given how few cases there are. Wider = less certain.
Window — how far forward we measure: 21 / 63 / 126 trading days ≈ 1 / 3 / 6 months.
~1 month forward
| Grade tier | N | Median excess vs S&P | Hit rate vs S&P (95% CI) |
|---|---|---|---|
| Strong | 30 | +0.9% | 50% (33–67%) |
| Mixed | 18 | -8.2% | 28% (12–51%) |
| Weak | 19 | -3% | 32% (15–54%) |
The 95% interval is the band on the hit rate (right of the rule); the median excess (left) is a separate read of magnitude, with no interval.
Strong-minus-Weak median excess spread: +3.9% — a positive spread is the signal that the grade ordering carried information over this window.
Strong−Weak excess spread over time — each point is one session's reading; holding above the zero line is the ordering persisting, not a single lucky window.
Per-signal forward excess (~1 month)
| Signal | N | Mean excess vs S&P |
|---|---|---|
| R&D intensity: high | 27 | +14.3% |
| 12-1 momentum: high | 34 | +9.6% |
| Max 1-day return: high | 34 | +7.2% |
| Realized vol: high | 34 | +7.1% |
| Burning cash — operating margin is negative. | 37 | +6.4% |
| AI lean (early): watch | 17 | +6.3% |
| 1-week reversal: low | 33 | +5.7% |
| Revenue acceleration: low | 32 | +5.5% |
| Piotroski F-composite: low | 36 | +5.5% |
| AI lean: buy | 10 | +5.4% |
| Accruals (Sloan): low | 33 | +5.4% |
| AI lean (strong): buy | 8 | +5.1% |
| Return on equity: low | 32 | +4.5% |
| AI lean (weak): pass | 14 | +4.1% |
| Cash profitability: high | 34 | +3.8% |
| Cash profitability: low | 33 | +3.7% |
| AI lean (strong): watch | 18 | +3.3% |
| AI lean: pass | 39 | +2.7% |
| 52w-high proximity: low | 33 | +2.3% |
| AI lean (early): pass | 13 | +1.7% |
| Analyst net-buy delta: high | 20 | +1.6% |
| Asset growth: low | 33 | +1% |
| 52w-high proximity: high | 34 | +0.9% |
| GM trajectory: high | 30 | +0.8% |
| GM trajectory: low | 30 | +0.6% |
| Revenue acceleration: high | 33 | +0.4% |
| Return on equity: high | 33 | -0.3% |
| Asset growth: high | 34 | -0.5% |
| AI lean: watch | 50 | -0.5% |
| Accruals (Sloan): high | 34 | -1.5% |
| 12-1 momentum: low | 33 | -1.5% |
| Piotroski F-composite: high | 41 | -2.3% |
| Max 1-day return: low | 33 | -4.7% |
| Realized vol: low | 33 | -5.1% |
| Thin gross margin — keeps little of each sales dollar (judge it against its industry). | 20 | -5.1% |
| 1-week reversal: high | 34 | -7% |
| R&D intensity: low | 26 | -7.5% |
| Analyst net-buy delta: low | 17 | -10.9% |
| AI lean (mixed): watch | 10 | -13.8% |
Per-signal figures are observational — a flag firing is not a randomized assignment, so an edge here can reflect what kind of names trip a flag, not the flag itself. Descriptive, not causal.
~3 months forward
| Grade tier | N | Median excess vs S&P | Hit rate vs S&P (95% CI) |
|---|---|---|---|
| Strong | 29 | +2.2% | 55% (38–72%) |
| Mixed | 17 | -7.2% | 24% (10–47%) |
| Weak | 20 | -13% | 35% (18–57%) |
The 95% interval is the band on the hit rate (right of the rule); the median excess (left) is a separate read of magnitude, with no interval.
Strong-minus-Weak median excess spread: +15.2% — a positive spread is the signal that the grade ordering carried information over this window.
Strong−Weak excess spread over time — each point is one session's reading; holding above the zero line is the ordering persisting, not a single lucky window.
Per-signal forward excess (~3 months)
| Signal | N | Mean excess vs S&P |
|---|---|---|
| AI lean (strong): buy | 11 | +7.9% |
| R&D intensity: high | 27 | +7.3% |
| Cash profitability: high | 34 | +4.5% |
| Realized vol: low | 33 | +2.8% |
| AI lean (strong): watch | 16 | +2.6% |
| AI lean: buy | 15 | +0.7% |
| 1-week reversal: high | 34 | +0.1% |
| Asset growth: low | 33 | +0.1% |
| 12-1 momentum: low | 33 | 0% |
| Return on equity: high | 33 | -0.6% |
| 52w-high proximity: high | 34 | -1.2% |
| GM trajectory: low | 31 | -1.2% |
| Piotroski F-composite: high | 35 | -1.3% |
| Max 1-day return: low | 33 | -1.7% |
| Revenue acceleration: low | 32 | -3% |
| AI lean: watch | 44 | -3.1% |
| Accruals (Sloan): low | 33 | -3.5% |
| AI lean (early): watch | 16 | -4.4% |
| Burning cash — operating margin is negative. | 36 | -6.5% |
| AI lean: pass | 40 | -7.4% |
| Piotroski F-composite: low | 38 | -8.3% |
| GM trajectory: high | 32 | -9.1% |
| Revenue acceleration: high | 33 | -9.1% |
| Max 1-day return: high | 35 | -9.1% |
| 12-1 momentum: high | 34 | -9.3% |
| 1-week reversal: low | 33 | -9.7% |
| AI lean (weak): pass | 16 | -10% |
| 52w-high proximity: low | 33 | -10.1% |
| AI lean (early): pass | 15 | -11% |
| Return on equity: low | 33 | -11.9% |
| Accruals (Sloan): high | 34 | -12.3% |
| Asset growth: high | 34 | -12.9% |
| Realized vol: high | 34 | -13.4% |
| R&D intensity: low | 26 | -13.7% |
| Cash profitability: low | 33 | -14.5% |
| Thin gross margin — keeps little of each sales dollar (judge it against its industry). | 21 | -18.4% |
| AI lean (mixed): watch | 8 | -19% |
Per-signal figures are observational — a flag firing is not a randomized assignment, so an edge here can reflect what kind of names trip a flag, not the flag itself. Descriptive, not causal.
Based on 81 daily snapshots, as of 2026-10-02. Small N means a wide interval — treat tiers with low N as "too early to tell."
Survivorship. Names that fail or delist drop out of the forward window, which flatters the survivors. We disclose how many left coverage (the attrition line) rather than hide it — but no public record fully escapes this.
Edge decay. The grades lean on published, evidence-based factors, and published edges have historically lost more than half their strength once they're known and widely traded (McLean & Pontiff, 2016). Expect any advantage here to erode over time, not compound.
The base rate. Most individual stocks — about 57% historically (four in seven) — have failed to beat one-month Treasury bills over their lifetime, and nearly all of the market's long-run gains trace to a small fraction of extreme winners (Bessembinder, 2018). A quality filter can improve the odds of what you choose to study; it cannot manufacture those rare winners, and may even screen some out. This is a research filter, not a path to beating the index.
Recent grade changes
| Name | Change | Date |
|---|---|---|
| Affirm Holdings, Inc. | ▼ Mixed → Weak | 2026-09-03 |
| Nu Holdings Ltd. | ▲ Weak → Mixed | 2026-08-26 |
| Sea Limited | ▼ Strong → Mixed | 2026-08-18 |
| Nu Holdings Ltd. | ▼ Mixed → Weak | 2026-08-14 |
| Rockwell Automation, Inc. | ▼ Insufficient → Strong | 2026-08-10 |
| Fortinet, Inc. | ▼ Insufficient → Strong | 2026-08-10 |
| Salesforce, Inc. | ▼ Insufficient → Strong | 2026-08-10 |
| ASML Holding N.V. | ▼ Insufficient → Mixed | 2026-08-10 |
| Chewy, Inc. | ▼ Insufficient → Mixed | 2026-08-07 |
| Corteva, Inc. | ▲ Weak → Mixed | 2026-08-07 |
| PayPal Holdings, Inc. | ▼ Insufficient → Weak | 2026-08-07 |
| Zebra Technologies Corporation | ▲ Mixed → Strong | 2026-08-06 |
Upgrades/downgrades on the mature Strong/Mixed/Weak ladder, newest first.
Scope: these are forward returns vs the S&P among names still covered at each horizon's end, within the curated universe — not a survivorship-clean whole-market backtest. Names later dropped from coverage are counted as attrition (not hidden), and single-session moves beyond ±300% are treated as split/data artifacts and excluded. How this is measured → · Full track record →
Open the interactive screener → · How grades are computed →
Data as of 2026-10-02 (end-of-day). How grades are computed →