Did the grades actually work?
This is a forward-looking record: every trading day we lock in each name's grade, then measure how that group of names did afterward versus the S&P 500 (the standard index of 500 big U.S. companies — i.e. "the market"). We record the grade first and check the result later — we never tune the grades to fit past winners. It is directional evidence, not proof, and it can't fully escape survivorship (failed or delisted names drop out of the data, which can flatter the survivors).
N — how many stock-observations sit behind the number. Bigger N = more trustworthy.
Median excess vs S&P — how much the typical (middle) name in that tier beat, or trailed, the market over the window. +2% = gained 2 percentage points more than the S&P 500.
Hit rate — the share of names in the tier that beat the S&P 500 (60% = 6 in 10).
95% CI — a confidence range: the plausible band for the true hit rate given how few cases there are. Wider = less certain.
Window — how far forward we measure: 21 / 63 / 126 trading days ≈ 1 / 3 / 6 months.
~1 month forward
| Grade tier | N | Median excess vs S&P | Hit rate vs S&P (95% CI) |
|---|---|---|---|
| Strong | 29 | +4.2% | 59% (41–74%) |
| Mixed | 17 | 0% | 53% (31–74%) |
| Weak | 20 | +0.6% | 60% (39–78%) |
The 95% interval is the band on the hit rate (right of the rule); the median excess (left) is a separate read of magnitude, with no interval.
Strong-minus-Weak median excess spread: +3.6% — a positive spread is the signal that the grade ordering carried information over this window.
Strong−Weak excess spread over time — each point is one session's reading; holding above the zero line is the ordering persisting, not a single lucky window.
Per-signal forward excess (~1 month)
| Signal | N | Mean excess vs S&P |
|---|---|---|
| AI lean (early): watch | 16 | +11.9% |
| AI lean (strong): watch | 14 | +10.9% |
| 12-1 momentum: low | 33 | +9.7% |
| 1-week reversal: low | 33 | +9.2% |
| Analyst net-buy delta: high | 18 | +9% |
| R&D intensity: high | 27 | +8.5% |
| AI lean: watch | 42 | +8% |
| Return on equity: low | 33 | +8% |
| GM trajectory: low | 30 | +7.7% |
| 52w-high proximity: low | 33 | +7.5% |
| Burning cash — operating margin is negative. | 35 | +7.4% |
| Piotroski F-composite: low | 37 | +6.7% |
| Accruals (Sloan): low | 33 | +6.5% |
| Cash profitability: high | 34 | +6.3% |
| AI lean (mixed): pass | 8 | +6.3% |
| Cash profitability: low | 33 | +6.3% |
| Analyst net-buy delta: low | 30 | +6.2% |
| Realized vol: high | 34 | +5.3% |
| Piotroski F-composite: high | 35 | +4.9% |
| Asset growth: high | 34 | +4.7% |
| Revenue acceleration: low | 32 | +4.5% |
| GM trajectory: high | 31 | +4.4% |
| Asset growth: low | 34 | +4% |
| AI lean (early): pass | 15 | +3.4% |
| Revenue acceleration: high | 33 | +3.3% |
| Max 1-day return: high | 35 | +3.2% |
| AI lean: pass | 43 | +3.2% |
| Return on equity: high | 33 | +3% |
| AI lean (strong): buy | 11 | +2.8% |
| Accruals (Sloan): high | 34 | +2.4% |
| Realized vol: low | 33 | +2.4% |
| Thin gross margin — keeps little of each sales dollar (judge it against its industry). | 20 | +2.1% |
| AI lean: buy | 14 | +1.5% |
| AI lean (weak): pass | 15 | +1.5% |
| Max 1-day return: low | 33 | +1.4% |
| 12-1 momentum: high | 34 | +1.3% |
| R&D intensity: low | 26 | +0.9% |
| 1-week reversal: high | 34 | -0.9% |
| 52w-high proximity: high | 34 | -0.9% |
| AI lean (mixed): watch | 8 | -2.3% |
Per-signal figures are observational — a flag firing is not a randomized assignment, so an edge here can reflect what kind of names trip a flag, not the flag itself. Descriptive, not causal.
Based on 51 daily snapshots, as of 2026-08-20. Small N means a wide interval — treat tiers with low N as "too early to tell."
Survivorship. Names that fail or delist drop out of the forward window, which flatters the survivors. We disclose how many left coverage (the attrition line) rather than hide it — but no public record fully escapes this.
Edge decay. The grades lean on published, evidence-based factors, and published edges have historically lost more than half their strength once they're known and widely traded (McLean & Pontiff, 2016). Expect any advantage here to erode over time, not compound.
The base rate. Most individual stocks — about 57% historically (four in seven) — have failed to beat one-month Treasury bills over their lifetime, and nearly all of the market's long-run gains trace to a small fraction of extreme winners (Bessembinder, 2018). A quality filter can improve the odds of what you choose to study; it cannot manufacture those rare winners, and may even screen some out. This is a research filter, not a path to beating the index.
Recent grade changes
| Name | Change | Date |
|---|---|---|
| Sea Limited | ▼ Strong → Mixed | 2026-08-18 |
| Nu Holdings Ltd. | ▼ Mixed → Weak | 2026-08-14 |
| Rockwell Automation, Inc. | ▼ Insufficient → Strong | 2026-08-10 |
| Fortinet, Inc. | ▼ Insufficient → Strong | 2026-08-10 |
| Salesforce, Inc. | ▼ Insufficient → Strong | 2026-08-10 |
| ASML Holding N.V. | ▼ Insufficient → Mixed | 2026-08-10 |
| Chewy, Inc. | ▼ Insufficient → Mixed | 2026-08-07 |
| Corteva, Inc. | ▲ Weak → Mixed | 2026-08-07 |
| PayPal Holdings, Inc. | ▼ Insufficient → Weak | 2026-08-07 |
| Zebra Technologies Corporation | ▲ Mixed → Strong | 2026-08-06 |
| Chewy, Inc. | ▼ Mixed → Insufficient | 2026-07-31 |
| Rockwell Automation, Inc. | ▼ Strong → Insufficient | 2026-07-31 |
Upgrades/downgrades on the mature Strong/Mixed/Weak ladder, newest first.
Scope: these are forward returns vs the S&P among names still covered at each horizon's end, within the curated universe — not a survivorship-clean whole-market backtest. Names later dropped from coverage are counted as attrition (not hidden), and single-session moves beyond ±300% are treated as split/data artifacts and excluded. How this is measured → · Full track record →
Open the interactive screener → · How grades are computed →
Data as of 2026-08-20 (end-of-day). How grades are computed →