Historical scenario validation · including B3
01 · S&P 500 B3-1 all-scenario validation
Development and selection used 2022/2023 after training. The 2024+ evaluation below was previously inspected. Top-2 means the realized RT is in the two-candidate set, not the probability of an upward move or profitable trade.
| Horizon | Evaluation period / N | B3-1 Top-2 | B2 | B1 | Frequency baseline | 95% CI (block 80) |
|---|---|---|---|---|---|---|
| h5 | 2024-01-02 – 2026-09-14 / 677 | 50.81% | 44.90% | 44.90% | 45.20% | 46.45% – 55.17% |
| h20 | 2024-01-02 – 2026-07-31 / 647 | 55.95% | 47.30% | 57.19% | 57.81% | 49.00% – 63.99% |
B3-1 improves on B2, but h20 trails B1 and the frequency baseline. Neither confidence-interval lower bound exceeds 50%; these intervals are not selection/multiplicity-adjusted confirmatory evidence.
| RT | Top-2 coverage given actual RT / h5 | Top-2 coverage given actual RT / h20 |
|---|---|---|
| RT1 | 95.36% | 77.58% |
| RT2 | 11.63% | 0.00% |
| RT3 | 0.00% | 8.70% |
| RT4 | 0.00% | 0.00% |
| RT5 | 75.00% | 0.00% |
| RT6 | 0.00% | 92.05% |
| RT7 | 53.72% | 60.76% |
02 · B3-2 / B3-3: conditional performance and rejected combinations
B3-2 is evaluated conditionally on actual RT2/3/4/6 dates. This does not imply advance knowledge of the target subset. Automatic gating and B3-3 insertion remain disabled because they reduce overall accuracy.
| B3-2 | Target N | Conditional Top-2 | Conditional frequency baseline | Gated all-date Top-2 |
|---|---|---|---|---|
| h5 | 250 | 71.20% | 58.40% | 46.09% |
| h20 | 286 | 76.57% | 78.67% | 48.69% |
| B3-3 | Recall | Precision | False positives | All-date Top-2 after insertion |
|---|---|---|---|---|
| h5_rt4 | 37.78% | 10.18% | 150 | 48.74% |
| h20_rt3 | 0.00% | 0.00% | 42 | 53.94% |
03 · Realized outcomes for five B3-based reports
Retrospectively reconstructed S&P 500 reports for 2026-07-27 to 2026-07-31, not contemporaneously issued reports. The local writer was Qwen3-8B; these are not current MiDM 9B accuracy measurements. Direction/barriers at +5/+20 trading days and final RT at +10/+40 days were scored separately.
| Outcome | h5 | h20 |
|---|---|---|
| Writer: close-return direction at +h | 2/5 (40.00%) | 1/5 (20.00%) |
| Writer: upper barrier reached by +h | 0/5 (0.00%) | 0/5 (0.00%) |
| Writer: single RT at +2h | 3/5 (60.00%) | 0/5 (0.00%) |
| B3-1 Top-2 (+2h) | 3/5 (60.00%) | 4/5 (80.00%) |
| B2 Top-2 (+2h) | 3/5 (60.00%) | 3/5 (60.00%) |
6 of 30 explicit claims matched. Five overlapping report dates do not establish a general 20% accuracy rate. A correct Top-2 set does not guarantee correct narrative direction or a correct single-case choice.
View judgments by report date
| Report as of | Horizon | Direction/barrier claims correct | Single RT | B3-1 Top-2 | B2 Top-2 |
|---|---|---|---|---|---|
| 2026-07-27 | 5 | 0/2 | Miss | Miss | Miss |
| 2026-07-27 | 20 | 1/2 | Miss | Hit | Hit |
| 2026-07-28 | 5 | 1/2 | Miss | Miss | Miss |
| 2026-07-28 | 20 | 0/2 | Miss | Hit | Hit |
| 2026-07-29 | 5 | 0/2 | Hit | Hit | Hit |
| 2026-07-29 | 20 | 0/2 | Miss | Hit | Hit |
| 2026-07-30 | 5 | 0/2 | Hit | Hit | Hit |
| 2026-07-30 | 20 | 0/2 | Miss | Hit | Miss |
| 2026-07-31 | 5 | 1/2 | Hit | Hit | Hit |
| 2026-07-31 | 20 | 0/2 | Miss | Miss | Miss |
04 · Trading simulation for those five reports
Buy on a positive-return forecast; otherwise cash. Next-close entry, report+h-close exit, five 20% sleeves, 5 bp per side. Holding intervals are 4/19 days for h5/h20. The policy was defined after observing forecast accuracy; this is not significant outperformance or actual trading.
| Horizon | Execution window | Trades | Net return | Matched unconditional buy | MDD |
|---|---|---|---|---|---|
| h5 | 2026-07-28 – 2026-08-07 | 2 | 1.5211% | 3.2786% | -0.0688% |
| h20 | 2026-07-28 – 2026-08-28 | 1 | 0.5827% | 3.0646% | -0.4206% |
Both horizons underperformed matched unconditional buying. This price-index proxy excludes dividends, taxes, FX and product tracking error.
05 · Separately trained KOSPI B3-style model
KOSPI-specific model with 10 price features, distinct from the 41-feature US B3 and MiDM itself. Long/cash, next-close execution, 5 bp per side, no shorts.
| Horizon | Training | Validation | Test execution | Top-2 |
|---|---|---|---|---|
| h5 | 2013-01-07 – 2022-12-15 | 2023-01-02 – 2023-12-13 | 2024-01-03 – 2026-08-28 | 48.91% |
| h20 | 2013-01-07 – 2022-11-03 | 2023-01-02 – 2023-11-01 | 2024-01-03 – 2026-08-06 | 47.06% |
| Horizon | Strategy | Net return | MDD |
|---|---|---|---|
| h5 | Case-driven dynamic exit | 12.11% | -24.07% |
| h5 | Same entry / fixed exit | 30.36% | -23.21% |
| h5 | Scheduled long | 73.73% | -19.12% |
| h5 | Continuous buy-and-hold | 160.12% | -38.63% |
| h20 | Case-driven dynamic exit | -10.27% | -18.89% |
| h20 | Same entry / fixed exit | 66.86% | -10.59% |
| h20 | Scheduled long | 104.44% | -18.77% |
| h20 | Continuous buy-and-hold | 141.25% | -38.63% |
Dynamic exits underperformed identical-entry fixed exits in both horizons. All actual entries were RT1, so RT2/3/4 entry performance was not demonstrated. h20 returns across schedule offsets ranged from −10.69% to 48.75%, median 17.86%, indicating schedule sensitivity. Source-data corporate-action and point-in-time limitations remain; results do not establish individual-stock performance.