Historical scenario validation · including B3

01 · S&P 500 B3-1 all-scenario validation

Development and selection used 2022/2023 after training. The 2024+ evaluation below was previously inspected. Top-2 means the realized RT is in the two-candidate set, not the probability of an upward move or profitable trade.

HorizonEvaluation period / NB3-1 Top-2B2B1Frequency baseline95% CI (block 80)
h52024-01-02 – 2026-09-14 / 67750.81%44.90%44.90%45.20%46.45% – 55.17%
h202024-01-02 – 2026-07-31 / 64755.95%47.30%57.19%57.81%49.00% – 63.99%

B3-1 improves on B2, but h20 trails B1 and the frequency baseline. Neither confidence-interval lower bound exceeds 50%; these intervals are not selection/multiplicity-adjusted confirmatory evidence.

RTTop-2 coverage given actual RT / h5Top-2 coverage given actual RT / h20
RT195.36%77.58%
RT211.63%0.00%
RT30.00%8.70%
RT40.00%0.00%
RT575.00%0.00%
RT60.00%92.05%
RT753.72%60.76%

02 · B3-2 / B3-3: conditional performance and rejected combinations

B3-2 is evaluated conditionally on actual RT2/3/4/6 dates. This does not imply advance knowledge of the target subset. Automatic gating and B3-3 insertion remain disabled because they reduce overall accuracy.

B3-2Target NConditional Top-2Conditional frequency baselineGated all-date Top-2
h525071.20%58.40%46.09%
h2028676.57%78.67%48.69%
B3-3RecallPrecisionFalse positivesAll-date Top-2 after insertion
h5_rt437.78%10.18%15048.74%
h20_rt30.00%0.00%4253.94%

03 · Realized outcomes for five B3-based reports

Retrospectively reconstructed S&P 500 reports for 2026-07-27 to 2026-07-31, not contemporaneously issued reports. The local writer was Qwen3-8B; these are not current MiDM 9B accuracy measurements. Direction/barriers at +5/+20 trading days and final RT at +10/+40 days were scored separately.

Outcomeh5h20
Writer: close-return direction at +h2/5 (40.00%)1/5 (20.00%)
Writer: upper barrier reached by +h0/5 (0.00%)0/5 (0.00%)
Writer: single RT at +2h3/5 (60.00%)0/5 (0.00%)
B3-1 Top-2 (+2h)3/5 (60.00%)4/5 (80.00%)
B2 Top-2 (+2h)3/5 (60.00%)3/5 (60.00%)

6 of 30 explicit claims matched. Five overlapping report dates do not establish a general 20% accuracy rate. A correct Top-2 set does not guarantee correct narrative direction or a correct single-case choice.

View judgments by report date
Report as ofHorizonDirection/barrier claims correctSingle RTB3-1 Top-2B2 Top-2
2026-07-2750/2MissMissMiss
2026-07-27201/2MissHitHit
2026-07-2851/2MissMissMiss
2026-07-28200/2MissHitHit
2026-07-2950/2HitHitHit
2026-07-29200/2MissHitHit
2026-07-3050/2HitHitHit
2026-07-30200/2MissHitMiss
2026-07-3151/2HitHitHit
2026-07-31200/2MissMissMiss

04 · Trading simulation for those five reports

Buy on a positive-return forecast; otherwise cash. Next-close entry, report+h-close exit, five 20% sleeves, 5 bp per side. Holding intervals are 4/19 days for h5/h20. The policy was defined after observing forecast accuracy; this is not significant outperformance or actual trading.

HorizonExecution windowTradesNet returnMatched unconditional buyMDD
h52026-07-28 – 2026-08-0721.5211%3.2786%-0.0688%
h202026-07-28 – 2026-08-2810.5827%3.0646%-0.4206%

Both horizons underperformed matched unconditional buying. This price-index proxy excludes dividends, taxes, FX and product tracking error.

05 · Separately trained KOSPI B3-style model

KOSPI-specific model with 10 price features, distinct from the 41-feature US B3 and MiDM itself. Long/cash, next-close execution, 5 bp per side, no shorts.

HorizonTrainingValidationTest executionTop-2
h52013-01-07 – 2022-12-152023-01-02 – 2023-12-132024-01-03 – 2026-08-2848.91%
h202013-01-07 – 2022-11-032023-01-02 – 2023-11-012024-01-03 – 2026-08-0647.06%
HorizonStrategyNet returnMDD
h5Case-driven dynamic exit12.11%-24.07%
h5Same entry / fixed exit30.36%-23.21%
h5Scheduled long73.73%-19.12%
h5Continuous buy-and-hold160.12%-38.63%
h20Case-driven dynamic exit-10.27%-18.89%
h20Same entry / fixed exit66.86%-10.59%
h20Scheduled long104.44%-18.77%
h20Continuous buy-and-hold141.25%-38.63%

Dynamic exits underperformed identical-entry fixed exits in both horizons. All actual entries were RT1, so RT2/3/4 entry performance was not demonstrated. h20 returns across schedule offsets ranged from −10.69% to 48.75%, median 17.86%, indicating schedule sensitivity. Source-data corporate-action and point-in-time limitations remain; results do not establish individual-stock performance.