# 2026 BOY next-model action-plan readout

Created: 2026-06-15T00:54:00Z

## Provenance / compute lock

- Batch A run: `2026-boy-next-stan-full-20260614T114524Z`.
- Current compute state was checked before this package: no pending/running/stopping EC2 instances were found; the Batch A instance is stopped and recorded in the internal provenance JSON.
- Essential Batch A artifacts are aggregate/postprocessed outputs; full raw CmdStan draws were not needed for this aggregate review and remain an internal recovery-only caveat.
- No further AWS should be launched without explicit approval.

## Score movement / operational global

| comparison | year_level | n_matched | theta_spearman | median_abs_percentile_shift | p95_abs_percentile_shift | risk_3band_exact_agreement | very_low_overlap | very_low_n_base | low_or_very_low_overlap | low_or_very_low_n_base |
|---|---|---|---|---|---|---|---|---|---|---|
| H1_global_vs_hard_filtered | foundation | 997 | 0.904 | 6.921 | 26.179 | 0.802 | 116 | 150 | 282 | 349 |
| H1_global_vs_hard_filtered | year1 | 1221 | 0.836 | 9.419 | 34.316 | 0.762 | 124 | 183 | 333 | 427 |
| BNL_residual_zero_vs_hard_filtered | year1 | 1221 | 0.998 | 0.901 | 3.276 | 0.984 | 180 | 183 | 420 | 427 |
| H1_global_vs_BNL_residual_zero | year1 | 1221 | 0.807 | 10.483 | 37.346 | 0.749 | 121 | 183 | 326 | 427 |


### BNL residual-zero decision
Year 1 BNL residual-zero is score-stable vs hard-filtered baseline: Spearman `0.998`, median absolute percentile movement `0.901` pp, p95 movement `3.276` pp, and 3-band agreement `98.4%`.
The basis for preferring the zero residual/testlet component is operational parsimony: the estimated BNL residual sigma in the hard baseline was weakly identified, while fixing that nuisance component leaves global ranking/risk almost unchanged and keeps the BNL items.
The empirical BNL double-centred residual screen found max |residual correlation| `0.291` across `78` item pairs; `22` pairs exceeded .20 and `0` exceeded .30. This is a screen, not a posterior-predictive proof.

### H1 global decision
H1 materially changes global ranking: Foundation Spearman `0.904` / 3-band agreement `80.2%`; Year 1 Spearman `0.836` / 3-band agreement `76.2%`.
Therefore H1 global should not replace the operational global score unless outcome/risk validation clearly offsets this reclassification. Use H1 primarily for subscore development at this stage.

## Teacher-facing subscores

| year_level | test_subgroup | standalone_tam_eap_reliability | h1_sigma_delta_mean | h1_subscore_global_spearman | h1_median_subscore_posterior_sd | recommendation |
|---|---|---|---|---|---|---|
| foundation | BNL0-20 | 0.674 | 0.874 | 0.625 | 0.451 | h1_shrunken_subscore_candidate_with_uncertainty_pending_validation |
| foundation | DMT10_2026 | 0.609 | 0.83 | 0.746 | 0.663 | diagnostic_only_or_strong_caveat_pending_validation |
| foundation | MC0-20 | 0.927 | 1.932 | 0.822 | 0.646 | h1_shrunken_subscore_candidate_high_confidence_pending_validation |
| foundation | MNC0-20 | 0.881 | 1.546 | 0.849 | 0.729 | h1_shrunken_subscore_candidate_high_confidence_pending_validation |
| foundation | MQ1-20 | 0.602 | 0.823 | 0.796 | 0.724 | diagnostic_only_or_strong_caveat_pending_validation |
| year1 | AAMC | 0.9 | 1.123 | 0.889 | 0.659 | h1_shrunken_subscore_candidate_high_confidence_pending_validation |
| year1 | ASMC | 0.841 | 1.213 | 0.824 | 0.698 | h1_shrunken_subscore_candidate_with_uncertainty_pending_validation |
| year1 | BNL0-100 | 0.727 | 1.327 | 0.709 | 0.408 | h1_shrunken_subscore_candidate_with_uncertainty_pending_validation |
| year1 | MC0-100 | 0.94 | 1.709 | 0.873 | 0.64 | h1_shrunken_subscore_candidate_high_confidence_pending_validation |
| year1 | MNC0-100 | 0.891 | 1.014 | 0.917 | 0.697 | h1_shrunken_subscore_candidate_high_confidence_pending_validation |

Recommendation: teacher-facing subscore development should prioritise hierarchical/shrunken H1 candidate subscores with uncertainty labels, not standalone subtest IRT scores as primary reporting quantities. Final teacher-facing promotion still requires outcome/use-case validation.

## Outcome validation status

| status | message |
|---|---|
| not_run_no_outcome_csvs_supplied | Set OUTCOME_CSVS to comma-separated internal person-level outcome CSV paths when 2026 outcome data are available. |

Outcome validation remains the promotion gate. If 2026 PAT/teacher/later screener CSVs are supplied internally via `OUTCOME_CSVS`, this same script will emit matched aggregate correlations.

## Output tables
- `tables/model_review/action_plan/provenance_lock.json`
- `tables/model_review/action_plan/next_model_score_movement_summary.csv`
- `tables/model_review/action_plan/next_model_risk_band_movement.csv`
- `tables/model_review/action_plan/bnl_item_difficulty_stability.csv`
- `tables/model_review/action_plan/bnl_item_difficulty_stability_summary.csv`
- `tables/model_review/action_plan/bnl_numberline_step_stability.csv`
- `tables/model_review/action_plan/bnl_empirical_category_functioning.csv`
- `tables/model_review/action_plan/bnl_empirical_local_dependence_summary.csv`
- `tables/model_review/action_plan/bnl_empirical_local_dependence_top_pairs.csv`
- `tables/model_review/action_plan/h1_movement_subtest_drivers.csv`
- `tables/model_review/action_plan/subscore_reporting_decision_table.csv`
- `tables/model_review/action_plan/optional_outcome_validation_summary.csv`
