# 2026 BOY continuous Number Line adjudication

Created: 2026-06-15T00:46:55Z

## What this does and does not decide

This is a non-IRT continuous-coordinate screen. It asks whether continuous accuracy/error contains enough extra stable signal to justify a formal model-based NL-only or full-battery Stan challenger. It does not by itself replace ordinal NL scoring.

## Continuous decision grid
| year_level | continuous_metric | alpha_delta_vs_current_ordinal | non_nl_spearman_delta_vs_current_ordinal | spearman_vs_current_ordinal | p95_percentile_shift_vs_current_ordinal | recommendation |
|---|---|---|---|---|---|---|
| foundation | continuous_accuracy | 0.101 | 0.035 | 0.931 | 21.82 | promising_continuous_challenger_consider_model_based_fit |
| foundation | negative_absolute_scaled_error | 0.101 | 0.035 | 0.931 | 21.838 | promising_continuous_challenger_consider_model_based_fit |
| foundation | absolute_signed_error_scaled_negative | 0.101 | 0.035 | 0.931 | 21.838 | promising_continuous_challenger_consider_model_based_fit |
| foundation | signed_error_scaled | -0.003 | -0.531 |  |  | diagnostic_bias_only_not_primary_score |
| year1 | continuous_accuracy | 0.093 | 0.061 | 0.931 | 21.817 | promising_continuous_challenger_consider_model_based_fit |
| year1 | negative_absolute_scaled_error | 0.093 | 0.061 | 0.931 | 21.781 | promising_continuous_challenger_consider_model_based_fit |
| year1 | absolute_signed_error_scaled_negative | 0.093 | 0.061 | 0.931 | 21.781 | promising_continuous_challenger_consider_model_based_fit |
| year1 | signed_error_scaled | 0.044 | -0.867 |  |  | diagnostic_bias_only_not_primary_score |


## Metric summary
| year_level | nl_metric | metric_family | complete_case_alpha | spearman_with_non_nl_composite | person_score_floor_rate | person_score_ceiling_rate |
|---|---|---|---|---|---|---|
| foundation | absolute_signed_error_scaled_negative | continuous | 0.807 | 0.369 | 0.001 | 0.001 |
| foundation | continuous_accuracy | continuous | 0.807 | 0.369 | 0.001 | 0.001 |
| foundation | negative_absolute_scaled_error | continuous | 0.807 | 0.369 | 0.001 | 0.001 |
| foundation | nl_80_90_95_4cat | ordinal_policy | 0.722 | 0.329 | 0.002 | 0.002 |
| foundation | nl_80_90_relaxed_3cat | ordinal_policy | 0.726 | 0.328 | 0.002 | 0.022 |
| foundation | nl_85_95_current_3cat | ordinal_policy | 0.706 | 0.333 | 0.006 | 0.002 |
| foundation | signed_error_scaled | continuous | 0.703 | -0.198 | 0.001 | 0.001 |
| year1 | absolute_signed_error_scaled_negative | continuous | 0.837 | 0.605 | 0.001 | 0.001 |
| year1 | continuous_accuracy | continuous | 0.837 | 0.605 | 0.001 | 0.001 |
| year1 | negative_absolute_scaled_error | continuous | 0.837 | 0.605 | 0.001 | 0.001 |
| year1 | nl_80_90_95_4cat | ordinal_policy | 0.77 | 0.559 | 0.002 | 0.001 |
| year1 | nl_80_90_relaxed_3cat | ordinal_policy | 0.782 | 0.581 | 0.002 | 0.008 |
| year1 | nl_85_95_current_3cat | ordinal_policy | 0.744 | 0.544 | 0.004 | 0.003 |
| year1 | signed_error_scaled | continuous | 0.788 | -0.323 | 0.001 | 0.001 |


## Recommendation

Continuous NL should be retained as a serious research/diagnostic contender if it improves reliability or non-NL composite alignment materially. If the gains are small and the continuous scores remain highly rank-correlated with the current ordinal score, prioritise ordinal `.80/.90` or 4-category challengers before an expensive continuous full-battery Stan run.

## Output tables
- `outputs/runs/irt-2026-boy-subtest-audit/latest/tables/model_review/numberline_continuous_adjudication/nl_continuous_metric_summary.csv`
- `outputs/runs/irt-2026-boy-subtest-audit/latest/tables/model_review/numberline_continuous_adjudication/nl_continuous_vs_ordinal_comparison.csv`
- `outputs/runs/irt-2026-boy-subtest-audit/latest/tables/model_review/numberline_continuous_adjudication/nl_continuous_target_residuals.csv`
- `outputs/runs/irt-2026-boy-subtest-audit/latest/tables/model_review/numberline_continuous_adjudication/nl_continuous_decision_grid.csv`
