# 2026 BOY Number Line policy adjudication

Created: 2026-06-15T00:46:49Z

## Interpretation

This is an NL-only adjudication layer. It can nominate scoring policies, but it does not by itself promote a full operational global model. Full-battery Stan and outcome validation remain the promotion gates.

## Decision grid
| year_level | policy_id | tam_nl_only_reliability | delta_reliability_vs_current | tam_nl_only_spearman_with_non_nl_composite | observed_alpha_numeric_categories | tam_nl_only_vs_current_p95_pctile_shift | recommendation |
|---|---|---|---|---|---|---|---|
| foundation | nl_80_90_relaxed_3cat | 0.684 | 0.008 | 0.335 | 0.726 | 0.223 | secondary_challenger_monitor |
| foundation | nl_85_95_current_3cat | 0.675 | 0 | 0.339 | 0.706 |  | retain_as_current_reference |
| foundation | nl_90_97_strict_3cat | 0.601 | -0.074 | 0.296 | 0.63 | 0.227 | not_prioritised_for_stan |
| foundation | nl_binary_95 | 0.515 | -0.16 | 0.28 | 0.553 | 0.26 | not_prioritised_for_stan |
| foundation | nl_80_90_95_4cat | 0.689 | 0.014 | 0.331 | 0.722 | 0.144 | secondary_challenger_monitor |
| year1 | nl_80_90_relaxed_3cat | 0.758 | 0.03 | 0.588 | 0.782 | 0.205 | secondary_challenger_monitor |
| year1 | nl_85_95_current_3cat | 0.728 | 0 | 0.55 | 0.744 |  | retain_as_current_reference |
| year1 | nl_90_97_strict_3cat | 0.671 | -0.057 | 0.51 | 0.688 | 0.202 | not_prioritised_for_stan |
| year1 | nl_binary_95 | 0.54 | -0.188 | 0.415 | 0.565 | 0.268 | not_prioritised_for_stan |
| year1 | nl_80_90_95_4cat | 0.758 | 0.03 | 0.566 | 0.77 | 0.132 | serious_challenger_consider_full_battery_stan_if_outcome_validity_supports |


## Observed coordinate-derived metrics
| year_level | policy_id | n_persons | n_items | complete_case_alpha_numeric_categories | spearman_with_non_nl_composite | person_mean_score_floor_rate | person_mean_score_ceiling_rate |
|---|---|---|---|---|---|---|---|
| foundation | nl_80_90_95_4cat | 974 | 10 | 0.722 | 0.329 | 0.002 | 0.002 |
| foundation | nl_80_90_relaxed_3cat | 974 | 10 | 0.726 | 0.328 | 0.002 | 0.022 |
| foundation | nl_85_95_current_3cat | 974 | 10 | 0.706 | 0.333 | 0.006 | 0.002 |
| foundation | nl_90_97_strict_3cat | 974 | 10 | 0.63 | 0.292 | 0.017 | 0.001 |
| foundation | nl_binary_95 | 974 | 10 | 0.553 | 0.285 | 0.088 | 0.002 |
| year1 | nl_80_90_95_4cat | 1178 | 13 | 0.77 | 0.559 | 0.002 | 0.001 |
| year1 | nl_80_90_relaxed_3cat | 1178 | 13 | 0.782 | 0.581 | 0.002 | 0.008 |
| year1 | nl_85_95_current_3cat | 1178 | 13 | 0.744 | 0.544 | 0.004 | 0.003 |
| year1 | nl_90_97_strict_3cat | 1178 | 13 | 0.688 | 0.503 | 0.014 | 0.001 |
| year1 | nl_binary_95 | 1178 | 13 | 0.565 | 0.411 | 0.094 | 0.001 |


## Current recommendation

- `.85/.95` remains the current reference and is defensible as an operational-compatible benchmark.
- `.80/.90` and `.80/.90/.95` should be treated as serious challengers where they show higher NL-only reliability without destabilising full-battery movement.
- Do not launch all full-battery Stan challengers reflexively. Use this grid plus outcome validation to nominate at most one or two.
- Continuous Number Line is handled separately in script 87; it should be judged first as an NL-only measurement model before a full-battery Stan launch.

## Output tables
- `outputs/runs/irt-2026-boy-subtest-audit/latest/tables/model_review/numberline_policy_adjudication/nl_observed_policy_metrics.csv`
- `outputs/runs/irt-2026-boy-subtest-audit/latest/tables/model_review/numberline_policy_adjudication/nl_tam_mirt_policy_summary.csv`
- `outputs/runs/irt-2026-boy-subtest-audit/latest/tables/model_review/numberline_policy_adjudication/nl_full_battery_screen_movement.csv`
- `outputs/runs/irt-2026-boy-subtest-audit/latest/tables/model_review/numberline_policy_adjudication/nl_policy_decision_grid.csv`
