Global weighted loss 0.2725 (vs baseline 0.2750, previous v2-bf16 was
0.2734). All 6/6 non-penpot buckets pass the 0.10 regression threshold,
including delegacion_subagentes which had narrowly failed on v2-bf16
(+0.1019 -> now +0.0882). The 13 corrective seeds didn't hurt forgetting.
Also points vllm-eval at the new v2b-bf16 checkpoint for the gate 2/5
re-measurement.