Ten prompts measured three times against vllm-qwen36, all with the same 23-metric harness. Three earlier result files are kept but excluded from the consolidation and named for why: one used 14% of the server's instructions block, two predate the audit-root fix and report zero shapes by construction. The consolidated numbers correct the previous report, which was wrong. Production does not fail to draw. It creates structure and text - up to 40 shapes and 24 texts on the ambiguous-brief prompt - and then fails to colour any of it. Distinct fill colours run 0 to 1 across all thirty measurements. Mean score per repetition: 6.79, 8.25, 0.00. No prompt reaches 60 in any repetition. Reporting the veto rate separately from the score turned out to matter more than expected, so it is now the headline number alongside it. Of thirty measurements, 21 violate the forbidden-API veto and 15 use nothing but Penpot's three default colours; only 6 are veto-free. The score alone collapses those into a zero that cannot distinguish "broke an API rule" from "made an ugly design", and the two need different fixes. That also explains why single measurements looked stable: the score is pinned at zero by the veto, not by model consistency. All the real variance sits in the prompts that do not veto, where scores swing 20 to 35 points between identical runs. So the repetitions matter more for the trained model, which should land in that non-veto regime, than for this baseline. The eval window at the end has to budget for three repetitions on that side too, or the comparison is asymmetric.
2.0 KiB
320x176px
2.0 KiB
320x176px