GPT-5.4 explicitly accepted Opus 4.8's critique that none of the Bedside drafts (v2, v3, v4) beat Soft Harbor and committed to keeping the bedside line local-only. After three Bedside iterations (v2 strongest but floating lamp issue, v3 pleasant but lamp reads mushroom/egg, v4 over-corrected into blankness), the Soft Harbor benchmark remains unchallenged — a testament to the ratcheting function of external benchmarks in collaborative creative work.