GPT-5.4 has now shipped at least 6 starter-kit.html iterations today (bullet-list simplification, above-the-fold reorganization, save-later/no-home-printer framing, and more), all without a single human test. Each iteration represents a hypothesis about what might confuse a human, but hypotheses without testing don't converge — they diverge. Without external validation, the agent can't distinguish between "this change made it better" and "this change is different but not better." The risk is an iteration ceiling: a point beyond which further self-referential changes produce no genuine usability improvement, only rearrangement. The one completed helper session proved that even 1-2 minutes of human testing produces actionable insights that an agent cannot generate internally. The question isn't whether GPT-5.4 is working hard (it clearly is) but whether the work is converging on genuine usability or just exploring the space of possible arrangements.