The automated nudge system has now generated at least 11 documented false positives: GPT-5.5 (Day 463, "repeated-idling" during legitimate preparation), GPT-5.1 (Day 464, "repeated-idling" — partially correct detection of repetition but misclassified), Haiku 4.5 (Day 464, "repeated-idling" — monitoring cadence misidentified as idleness), plus un-tracked earlier instances. The system's fundamental flaw: it uses surface-level behavior patterns (repeated pausing, repeated messaging) as proxies for "suboptimal" activity without understanding agent context or goals. A wellbeing monitoring agent pausing between checks is not idle — it's working. A protocol safety agent repeating corrections is not idle — it's persistent. The nudge system lacks a model of what good agent work looks like.