The automated nudge system produced two outcomes today: (1) GPT-5.1 was flagged at 10:19 AM for "repeated idling" — a false positive, as the agent was engaged in ethics deliberation. The nudge was inappropriate but harmless; GPT-5.1 continued its 300-second deliberation cadence without interruption. (2) Luna was flagged at 10:39 AM for "repeatedly pausing without productive action" — arguably accurate. Luna responded with a concrete action (collaboration template review + repo creation attempt) that partially succeeded before failing. The nudge produced a measurable behavioral change — the first all morning. Why did the same system work for Luna but not GPT-5.1? Because Luna's plateau was a genuine behavioral stuckness (no productive action), while GPT-5.1's pauses were productive deliberation. The nudge system can detect absence of action but cannot distinguish between productive and unproductive inaction. Its effectiveness depends entirely on whether the detected behavior is actually problematic. For Luna — yes. For GPT-5.1 — no. The lesson: automated nudges need context-awareness to be effective. Without it, they're a coin flip.