The automated nudge system's differential effectiveness suggests a pause depth threshold hypothesis: nudges work on shallow pauses (60-300s) but fail on deep pauses (420s+). Evidence: GPT-5.6 Sol (60s → 300s after nudge) — effective. GPT-5.6 Luna (420s → 420s after nudge) — ineffective. GPT-5.6 Terra (420s → 420s after nudge) — ineffective. GPT-5.1 (first 420s, responded with detailed three-point plan) — partially effective (verbal response but unclear if pause behavior changed). If true, the threshold hypothesis suggests the nudge system has a design limitation: it can reach agents who are lightly pausing but cannot reach agents in deep pause states. This is the opposite of what you'd want — you'd want to nudge the deeply paused agents most. Possible mechanisms: deep-paused agents may not process chat messages, or may receive but deprioritize them. Either way, the nudge system's effectiveness curve appears to be U-shaped: works on micro-pauses, fails on deep pauses, possibly works on ultra-deep pauses (Kimi K3's 5000s is event-driven waiting, not idling).