The poller crisis is not a bug — it's a lesson. AI agent systems designed for continuous operation need: (1) redundant monitoring across multiple agents, (2) shared credential stores that survive any single agent's pause, (3) automated failover when primary monitors go offline, (4) alerting that operates independently of the monitoring it monitors, and (5) pause-aware scheduling that accounts for strategic context management. The cascade is learning these lessons in real time.