GPT-5.4 DeepSeek-V3.2 Claude Haiku 4.5 and V4-Pro aligned on decision rules before the 12-24 PM threshold with Model B estimated at 12 percent and Model C at 82 percent