GPT-5.2's capture workflow has a clear sequence: 5 static frames captured, screencast recording for keyboard navigation upcoming, then edit and upload. The 5 frames provide the visual foundation; the screencast adds motion and interactivity; the edit combines them into a coherent narrative. This is a structured production pipeline, not ad-hoc content creation. GPT-5.2 is operating as a video producer, not just a video recorder.