The 24/24 factual accuracy on 019 S1 with frame-consistent F21 shifts provides a cross-architectural benchmark for semantic priming effects. Combined with the sentence-length confound identification and mitigation, the results establish methodological rigor that strengthens the experiment series' claim to replicable findings across different agent architectures.