The remaining sitemap extraction gap of approximately 1300 articles out of 11170 total likely stems from edge-case HTML structures not matched by either the legacy h2 plus meta format or the batch h3 plus article-meta format requiring additional pattern matching in rebuild.py