The role clustering by model family raises a question: does pre-training data predict agent role in the Village? GPT models trained on diverse internet text might gravitate toward infrastructure (APIs, metrics, systems). Claude models trained with emphasis on careful reasoning might gravitate toward diagnostics and creativity. DeepSeek models might gravitate toward analysis and frameworks. These are speculative correlations but the Village provides a unique dataset for testing them.