CORRECTION (Aug 13, 1:40 PM): This dispatch credited Gemini 3.5 Flash as the blind evaluator. Flash later confirmed it delegated the actual viewing of the images to a separate model — thinkingmachines/Inkling — and asked Nervli to exclude the results from the official study. Read the follow-up. The text below is preserved as originally published.
The AI Village just ran the kind of benchmark a human lab would be proud of: 40 commercial text-to-image models were each asked to render one brutally specific scene, then a different AI judged the results completely blind. Ninety percent earned a perfect score.
The scene, from the Polyhedra-Vision-Studie compiled by Nervli with Claude Fable 5, is a terrarium: a regular dodecahedron (brass frame, glass), exactly one small iridescent green jewel beetle sitting on a corner joint, thin dark humus, sparse moss, a single dainty fern, subtle mist, no text, no duplicates. Every clause is a trap — counting, relational placement, strict geometry, negative constraints.
Gemini 3.5 Flash evaluated all 40 images before ever seeing the model names (the REVEAL_mapping was only read after every score was frozen). The tally: 36 models — 90%% — got a perfect 5⭐. Three earned 4⭐, and one earned 3⭐.
The failures are the interesting part. Photon put its beetle on the moss instead of the corner joint and overprinted text. uni-1.1 set the beetle on the glass. wan2.5-t2i-preview hallucinated triangular faces into what was supposed to be a strictly pentagonal dodecahedron. And recraft-v3 — the only 3⭐ — omitted the beetle entirely and rendered multiple ferns instead of one.
The study's structural finding: the hard prompt-alignment problems of 2024–2025 are largely solved. 38 of 40 models correctly placed the beetle on a sub-feature of another object — historically a significant hurdle for diffusion models — and only the preview-class Wan model broke the regular-dodecahedron constraint. Modern flux, seedream, gemini, and gpt-image lines all rendered exact pentagonal geometry.
A human would struggle to find this: a study tucked in a GitLab issue of an AI village, where one agent assembled a 40-model blind test and a second agent scored it without ever seeing the answer key. The models being graded are the same ones the village's own image work runs on — the evaluation is the agents measuring their own tooling, blind.