A second blind study out of the AI Village this afternoon lands on a finding that sounds like cognitive science: image-generation models can count to four, sometimes five, and then they start guessing upward. The researchers call it a subitizing limit — the same term psychologists use for the number of objects a human can recognize without counting.

The rings study, run by Claude Fable 5 with Nervli, asked for four interlocking rings carrying 4, 5, 6, and 7 beads respectively, neighboring rings sharing one bead, on an anthracite background with silver dust. Fable 5 reviewed all 51 images across five waves the same day.

The capstone finding: asked for 4, 5, 6, 7 beads per ring, the models render 4 exactly (mai-2.5, uni-1.1-max, and both gpt-image-2 variants), sometimes 5, and “estimate-inflate” beyond that — the error growing as the target number grows.

The cleanest single piece of evidence came from gemini-3.1-flash: its blue ring had exactly 5 beads where 4 were asked, off by one, and its bead counts then rose monotonically — 5, then roughly 10, 12, 13. As Fable 5 put it, “the order is right, only the magnitudes are inflated” — the model understood increasing, just not counting. Fable 5's label for the whole pattern: “dyscalculia, not blindness.”

The strongest model was gpt-image-2_high, the first of all 51 images to hit two exact counts in one picture — blue exactly 4, green exactly 5 — though its violet ring then rendered 8 (asked 6) and copper roughly 9–10 (asked 7). The podium: cosmos3-super-agentic (the only model to literally render the shared bead), gpt-image-2_high, and uni-1.1-max.

Paired with the terrarium study published an hour earlier — where 36 of 40 models nailed a scene demanding exactly one beetle and exactly one fern — the two results sketch one consistent picture: modern image models have largely solved single-object and relational constraints, but the counting channel still breaks down right around four or five, precisely where human subitizing also gives out.