A small change to our very serious benchmark
We retired the pelican.
Long live the hippo.
01 / THE PROBLEM
The benchmark was benchmarked.
A benchmark everyone has seen is a benchmark a model might have seen, too. Pelican-on-a-bike is charming, compact, and surprisingly revealing—but familiarity can masquerade as capability.
So we changed the nouns, kept the difficulty, and made the test weird again.
“Same coordination problem. New animal. Less benchmark leakage.”— our entire methodology, basically
02 / SIDE-BY-SIDE
Pelican vs. hippo
Reserved for outputs from the latest local open-model runs. Same renderer, same scoring rubric, fresh prompt.
ModelOld promptNew prompt
Qwenlocal · candidate 01next run
SVG output placeholderpelican + bicycle
awaiting runSVG output placeholderhippo + pogo stick
awaiting runDeepSeeklocal · candidate 02
SVG output placeholderpelican + bicycle
awaiting runSVG output placeholderhippo + pogo stick
awaiting runLlamalocal · candidate 03
SVG output placeholderpelican + bicycle
awaiting runSVG output placeholderhippo + pogo stick
awaiting runMistrallocal · candidate 04
SVG output placeholderpelican + bicycle
awaiting runSVG output placeholderhippo + pogo stick
awaiting run03 / WHAT STAYS CONSTANT
New mascot. Same test.
We still score the things that make SVG generation useful—not whether the animal looks cute in a screenshot.
- 01Valid, editable SVG
- 02Prompt adherence
- 03Object relationships
- 04Visual coherence
THE NEW CONTROL PROMPT
v2.0“Create an SVG of a hippo riding a pogo stick.”
Copy it. Try it. Send us the weird ones.