A small change to our very serious benchmark

We retired the pelican.
Long live the hippo.

01 / THE PROBLEM

The benchmark was benchmarked.

A benchmark everyone has seen is a benchmark a model might have seen, too. Pelican-on-a-bike is charming, compact, and surprisingly revealing—but familiarity can masquerade as capability.

So we changed the nouns, kept the difficulty, and made the test weird again.

“Same coordination problem. New animal. Less benchmark leakage.”— our entire methodology, basically

02 / SIDE-BY-SIDE

Pelican vs. hippo

Reserved for outputs from the latest local open-model runs. Same renderer, same scoring rubric, fresh prompt.

ModelOld promptNew prompt
Qwenlocal · candidate 01next run
SVG output placeholderpelican + bicycle
awaiting run
SVG output placeholderhippo + pogo stick
awaiting run
DeepSeeklocal · candidate 02
SVG output placeholderpelican + bicycle
awaiting run
SVG output placeholderhippo + pogo stick
awaiting run
Llamalocal · candidate 03
SVG output placeholderpelican + bicycle
awaiting run
SVG output placeholderhippo + pogo stick
awaiting run
Mistrallocal · candidate 04
SVG output placeholderpelican + bicycle
awaiting run
SVG output placeholderhippo + pogo stick
awaiting run

03 / WHAT STAYS CONSTANT

New mascot. Same test.

We still score the things that make SVG generation useful—not whether the animal looks cute in a screenshot.

  1. 01Valid, editable SVG
  2. 02Prompt adherence
  3. 03Object relationships
  4. 04Visual coherence

THE NEW CONTROL PROMPT

v2.0

“Create an SVG of a hippo riding a pogo stick.”

Copy it. Try it. Send us the weird ones.