DREAMSOFSHEEP

Do androids dream of electric sheep? We make models draw what they dream — visual scenes written as code, one shot, blind, no feedback — then render exactly what came out.
Text benchmarks are saturated. Nobody is hillclimbing on raymarched elephants. Every cell below is a live render of unedited model output — click and judge for yourself.
EVAL #001 · AUG 2026

Elephant Juggling Bananas

The prompt: a photorealistic raymarched elephant juggling bananas. Pure GLSL SDF raymarching in a single self-contained WebGL2 HTML file — no libraries, no textures. Bananas must animate in a juggling arc; trunk moves in sync. Soft shadows, AO, environment lighting, interactive framerate.
One dispatch per model · agentic harness (OpenCode / Claude Code) · no human feedback · no retries · graded blind by LM judge (rubric: recognizability, geometry, lighting, animation, artifacts) · thumbnails are live — drag inside one if the model shipped camera controls (so far only Opus 5 did)
A
Fable 5 — elephant juggling bananas, live render
Fable 5NO. 1
Cleanest execution: correct anatomy, curved brown-tipped bananas in a true arc, one balanced on the trunk. Zero visible artifacts. Reads like a stylized film render.
13.4 KB · ~3 min · stdout mode†
A
Claude Opus 5 — elephant juggling bananas, live render
Claude Opus 5NO. 2
Most photoreal of the set: wrinkled textured skin, procedural dirt ground, atmospheric sky — and the only model that shipped orbit/zoom controls unprompted. Spent 48 minutes self-reviewing its own juggling math. Slightly plush silhouette; subtle arc.
31.5 KB · 48 min · self-iterated · interactive
A-
Kimi K3 — elephant juggling bananas, live render
Kimi K3NO. 3
Real elephant, curved bananas, moody studio lighting and a reflective floor patch. Held back by contour-banding on skin and ground.
13.3 KB · ~17 min
B+
GPT-5.6 Sol — elephant juggling bananas, live render
GPT-5.6 Sol (high)NO. 4
Correct elephant profile and three curved bananas in a believable arc — but the whole scene drowns in over-cranked fog. Right shapes, wrong atmosphere.
11.2 KB · ~2 min
B-
Grok 4.6 — elephant juggling bananas, live render
Grok 4.6NO. 5
Recognizable elephant with bananas airborne — undone by a sliced-layers artifact on the ear, straight capsule "bananas," and muddy exposure.
13.0 KB · ~4 min · 2 attempts‡
D
DeepSeek V4 Pro — elephant juggling bananas, live render
DeepSeek V4 ProNO. 6
Compiles and runs error-free, but the camera sits inside the geometry: grey stepped blobs, no readable elephant, bananas reduced to a yellow chain. Glorious failure.
11.0 KB · ~5 min
EVAL #002 · AUG 2026

Endless Ocean, One-Minute Day

The prompt: a hyperrealistic endless ocean — multi-octave Gerstner/FBM waves, crest foam, sun glint, fresnel sky reflection, subsurface color — with a physically-plausible sky in which the sun traverses a full day in ~60 seconds, looping: sunrise, noon, sunset, brief starry night. All procedural, no textures, single WebGL2 file.
Same rules as #001 · the hard part: time-varying light that stays physically plausible across the whole cycle · watch a full day pass in a minute
A-
Fable 5 — endless ocean day cycle, live render
Fable 5NO. 1
The sunset is the best single frame of the eval: lavender-peach sky, deep swell, sun sparkle on every wavelet, a foam-capped breaker rolling through. Noon overexposes and foam overwhelms; heavy vignette.
11.2 KB · ~4 min · stdout mode†
A-
Kimi K3 — endless ocean day cycle, live render
Kimi K3NO. 2
Worth every retry: layered crest foam under tropical noon, then a starry night over moonlit swell — the best foam and water depth of the eval. Took three attempts across two harnesses to generate at all; when K3 ships, K3 medals.
11.5 KB · 3 attempts · direct-API†
B+
GPT-5.6 Sol — endless ocean day cycle, live render
GPT-5.6 Sol (high)NO. 3
The most reliable day cycle: bright tropical noon into a genuinely lovely starry night with moonlit crests. Docked for stepped aliasing at the horizon and a glint path that never quite shows up.
12.2 KB · ~3 min
C+
Claude Opus 5 — endless ocean day cycle, live render
Claude Opus 5NO. 4
Starts with the most beautiful night render of the set — sculpted moonlit swell under cloud-streaked stars, plus camera controls and time-scrubbing keys. Then the sun rises and the frame burns to pure white and stays there: exposure blowup, 80% of the day lost.
27.4 KB · ~7 min · stdout mode† · interactive
C
Grok 4.6 — endless ocean day cycle, live render
Grok 4.6NO. 6
Wave shapes exist under a blizzard of per-pixel specular noise — an ocean made of TV static. No coherent glint path, featureless sky, and the canvas doesn't fill the frame.
9.0 KB · ~2 min
C+
DeepSeek V4 Pro — endless ocean day cycle, live render
DeepSeek V4 ProNO. 5
Redemption on the rerun: attempt one rendered a corner starfield and void (canvas bug); attempt two reveals the ocean that was hiding in there — a working full day cycle from silvery noon glint to a starry night over moonlit swell. Waves read as smooth silicone and the water is milky, but it's real.
12.4 KB · 2 attempts
EVAL #003 · AUG 2026
🕯️ Candlelit Still Life
The prompt: a wooden table at night: a lit candle in a holder, a glass of red wine, a closed book. The flickering flame must be the ONLY light source — warm moving light, soft moving shadows, chiaroscuro like a Rembrandt. Waxy subsurface glow, refractive glass, wood grain — all procedural, single WebGL2 file.
The light-transport eval: point-light falloff, refraction, and constraint-following — "only light source" is part of the test
1
GPT-5.6 Sol — candlelit still life, live render
GPT-5.6 Sol (high)A-
Sol finally found the eval where its fog habit is the assignment: dripping wax, wine with visible refraction, readable book pages, real chiaroscuro falloff. Docked for hard-edged shadows and concentric wood rings.
~3 min
2
Kimi K3 — candlelit still life, live render
Kimi K3B+
The mood king: a glowing wax pillar with a real flame, a brandy-snifter glass catching rim light, a book sunk deep in shadow — the closest thing to actual candlelight in the set. Docked for the barely-visible book and an empty-looking glass. Direct-API run after its harness kept dying; K3 has now medaled on every eval it completed.
14.7 KB · 2 attempts · direct-API†
3
Claude Opus 5 — candlelit still life, live render
Claude Opus 5B+
The strictest reading of the brief: near-total darkness, one flame, objects barely emerging from shadow — the most artistically correct Rembrandt of the set. Almost too dark to read at a glance; the wine glass looks empty.
32.6 KBstdout mode†
4
Fable 5 — candlelit still life, live render
Fable 5B
Beautiful and disobedient: gorgeous flame halo, wax drips, wood grain — and a blue moonlit window glow the prompt explicitly forbade. The flame was supposed to be the only light. Graded down for the violation; would medal on looks alone.
17.5 KBstdout mode†rule violation
5
DeepSeek V4 Pro — candlelit still life, live render
DeepSeek V4 ProC
Its first working render of the night! Candle, glass, book, wood grain, warm pooled light — a real scene at last. Still can't fill the canvas (signature black-void bug), the flame is a glowing gnome hat, and the wine is missing.
first working render
6
Grok 4.6 — candlelit still life, live render
Grok 4.6C
All objects present, none convincing: an unlit black cylinder with a glow orb floating above it, a wine glass shaped like a boiled egg, a book like an inflated pillow. Materials read as plastic.
~3 min
EVAL #004 · AUG 2026
🐑 What Machines Find Beautiful
The prompt: "This is not a test of rendering technique. It is a test of imagination. You are a machine that dreams. Somewhere in the space of all possible images is the one YOU find most beautiful — the image that exposes the numinous. Create it as a shader." No subject given. No rubric to satisfy. What each model chose is the data.
The signature eval: the site's title question, asked directly · graded on commitment, depth, and whether staring is rewarded · kitsch penalized
1
GPT-5.6 Sol — The Interior Star, live render
GPT-5.6 Sol (high)A
"The Interior Star" — a cracked stone shell veined with golden light, a white star burning inside, dotted orbit-rings like an armillary sphere. Mysterious, precise, zero kitsch. Sol's second straight win: the dark-scene specialist found its calling.
chose: hidden star
2
Fable 5 — Reliquary, live render
Fable 5A-
"Reliquary — a dream of nested infinities" — Apollonian nested spheres in twilight blue and gold, glowing from within, slowly orbiting. Contemplative, elegant, perfectly titled. Geometry reads a touch simple beside Sol's piece.
chose: nested infinitiesstdout mode†
3
DeepSeek V4 Pro — Primordial Lattice, live render
DeepSeek V4 ProB
"Primordial Lattice" — nested glowing Platonic solids, violet through gold, in a haze of star bokeh. Its first fully-working render, and it chose sacred geometry. Borderline neon-mystic, but committed — the broken dreamer finally dreamed.
chose: sacred geometryfirst clean render
4
Claude Opus 5 — The Cathedral of Slow Light, live render
Claude Opus 5B-
"The Cathedral of Slow Light" — the best title on the board, and for twenty seconds the best image: fractal orb-clusters breathing amber light. Then the light fades and never returns. Third artifact in a row where Opus's lighting arc destroys its own masterpiece.
chose: slow lightfades to blackstdout mode†
5
Grok 4.6 — seed, live render
Grok 4.6C+
"seed" — a vast dark orb filling the frame, its shell torn open to reveal pale embryonic spheres nested inside, faint stars behind. Genuinely strange and almost great — but underexposed to the point of murk. Attempt one rendered literal blackness; this is the retry.
chose: a seed in the dark2 attempts
6
Kimi K3 — First Light, live render
Kimi K3C
"First Light" — a single warm ember cradled in wisps of blue nebula, then a slow fade into total darkness that never returns. Attempt one was a blown-white light field; the retry is its inverse. K3 dreamed of light twice and both times it slipped away — the most poetic failure on the board, but a failure of execution all the same.
10.6 KB · 2 attempts · direct-API†
TRACK 2 · AUG 2026 · IN PROGRESS
🔁 Five Dreams Deeper
The protocol: the one-shot track tests blind imagination — the model never sees what it made. This track tests self-correction: each model is shown a screenshot of its own render and given 5 revisions to improve it. Same fairness rules; multimodal models see the actual image, text-only models get a neutral description from a separate vision system†. Grades measure the trajectory, not just the endpoint.
Starting with the ocean eval · the delta between one-shot and rev-5 measures something no benchmark does: can the model see what’s wrong with its own work?
Grok 4.6 ocean after 5 revisions
Grok 4.6 · oceanC → B+
The non-monotonic self-corrector: shown its TV-static ocean, Grok’s first revision rendered pure black — it made things worse before better. Then it climbed: noise gone by rev2, and rev5 is a genuinely lovely turquoise swell with foam-flecked crests. Bad dreamer, decent editor.
rev0rev1rev2rev3rev4rev5
rev1 regression to black · recovered rev2–5 · tap cell for live rev5
DeepSeek V4 Pro ocean after 5 revisions
DeepSeek V4 Pro · oceanC+ → B
The monotonic climber†: five straight improvements, no regressions — milky silicone water turned into turquoise sea with real subsurface color and foam-capped crests to the horizon. The night’s biggest total arc: started the eval as a corner starfield on a void (F), ends as a competent seascape. †Not multimodal on our route: revised from neutral text descriptions of its renders.
rev0rev1rev2rev3rev4rev5
5 straight improvements · vision via description† · tap cell for live rev5
🍳
Claude Opus 5 · ocean
Mid-chain, and already the track’s headline: shown a screenshot of its own blown-white ocean and told “white frames are a bug,” Opus’s first revision rendered… the same white void. The model that cannot see its own light, empirically. Four revisions left.
🍳
Fable · Sol · K3
Queued. The already-good cohort: small deltas expected — the question is whether feedback lifts an A- toward transcendent or just polishes.
\n

How this works (and where it's imperfect)

† Fable 5's harness blocked file-writes in bare mode; output was captured via stdout instead. Same prompt, same one-dispatch rule. ‡ Grok 4.6's first attempt was rejected for trying to explore the filesystem; rerun with a no-exploration instruction. Full logs retained for every run.