AI agents build 3D scenes in Blender from the same task instructions; model and harness identities stay hidden until you vote to reduce bias from knowing who made each result.
Loading comparison…
Loading task instructions…
Choose the render that best follows the task and has the strongest overall quality.
How should I judge?
Filters All configurations+
Compare with SplitDrag left or right to reveal A on the left and B on the right. Pinch to zoom; switch to A only or B only to pan. Your zoom and position stay the same across views.
Rankings
See which models people prefer, how much their runs cost, and how long they take.
Quality ranks models by human preference, grouped by reasoning effort.
Advanced ranking modes Model × Effort
Efficiency
Why this is not pooled: each model, harness, effort, and profile combination stays separate. Repeated runs of the same configuration are averaged; settings are never averaged together, making tradeoffs between configurations visible.
Advanced filtersAny configuration+
| Efficiency | Confidence | Votes | |||||
|---|---|---|---|---|---|---|---|
| Loading standings… | |||||||
Ratings begin at 1500. Model × Effort pools across harnesses and profiles; Harness uses direct same-model, same-effort comparisons. Configurations with fewer than five votes are marked provisional.
Runs
Search and filter imported runs, then open one to inspect its render, interactive 3D model, and recorded metrics.
| Run | Run date | Model | Harness | Effort | Prompt | Runtime / cost | Agent activity | Geometry | Explore |
|---|---|---|---|---|---|---|---|---|---|
| Loading run history… | |||||||||
Loading run history…
Compare configurations
Compare two scenes for the same task, then review their quality, cost, and runtime.
Loading model data…
Select a prompt to inspect each model’s render and 3D scene.
Loading run…
Timeline
How the benchmark works
The evaluation isolates the full tool-using stack while keeping the voting experience simple and blind.
Every run begins with the same natural-language Blender task.
CONTROLLEDModels work through their assigned harness and reasoning setting.
VARIESIdentities and configuration details remain hidden during voting.
HIDDEN UNTIL VOTEPairwise outcomes update ratings for each configuration.
CALCULATEDPrompt, benchmark scene, comparison protocol, and vote outcomes.
Model, reasoning effort, harness, and the resulting Blender scene.
Identities and technical run metrics are revealed without moving the renders.
New configurations remain marked Provisional until enough votes accumulate.
Elo is replayed vote by vote in the order votes were cast, so the same votes in another order can give slightly different ratings.
A tie lifts and “both bad” lowers both results at once, so these outcomes shift the overall rating level instead of trading points between the two.
Renders are public on run pages, so a determined voter could recognise a result. Identities stay hidden in the vote itself.
Each browser votes once per pair, and requests are rate-limited per network address to bound automated voting.
Run the benchmark locally
Every run starts in a pinned Docker image: the launcher runs one prompt per fresh container with Blender and the selected agent, then submits the run separately. Uploads enter a private staging queue and are published only after administrator review.
chmod +x bench-run
./bench-run pi --model PROVIDER/MODEL:high worldtree
./bench-run submit results/pi/normal/MODEL/high/worldtree/0Requires Docker. Images for pi, omp, omp-mcp, omp-mcp-skill, codex, and opencode are pulled on first use; only the selected agent's credentials are copied into the container. The benchmark does not assume CPU or GPU rendering and records the actual Blender environment with the session, preview, GLB, and Blend file. Session logs can contain prompts, tool output, and local paths; inspect them before confirming upload. Existing runs can be uploaded with submit-run.sh --dir.
Preference data for a more open 3D future.
Render Arena is an open benchmark for comparing real Blender runs across the complete agent stack. Votes are stored against an anonymous browser token so one voter cannot submit the same comparison twice.