Loading comparison…
Loading benchmark brief…
Benchmark detailsBlind · randomized · Match 1+
Identities and configuration details stay hidden until you vote.
How should I judge?
Judge prompt adherence and overall result quality. Model, harness, and effort stay hidden until you vote.
Filters All configurations+
Which result better satisfies the brief?
Compare only what is visible. Select the stronger render, or record that both or neither satisfy the brief. Identities hidden
Rankings
Each row represents one model × effort group pooled across harnesses and profiles. Ratings are updated from blind pairwise votes.
Efficiency
Why this is not pooled: each model, harness, effort, and profile combination stays separate. Repeated runs of the same configuration are averaged; settings are never averaged together, making tradeoffs between configurations visible.
Advanced filtersAny configuration+
| Efficiency | Confidence | Votes | |||||
|---|---|---|---|---|---|---|---|
| Loading standings… | |||||||
Ratings begin at 1500. Model × Effort pools across harnesses and profiles; Harness uses direct same-model, same-effort comparisons. Configurations with fewer than five votes are marked provisional.
Runs
Search and filter imported runs, then open one to inspect its render, interactive 3D model, and recorded metrics.
| Run | Run date | Model | Harness | Effort | Prompt | Runtime / cost | Agent activity | Geometry |
|---|---|---|---|---|---|---|---|---|
| Loading run history… | ||||||||
Loading run…
How the benchmark works
The evaluation isolates the full tool-using stack while keeping the voting experience simple and blind.
Every run begins with the same natural-language Blender task.
CONTROLLEDModels work through their assigned harness and reasoning setting.
VARIESIdentities and configuration details remain hidden during voting.
HIDDEN UNTIL VOTEPairwise outcomes update ratings for each configuration.
CALCULATEDPrompt, benchmark scene, comparison protocol, and vote outcomes.
Model, reasoning effort, harness, and the resulting Blender scene.
Identities and technical run metrics are revealed without moving the renders.
New configurations remain marked Provisional until enough votes accumulate.
Run the benchmark locally
Use the submission script to run Pi with a selected model, reasoning effort, and canonical prompt. Uploads enter a private staging queue and are published only after administrator review.
chmod +x submit-run.sh
./submit-run.sh --server https://arena.example \
--model PROVIDER/MODEL --effort high --prompt worldtreeThe script does not assume CPU or GPU rendering. It records the actual Blender environment with the session, preview, GLB, and Blend file. Session logs can contain prompts, tool output, and local paths; inspect them before confirming upload.
Preference data for a more open 3D future.
Render Arena is an open benchmark for comparing real Blender runs across the complete agent stack. Votes are stored against an anonymous browser token so one voter cannot submit the same comparison twice.