Benchmark detailsBlind · randomized · Match 1+
Identities and configuration details stay hidden until you vote.
How should I judge?
Judge how well each render follows the task and its overall quality. Model, harness, and effort stay hidden until you vote.
Filters All configurations+
Which result best matches the task?
Compare only what you can see. Choose the stronger render, or record that both or neither meet the task. Identities hidden
Loading comparison…
Loading task instructions…
Rankings
Each row represents one model × effort group pooled across harnesses and profiles. Ratings are updated from blind pairwise votes.
Efficiency
Why this is not pooled: each model, harness, effort, and profile combination stays separate. Repeated runs of the same configuration are averaged; settings are never averaged together, making tradeoffs between configurations visible.
Advanced filtersAny configuration+
| Efficiency | Confidence | Votes | |||||
|---|---|---|---|---|---|---|---|
| Loading standings… | |||||||
Ratings begin at 1500. Model × Effort pools across harnesses and profiles; Harness uses direct same-model, same-effort comparisons. Configurations with fewer than five votes are marked provisional.
Runs
Search and filter imported runs, then open one to inspect its render, interactive 3D model, and recorded metrics.
| Run | Run date | Model | Harness | Effort | Prompt | Runtime / cost | Agent activity | Geometry | Explore |
|---|---|---|---|---|---|---|---|---|---|
| Loading run history… | |||||||||
Compare two models
Choose two models to compare their blind-vote ratings, run coverage, efficiency, and strongest effort settings side by side.
Loading model data…
Select a prompt to inspect each model’s render and 3D scene.
Loading run…
Timeline
How the benchmark works
The evaluation isolates the full tool-using stack while keeping the voting experience simple and blind.
Every run begins with the same natural-language Blender task.
CONTROLLEDModels work through their assigned harness and reasoning setting.
VARIESIdentities and configuration details remain hidden during voting.
HIDDEN UNTIL VOTEPairwise outcomes update ratings for each configuration.
CALCULATEDPrompt, benchmark scene, comparison protocol, and vote outcomes.
Model, reasoning effort, harness, and the resulting Blender scene.
Identities and technical run metrics are revealed without moving the renders.
New configurations remain marked Provisional until enough votes accumulate.
Run the benchmark locally
Use the submission script to run Pi or OMP with a selected model, reasoning effort, and canonical prompt. Uploads enter a private staging queue and are published only after administrator review.
chmod +x submit-run.sh
./submit-run.sh --server https://arena.example \
--model PROVIDER/MODEL --effort high --prompt worldtreeThe script does not assume CPU or GPU rendering. It records the actual Blender environment with the session, preview, GLB, and Blend file. Session logs can contain prompts, tool output, and local paths; inspect them before confirming upload.
Preference data for a more open 3D future.
Render Arena is an open benchmark for comparing real Blender runs across the complete agent stack. Votes are stored against an anonymous browser token so one voter cannot submit the same comparison twice.