choose which configurations can be matched
identities revealed · vote counted toward rankings
[4] judge by+
[5] filters: all +
identities sealed until vote
Loading comparison…
Loading task instructions…
best follows the task · strongest overall quality
Reference image
Judge fidelity to the reference: shape, proportions, materials, lighting, and composition. Text and reference-image Elo scores are independent.
[4] / search · filters: all +
Option counts show matching configurations with the other filters applied. Zero-match options are unavailable.
Keyboard shortcuts
- R
- Reset filters
- /
- Focus search
- J / K
- Select next / previous row
- Enter
- Open the selected row when page background has focus
- N / P
- Next / previous page
Shortcuts are inactive while typing in a field, while a dialog or navigation menu is open, or while Ctrl, Alt or Command is held.
rankings · quality
Quality ranks models by human preference, grouped by reasoning effort.
Efficiency
higher is betterWhy this is not pooled: each model, harness, effort, and profile combination stays separate. Repeated runs of the same configuration are averaged; settings are never averaged together, making tradeoffs between configurations visible.
W = wins · L = losses · T = tie / both good · B = both bad. Zero-vote harnesses are unranked; 1500 is their initial rating.
| cost/run · tokens | confidence | |||||
|---|---|---|---|---|---|---|
| Loading standings… | ||||||
Ratings begin at 1500. Model × Effort pools across harnesses and profiles; Harness uses direct same-model, same-effort comparisons. Configurations with fewer than five votes are marked provisional.
[3] filters: all +
Option counts show matching runs with the other filters applied. Zero-match options are unavailable.
Keyboard shortcuts
- R
- Reset filters
- /
- Focus search
- J / K
- Select next / previous run
- Enter
- Open selected run when page background has focus
- C
- Compare selected run
- N / P
- Next / previous page
Shortcuts are inactive while typing in a field, while a dialog or navigation menu is open, or while Ctrl, Alt or Command is held.
runs
search and filter imported runs, then open one to inspect its render, 3d model and recorded metrics.
| run · date | model · harness | effort | prompt | time · cost | agent activity | geometry | compare |
|---|---|---|---|---|---|---|---|
| Loading run history… | |||||||
Loading run history…
effort & harness
effort & harness
Prompt availability uses each side’s selected effort and harness.
select a prompt to inspect each side’s render and 3d scene.
Keyboard shortcuts
- S
- Swap configurations A and B
Shortcuts are inactive while typing in a field, while a dialog or navigation menu is open, or while Ctrl, Alt or Command is held.
compare configurations
loading model data…
loading prompt…
the exact instructions given to the agent. the benchmark task replaces $@.
loading model…
loading prompt…
Loading run…
[8] session timeline · loading trace…
Each side must have this many votes for its model, effort, harness and tool configuration. Votes are pooled across prompts. The default of 10 excludes setups with fewer votes.
Checking how many comparisons this filter hides…
what changes the result
each comparison holds every other setting fixed and changes one: reasoning effort, harness, blender mcp or skills. elo differences are in rating points; cost, tokens and runtime are geometric-mean ratios of the changed setup against the base (2× means twice as much). the smaller line under each value is the range across the matched setups.
Loading insights…
how the benchmark works
the evaluation isolates the full tool-using stack while keeping the voting experience simple and blind.
- 01[controlled]standardized task
every run begins with the same natural-language blender task.
- 02[varies]agent builds
models work through their assigned harness and reasoning setting.
- 03[hidden until vote]blind comparison
identities and configuration details remain hidden during voting.
- 04[calculated]stack-level elo
pairwise outcomes update ratings for each configuration.
- controlled
- prompt, benchmark scene, comparison protocol, and vote outcomes.
- what varies
- model, reasoning effort, harness, and the resulting blender scene.
- after voting
- identities and technical run metrics are revealed without moving the renders.
- low samples
- new configurations remain marked provisional until enough votes accumulate.
- rating order
- Elo is replayed in vote order. Reordering the same outcomes can change ratings. Inspect current ratings and vote records.
- shared outcomes
- Tie / both good and both bad can shift the overall rating level through the reference adjustment below. They do not just exchange points. Compare W/L/T/B outcome counts.
- blindness limits
- Identities are hidden during voting, but public renders can be recognized. Browse the public runs and renders.
- uneven coverage
- Models and configurations have different numbers of runs and votes. Matched comparisons can cover only a few setups. Inspect matched setups, ranges and vote-filter exclusions.
- vote limits
- Each browser votes once per pair, and requests are rate-limited per network address. This does not establish that every vote comes from a different person. Inspect recorded vote totals; individual voter identities and a raw vote-event export are not public.
Every entry starts at 1500 Elo. For eligible same-task votes between different entries, the expected score for A and the rating update are:
EA = 1 / (1 + 10(RB − RA) / 400)
Δ = round(24 × (SA − EA))
R′A = RA + Δ; R′B = RB − Δ
K-factor = 24. R is the rating before the vote. SA is 1 when A wins, 0 when B wins, and 0.5 for tie / both good or both bad. Updates are rounded to integer Elo points.
Shared outcomes: the tie button records “both good.” After the pairwise update, each entry receives a reference adjustment for both good (Q = 1) or both bad (Q = 0):
Eref(R′) = 1 / (1 + 10(1500 − R′) / 400)
R″ = R′ + round(8 × (Q − Eref(R′)))
Reference K-factor = 8. At equal ratings of 1500, an A win produces 1512 / 1488; both good produces 1504 / 1504; both bad produces 1496 / 1496.
Ranking pools: full stack rates model × effort × harness × profile. Model × effort pools across harnesses and profiles and ignores votes within the same model × effort entry. Harness ratings use same-model, same-effort, same-profile votes across different harnesses and exclude shared outcomes entirely. Profile ratings use same-model, same-effort, same-harness votes across different profiles; shared outcomes use the half-point update without the reference adjustment.
Win rate is wins / (wins + losses); both-good and both-bad outcomes are excluded. Explore the ranking pools or inspect full-stack ratings as JSON.
Confidence intervals are not currently calculated. There are no 95% intervals or significance tests for Elo ranks. The “provisional” label means fewer than five votes; “established” means at least five and is not a statistical confidence guarantee.
Insights hold other configuration factors fixed and summarize differences across matched setups. Elo uses an arithmetic mean of rating differences. Cost, tokens and runtime use a geometric mean of available positive ratios: exp(mean(log(after / before))). Missing measurements are excluded from that metric.
Ranges show the smallest and largest observed setup values, not confidence intervals. Setups can share models and runs, and votes can come from repeat voters, so observations are not necessarily independent. Small rating gaps do not establish a reliable ordering.
Inspect vote counts and provisional ratings · Explore matched setup counts and ranges · Inspect run coverage, costs and runtimes · download matched-comparison data as JSON.
pi · omp · omp-mcp
omp-mcp-skill · codex · opencode
claude-code · claude-code-mcp
images pulled on first use
run the benchmark locally
every run starts in a pinned docker image: one prompt per fresh container with blender and the selected agent, then the run is submitted separately.
You need a terminal, Bash, Python 3, and Docker installed and running. Configure credentials for the agent and model provider you plan to use. Model API calls may incur charges. Blender runs inside the downloaded container.
Download bench-run, open a terminal in its download folder, and run the steps below one at a time. PROVIDER/MODEL:high is a placeholder: replace it with your provider, model identifier and a supported reasoning effort. For a model without reasoning, use --model PROVIDER/MODEL --no-reasoning.
Already have a run directory? Download submit-run.sh and use submit-run.sh --dir <path-to-run>.
Before uploading: session logs can contain prompts, tool output and local paths. Inspect the completed run directory before running the submit step.
Edit the highlighted model value before running. Run each step separately.
# 1. Make the downloaded launcher executable
$ chmod +x bench-run
# 2. Replace PROVIDER/MODEL:high, then generate one run
$ ./bench-run pi --model PROVIDER/MODEL:high --result-file benchmark-run-path.txt worldtree
# 3. Read the actual output directory; inspect its files and logs
$ cat benchmark-run-path.txt
# 4. After inspection, submit that exact run directory
$ ./bench-run submit "$(cat benchmark-run-path.txt)"- images for pi, omp, codex, opencode and claude-code are pulled on first use.
- only the selected agent’s credentials are copied into the container.
- no cpu or gpu rendering is assumed; the actual blender environment is recorded with the session, preview, glb and blend file.
- the launcher prints the run directory and saves it to
benchmark-run-path.txt. Its numbered suffix increases for each new run; the submit command reads the saved path.
- run locallybench-run executes the prompt in a fresh container
- inspect, then submitreview the files in the saved run directory, then upload it with bench-run submit
- private staging queueuploads wait here, not yet public
- administrator review publishedthe run appears in history and enters voting
preference data for a more open 3D future.
Render Arena is an open benchmark for comparing real Blender runs across the complete agent stack. Votes are stored against an anonymous browser token so one voter cannot submit the same comparison twice. Read more