BENCHMARK 002MATCH / 0001
TASK

Loading comparison…

Loading benchmark brief…

Evaluation Blind pairwise
Order Randomized
Progress001 / —
Benchmark detailsBlind · randomized · Match 1
ProgressMatch 1 / —
MethodBlind pairwise comparison
RandomizationSides randomized

Identities and configuration details stay hidden until you vote.

How should I judge?

Judge prompt adherence and overall result quality. Model, harness, and effort stay hidden until you vote.

Brief adherenceRequested elements are present and convincingly combined.
CompositionThe scene is clear, intentional, and readable at a glance.
Finish qualityMaterials, lighting, detail, and completeness feel resolved.
Filters All configurations
BLIND PAIRWISE COMPARISON

Which result better satisfies the brief?

Compare only what is visible. Select the stronger render, or record that both or neither satisfy the brief. Identities hidden

Anonymous Blender result A Anonymous Blender result B
Pinch or drag · both views stay aligned
AA / RENDER #—
2D / RGB
2D100%Anonymous Blender result A
BB / RENDER #—
2D / RGB
2D100%Anonymous Blender result B
Record outcome
LIVE STANDINGS02

Rankings

Each row represents one model × effort group pooled across harnesses and profiles. Ratings are updated from blind pairwise votes.

configurations
total votes
Liveupdates
Ranking view
Show
Advanced filtersAny configuration
Configuration rankings
EfficiencyConfidenceVotes
Loading standings…

Ratings begin at 1500. Model × Effort pools across harnesses and profiles; Harness uses direct same-model, same-effort comparisons. Configurations with fewer than five votes are marked provisional.

RUN EXPLORER03

Runs

Search and filter imported runs, then open one to inspect its render, interactive 3D model, and recorded metrics.

Loading runs…total spent
Imported run history
RunRun dateModelHarnessEffortPromptRuntime / costAgent activityGeometry
Loading run history…
RESEARCH METHOD04

How the benchmark works

The evaluation isolates the full tool-using stack while keeping the voting experience simple and blind.

01Standardized brief

Every run begins with the same natural-language Blender task.

CONTROLLED
02Agent builds

Models work through their assigned harness and reasoning setting.

VARIES
03Blind comparison

Identities and configuration details remain hidden during voting.

HIDDEN UNTIL VOTE
04Stack-level Elo

Pairwise outcomes update ratings for each configuration.

CALCULATED
CONTROLLED

Prompt, benchmark scene, comparison protocol, and vote outcomes.

WHAT VARIES

Model, reasoning effort, harness, and the resulting Blender scene.

AFTER VOTING

Identities and technical run metrics are revealed without moving the renders.

LOW SAMPLES

New configurations remain marked Provisional until enough votes accumulate.

CONTRIBUTE A RUN05

Run the benchmark locally

Use the submission script to run Pi with a selected model, reasoning effort, and canonical prompt. Uploads enter a private staging queue and are published only after administrator review.

QUICK START
chmod +x submit-run.sh
./submit-run.sh --server https://arena.example \
  --model PROVIDER/MODEL --effort high --prompt worldtree

The script does not assume CPU or GPU rendering. It records the actual Blender environment with the session, preview, GLB, and Blend file. Session logs can contain prompts, tool output, and local paths; inspect them before confirming upload.

Download submission script
ABOUT THE LAB06

Preference data for a more open 3D future.

Render Arena is an open benchmark for comparing real Blender runs across the complete agent stack. Votes are stored against an anonymous browser token so one voter cannot submit the same comparison twice.

Interactive 3D preview
Render detail