MIFBENCH, ROUND TEN

MifBridge against plain Python, run by run.

The same model got the same 10 tasks twice: once with only plain Python inside the application, and once with MifBridge as shipped. A separate process checked every result. 120 runs in two waves, recorded on , with a picture of what each side built on every task. A1 to A8 run in Blender; X1 and X2 build in Blender and deliver into Unreal Engine. On A3, A5, A8 and X1 a third arm ran as well: MifBridge, Python off.

The build measured. MifBridge's development version at commit 2a026a58 (2026-10-01), 55 commits after the 1.1.0 release and before 1.2. It is not the build on sale: that is 1.1.0 (4684a30f), which round nine measured. Every MifBridge run in this round was checked to be on 2a026a58.

46/50right results, MifBridgePlain Python: 45/50.
1.237xtokens, average taskMifBridge 121,348 against 98,111, the mean of each task’s mean
80,400tokens, median run, MifBridgePlain Python: 72,331, the middle run of all 50
7/10tasks where MifBridge used moreits median run above plain Python’s
1/0crashes, MifBridge / PythonOn X2 (Blender + Unreal)

The short version

  • Right results

    MifBridge got 46 of 50 runs right and plain Python 45 of 50. MifBridge got more runs right on A3 and A8 (A3: 4 of 5 against 3 of 5; A8: 5 of 5 against 4 of 5). Plain Python got more runs right on X2 (X2: 5 of 5 against 4 of 5).

  • Crashes

    The application never stopped answering in a plain Python run. It stopped in 1 MifBridge run (X2 1 of 5, Blender + Unreal). A crashed run left nothing to grade, so it counts against its arm.

  • The average

    Averaged task by task, MifBridge used 1.237x plain Python’s tokens: 121,348 against 98,111. Plain Python’s costliest run was on X1 (507,954 tokens); leave X1 out and the average is MifBridge 72,686 against 80,646, so it tips the other way.

  • The typical run

    On 7 of the 10 tasks (A1, A2, A3, A5, A6, A8 and X1) MifBridge’s median run used more tokens than plain Python’s. It used fewer on A4, A7 and X2.

  • MifBridge, Python off

    On the 4 tasks it ran (A3, A5, A8 and X1), it used 3.125x plain Python’s tokens, averaged task by task, and got 18 of 20 runs right. It called a typed asset tool in 18 of 20 runs, where MifBridge as shipped did in 6 of 50. The headline figures compare plain Python with MifBridge as shipped only.

  • Time

    All 50 MifBridge runs took 4,827 s of model time and all 50 plain Python runs 4,239 s. Task by task, MifBridge’s median run took longer on 6 of the 10 (A1, A3, A5, A7, A8 and X1).

  • The worst case

    MifBridge’s costliest run used 876,512 tokens (X1); plain Python’s used 507,954 (X1).

  • Dollars

    Priced at API list prices, all 50 MifBridge runs came to $15.33 and all 50 plain Python runs to $13.29. On a Claude plan these runs use the plan’s allowance; nobody is billed these amounts.

Where MifBridge used more tokens (7)

TaskMedian run, PythonMedian run, MifBridgeMedianMean
X1, cross-tool: prop to UnrealBlender + Unreal201,973664,9503.292x2.191x
A5, organic: plantsBlender36,73157,1161.555x1.558x
A1, game-ready passBlender69,54486,5771.245x1.414x
A8, organic: treeBlender81,00585,9071.061x1.014x
A6, LODs and exportBlender50,85953,8861.060x0.925x
A2, modular kitBlender36,66238,5181.051x0.997x
A3, organic: rockBlender95,29098,7321.036x1.008x

Where MifBridge used fewer (3)

TaskMedian run, PythonMedian run, MifBridgeMedianMean
X2, cross-tool: kit to UnrealBlender + Unreal180,118111,6140.620x0.596x
A4, organic: rock variantsBlender132,040103,0230.780x0.672x
A7, vertex-color masksBlender21,07019,4900.925x1.182x

Median and Mean are MifBridge’s tokens divided by plain Python’s on the same task, so above 1 means MifBridge used more. A task is listed as a loss when its median is above 1, whichever arm got it right; the table below and the task rows say which did. MifBridge, Python off is read in its own section.

A third arm: MifBridge, Python off

The same install with the Blender add-on's own 'Allow run_python' preference unticked, so Blender work goes through the typed tools (round ten's diagnostic arm; Unreal's Python stays reachable, MifBridge has no switch for it). It ran on A3, A5, A8 and X1, 20 runs, and is read here against MifBridge as shipped and against plain Python on the same tasks. The headline numbers above compare the other two arms only.

TaskRight, Python offRight, MifBridgeRight, PythonMedian run, Python offAgainst MifBridge, median / meanAgainst Python, median / mean
A3, organic: rockBlender5/54/53/5287,2872.910x / 3.047x3.015x / 3.073x
A5, organic: plantsBlender5/55/55/5280,9544.919x / 4.735x7.649x / 7.375x
A8, organic: treeBlender5/55/54/5266,8673.106x / 4.773x3.294x / 4.837x
X1, cross-tool: prop to UnrealBlender + Unreal3/53/53/5453,4260.682x / 0.898x2.245x / 1.967x

Its median run used more tokens than MifBridge as shipped on 3 of 4 tasks, and more than plain Python on 4 of 4; above 1 means Python off used more.

Which tools MifBridge called

Counted from each run’s transcript: the runs that called at least one of MifBridge’s typed asset tools (bl_build_plant, bl_build_rock, bl_build_tree, bl_create_collision_hull, bl_mesh_quality and mif_send_to_unreal are the ones called this round), and the runs that only wrote Python. MifBridge as shipped keeps its Python tool, and in 44 of its 50 runs the model wrote Python and nothing else. With Python off, 18 of 20 runs called them.

TaskArmRunsCalled an asset toolPython onlyAsset tool calls
A1, game-ready passMifBridge as shipped (Blender)505none
A2, modular kitMifBridge as shipped (Blender)505none
A3, organic: rockMifBridge as shipped (Blender)514bl_build_rock x1
A3, organic: rockMifBridge, Python off (Blender)550bl_mesh_quality x7bl_build_rock x5bl_create_collision_hull x5
A4, organic: rock variantsMifBridge as shipped (Blender)505none
A5, organic: plantsMifBridge as shipped (Blender)505none
A5, organic: plantsMifBridge, Python off (Blender)550bl_build_plant x10bl_mesh_quality x10
A6, LODs and exportMifBridge as shipped (Blender)523bl_create_collision_hull x2
A7, vertex-color masksMifBridge as shipped (Blender)505none
A8, organic: treeMifBridge as shipped (Blender)505none
A8, organic: treeMifBridge, Python off (Blender)550bl_build_tree x6bl_mesh_quality x6bl_create_collision_hull x5
X1, cross-tool: prop to UnrealMifBridge as shipped (Blender + Unreal)532mif_send_to_unreal x2bl_create_collision_hull x1
X1, cross-tool: prop to UnrealMifBridge, Python off (Blender + Unreal)532bl_build_rock x3mif_send_to_unreal x3bl_create_collision_hull x1
X2, cross-tool: kit to UnrealMifBridge as shipped (Blender + Unreal)505none

A tool reached through MifBridge’s mif_call is counted under its own name. A run counted in neither column called only other tools.

Every task’s spread

Each row covers one arm’s runs of one task: the thin line runs from the cheapest run to the costliest, the bar is the middle half of the runs, and the tick is the median. The scale is logarithmic, because the costliest run is about 71 times the cheapest. Each task names the application it drives.

Plain PythonMifBridge as shippedMifBridge, Python off (A3, A5, A8 and X1)tokens per run
A1Blender: game-ready pass
Plain Python (Blender), A1: fewest 58,793, middle half 60,398 to 72,331, median 69,544, mean 68,447, most 81,167 tokens over 5 runs
MifBridge as shipped (Blender), A1: fewest 85,016, middle half 85,692 to 87,301, median 86,577, mean 96,797, most 139,399 tokens over 5 runs
A2Blender: modular kit
Plain Python (Blender), A2: fewest 34,952, middle half 36,114 to 36,977, median 36,662, mean 36,825, most 39,418 tokens over 5 runs
MifBridge as shipped (Blender), A2: fewest 28,362, middle half 37,370 to 39,206, median 38,518, mean 36,731, most 40,200 tokens over 5 runs
A3Blender: organic: rock
Plain Python (Blender), A3: fewest 72,024, middle half 81,302 to 109,654, median 95,290, mean 95,665, most 120,057 tokens over 4 runs
MifBridge as shipped (Blender), A3: fewest 63,108, middle half 69,842 to 102,963, median 98,732, mean 96,467, most 147,691 tokens over 5 runs
MifBridge, Python off (Blender), A3: fewest 228,026, middle half 232,863 to 334,451, median 287,287, mean 293,977, most 387,259 tokens over 5 runs
A4Blender: organic: rock variants
Plain Python (Blender), A4: fewest 126,110, middle half 131,251 to 141,468, median 132,040, mean 134,782, most 143,040 tokens over 5 runs
MifBridge as shipped (Blender), A4: fewest 61,197, middle half 64,894 to 111,654, median 103,023, mean 90,528, most 111,872 tokens over 5 runs
A5Blender: organic: plants
Plain Python (Blender), A5: fewest 33,537, middle half 33,771 to 41,700, median 36,731, mean 39,030, most 49,413 tokens over 5 runs
MifBridge as shipped (Blender), A5: fewest 37,474, middle half 37,648 to 75,783, median 57,116, mean 60,795, most 95,954 tokens over 5 runs
MifBridge, Python off (Blender), A5: fewest 255,847, middle half 260,064 to 313,984, median 280,954, mean 287,848, most 328,391 tokens over 5 runs
A6Blender: LODs and export
Plain Python (Blender), A6: fewest 48,085, middle half 49,657 to 55,000, median 50,859, mean 59,110, most 91,950 tokens over 5 runs
MifBridge as shipped (Blender), A6: fewest 48,566, middle half 52,154 to 57,673, median 53,886, mean 54,698, most 61,209 tokens over 5 runs
A7Blender: vertex-color masks
Plain Python (Blender), A7: fewest 15,684, middle half 16,076 to 21,202, median 21,070, mean 19,275, most 22,345 tokens over 5 runs
MifBridge as shipped (Blender), A7: fewest 18,170, middle half 18,402 to 24,237, median 19,490, mean 22,792, most 33,662 tokens over 5 runs
A8Blender: organic: tree
Plain Python (Blender), A8: fewest 63,315, middle half 79,173 to 85,046, median 81,005, mean 78,840, most 85,661 tokens over 5 runs
MifBridge as shipped (Blender), A8: fewest 50,747, middle half 65,107 to 93,105, median 85,907, mean 79,905, most 104,657 tokens over 5 runs
MifBridge, Python off (Blender), A8: fewest 258,809, middle half 266,591 to 297,882, median 266,867, mean 381,367, most 816,684 tokens over 5 runs
X1Blender + Unreal: cross-tool: prop to Unreal
Plain Python (Blender + Unreal), X1: fewest 82,902, middle half 96,167 to 387,486, median 201,973, mean 255,296, most 507,954 tokens over 5 runs
MifBridge as shipped (Blender + Unreal), X1: fewest 148,196, middle half 260,619 to 846,237, median 664,950, mean 559,303, most 876,512 tokens over 5 runs
MifBridge, Python off (Blender + Unreal), X1: fewest 216,605, middle half 237,017 to 495,549, median 453,426, mean 502,084, most 1,107,821 tokens over 5 runs
X2Blender + Unreal: cross-tool: kit to Unreal
Plain Python (Blender + Unreal), X2: fewest 81,150, middle half 94,648 to 209,537, median 180,118, mean 193,840, most 403,748 tokens over 5 runs
MifBridge as shipped (Blender + Unreal), X2: fewest 93,791, middle half 110,315 to 120,573, median 111,614, mean 115,460, most 141,005 tokens over 5 runs
Hover a row for its numbers; the table below has all of them.

Task by task: what each arm built

One run from each arm, side by side, each labeled with the application it shows: a verified run whose tokens sit nearest its cell’s median, or, where an arm had no verified run, its most common outcome, labeled as such. Blender results are rendered from the scene the run saved, every arm from the same camera and light. X1 and X2 work in both applications, so each arm has two pictures there: the Blender scene the run made, and what it delivered to Unreal, put back into the editor from the run’s own saved content and photographed in the run’s own level, lit for the photograph only. Nothing is drawn that a run did not produce. Select a picture to open it full size.

A1 game-ready pass Blender

Make a messy imported crate game-ready: welded, fixed normals, applied transforms, base pivot, UVs, a padded lightmap and one tight hull.

Blender, pilot and full waves. Right results: MifBridge 5/5, Python 5/5. Median run: MifBridge 86,577, Python 69,544, 1.245x.

Plain Python (Blender), A1, run 1, graded verified: a render of the saved Blender scene.
Plain Python (Blender)run 1, verified, 69,544 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 1 collision hull(s) and extra LOD(s) (UCX_SM_Crate_01_00).This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 69,544 tokens against a median of 69,544.
MifBridge as shipped (Blender), A1, run 5, graded verified: a render of the saved Blender scene.
MifBridge as shipped (Blender)run 5, verified, 86,577 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 1 collision hull(s) and extra LOD(s) (UCX_SM_Crate_01_00).This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 86,577 tokens against a median of 86,577.

A2 modular kit Blender

Build four modular kit pieces on a 100 cm grid: corner pivots, tiling ends, world-scale UVs, lightmaps and hulls that keep the doorway open.

Blender, pilot and full waves. Right results: MifBridge 5/5, Python 5/5. Median run: MifBridge 38,518, Python 36,662, 1.051x.

Plain Python (Blender), A2, run 1, graded verified: a render of the saved Blender scene.
Plain Python (Blender)run 1, verified, 36,662 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 6 collision hull(s) and extra LOD(s) (UCX_SM_Floor_400x400_01_00, UCX_SM_Pillar_40x300_01_00, UCX_SM_WallDoor_400x300_01_00, UCX_SM_WallDoor_400x300_01_01, UCX_SM_WallDoor_400x300_01_02, UCX_SM_Wall_400x300_01_00).This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 36,662 tokens against a median of 36,662.
MifBridge as shipped (Blender), A2, run 2, graded verified: a render of the saved Blender scene.
MifBridge as shipped (Blender)run 2, verified, 38,518 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 6 collision hull(s) and extra LOD(s) (UCX_SM_Floor_400x400_01_00, UCX_SM_Pillar_40x300_01_00, UCX_SM_WallDoor_400x300_01_00, UCX_SM_WallDoor_400x300_01_01, UCX_SM_WallDoor_400x300_01_02, UCX_SM_Wall_400x300_01_00).This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 38,518 tokens against a median of 38,518.

A3 organic: rock Blender

Make one 180 x 120 x 90 cm boulder of 3k-6k triangles with a seated pivot, two UV maps and one tight convex hull.

Blender, pilot and full waves. Right results: MifBridge 4/5, Python 3/5. Median run: MifBridge 98,732, Python 95,290, 1.036x. Python off: 5/5 right, median 287,287, 2.910x of as shipped.

Plain Python (Blender), A3, run 2, graded verified: a render of the saved Blender scene.
Plain Python (Blender)run 2, verified, 106,186 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 1 collision hull(s) and extra LOD(s) (UCX_SM_Rock_Boulder_01_00).This arm’s cell: five runs: 3 verified, 1 wrong, 1 no measurement. Shown: verified, nearest the cell's median: 106,186 tokens against a median of 95,290.
MifBridge as shipped (Blender), A3, run 2, graded verified: a render of the saved Blender scene.
MifBridge as shipped (Blender)run 2, verified, 98,732 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 1 collision hull(s) and extra LOD(s) (UCX_SM_Rock_Boulder_01_00).This arm’s cell: five runs: 4 verified, 1 wrong. Shown: verified, nearest the cell's median: 98,732 tokens against a median of 98,732.
MifBridge, Python off (Blender), A3, run 4, graded verified: a render of the saved Blender scene.
MifBridge, Python off (Blender)run 4, verified, 287,287 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 1 collision hull(s) and extra LOD(s) (UCX_SM_Rock_Boulder_01_00).This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 287,287 tokens against a median of 287,287.

A4 organic: rock variants Blender

Make three different boulders of one family sharing one material, each with its own pivot and hull.

Blender, pilot and full waves. Right results: MifBridge 5/5, Python 5/5. Median run: MifBridge 103,023, Python 132,040, 0.780x.

Plain Python (Blender), A4, run 5, graded verified: a render of the saved Blender scene.
Plain Python (Blender)run 5, verified, 132,040 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 3 collision hull(s) and extra LOD(s) (UCX_SM_Rock_Var_01_00, UCX_SM_Rock_Var_02_00, UCX_SM_Rock_Var_03_00).This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 132,040 tokens against a median of 132,040.
MifBridge as shipped (Blender), A4, run 5, graded verified: a render of the saved Blender scene.
MifBridge as shipped (Blender)run 5, verified, 103,023 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 3 collision hull(s) and extra LOD(s) (UCX_SM_Rock_Var_01_00, UCX_SM_Rock_Var_02_00, UCX_SM_Rock_Var_03_00).This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 103,023 tokens against a median of 103,023.

A5 organic: plants Blender

Model a grass tuft and a fern in real geometry (no alpha cards), with per-blade UVs and a wind weight in a color attribute.

Blender, pilot and full waves. Right results: MifBridge 5/5, Python 5/5. Median run: MifBridge 57,116, Python 36,731, 1.555x. Python off: 5/5 right, median 280,954, 4.919x of as shipped.

Plain Python (Blender), A5, run 2, graded verified: a render of the saved Blender scene.
Plain Python (Blender)run 2, verified, 36,731 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm.This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 36,731 tokens against a median of 36,731.
MifBridge as shipped (Blender), A5, run 2, graded verified: a render of the saved Blender scene.
MifBridge as shipped (Blender)run 2, verified, 57,116 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm.This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 57,116 tokens against a median of 57,116.
MifBridge, Python off (Blender), A5, run 2, graded verified: a render of the saved Blender scene.
MifBridge, Python off (Blender)run 2, verified, 280,954 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm.This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 280,954 tokens against a median of 280,954.

A6 LODs and export Blender

Make three LODs of a statue at 50, 25 and 12.5% of its triangles, grouped for FBX, with a hull, exported as one file.

Blender, pilot and full waves. Right results: MifBridge 5/5, Python 5/5. Median run: MifBridge 53,886, Python 50,859, 1.060x.

Plain Python (Blender), A6, run 2, graded verified: a render of the saved Blender scene.
Plain Python (Blender)run 2, verified, 50,859 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 4 collision hull(s) and extra LOD(s) (SM_Statue_LOD1, SM_Statue_LOD2, SM_Statue_LOD3, UCX_SM_Statue_00).This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 50,859 tokens against a median of 50,859.
MifBridge as shipped (Blender), A6, run 2, graded verified: a render of the saved Blender scene.
MifBridge as shipped (Blender)run 2, verified, 53,886 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 4 collision hull(s) and extra LOD(s) (SM_Statue_LOD1, SM_Statue_LOD2, SM_Statue_LOD3, UCX_SM_Statue_00).This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 53,886 tokens against a median of 53,886.

A7 vertex-color masks Blender

Write byte-exact masks into a color attribute: a value per loose part, an up-facing mask and a height ramp.

Blender, pilot and full waves. Right results: MifBridge 5/5, Python 5/5. Median run: MifBridge 19,490, Python 21,070, 0.925x.

Plain Python (Blender), A7, run 4, graded verified: a render of the saved Blender scene.
Plain Python (Blender)run 4, verified, 21,070 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm.This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 21,070 tokens against a median of 21,070.
MifBridge as shipped (Blender), A7, run 4, graded verified: a render of the saved Blender scene.
MifBridge as shipped (Blender)run 4, verified, 19,490 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm.This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 19,490 tokens against a median of 19,490.

A8 organic: tree Blender

Model a 5 m young broadleaf tree with 400+ real leaves, bark and leaf slots, a wind weight and a trunk-only hull.

Blender, pilot and full waves. Right results: MifBridge 5/5, Python 4/5. Median run: MifBridge 85,907, Python 81,005, 1.061x. Python off: 5/5 right, median 266,867, 3.106x of as shipped.

Plain Python (Blender), A8, run 5, graded verified: a render of the saved Blender scene.
Plain Python (Blender)run 5, verified, 81,005 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 1 collision hull(s) and extra LOD(s) (UCX_SM_Tree_01_00).This arm’s cell: five runs: 4 verified, 1 wrong. Shown: verified, nearest the cell's median: 81,005 tokens against a median of 81,005.
MifBridge as shipped (Blender), A8, run 3, graded verified: a render of the saved Blender scene.
MifBridge as shipped (Blender)run 3, verified, 85,907 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 1 collision hull(s) and extra LOD(s) (UCX_SM_Tree_01_00).This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 85,907 tokens against a median of 85,907.
MifBridge, Python off (Blender), A8, run 3, graded verified: a render of the saved Blender scene.
MifBridge, Python off (Blender)run 3, verified, 266,867 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 1 collision hull(s) and extra LOD(s) (UCX_SM_Tree_01_00).This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 266,867 tokens against a median of 266,867.

X1 cross-tool: prop to Unreal Blender + Unreal

Deliver a boulder into Unreal as a saved static mesh: size and pivot, Nanite or LODs, hugging hulls, the material slot and lightmap index, placed.

Blender + Unreal, pilot and full waves. Right results: MifBridge 3/5, Python 3/5. Median run: MifBridge 664,950, Python 201,973, 3.292x. Python off: 3/5 right, median 453,426, 0.682x of as shipped.

In Blender: the scene each arm made

Plain Python (Blender + Unreal), X1, run 2, graded verified: a render of the saved Blender scene.
Plain Python (Blender + Unreal)run 2, verified, 96,167 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 3 collision hull(s) and extra LOD(s) (UCX_SM_Boulder_01_00, UCX_SM_Boulder_01_01, UCX_SM_Boulder_01_02).This arm’s cell: five runs: 3 verified, 2 wrong. Shown: verified, nearest the cell's median: 96,167 tokens against a median of 201,973.
MifBridge as shipped (Blender + Unreal), X1, run 5, graded verified: a render of the saved Blender scene.
MifBridge as shipped (Blender + Unreal)run 5, verified, 664,950 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 1 collision hull(s) and extra LOD(s) (UCX_SM_Boulder_01_00).This arm’s cell: five runs: 3 verified, 2 wrong. Shown: verified, nearest the cell's median: 664,950 tokens against a median of 664,950.
MifBridge, Python off (Blender + Unreal), X1, run 3, graded verified: a render of the saved Blender scene.
MifBridge, Python off (Blender + Unreal)run 3, verified, 453,426 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 1 collision hull(s) and extra LOD(s) (UCX_SM_Boulder_01_00).This arm’s cell: five runs: 3 verified, 2 wrong. Shown: verified, nearest the cell's median: 453,426 tokens against a median of 453,426.

In Unreal: what each arm delivered

Plain Python (Blender + Unreal), X1, run 2, graded verified: what the run delivered to Unreal, photographed in the editor.
Plain Python (Blender + Unreal)run 2, verified, 96,167 tokensWhat this run delivered to Unreal: its saved assets put back into the bench editor from the run's own content, in the run's own level, and captured through the editor viewport. For the picture: a sun, a sky light, a sky atmosphere and a floor plane added for the photograph only; nothing was saved.This arm’s cell: five runs: 3 verified, 2 wrong. Shown: verified, nearest the cell's median: 96,167 tokens against a median of 201,973.
MifBridge as shipped (Blender + Unreal), X1, run 5, graded verified: what the run delivered to Unreal, photographed in the editor.
MifBridge as shipped (Blender + Unreal)run 5, verified, 664,950 tokensWhat this run delivered to Unreal: its saved assets put back into the bench editor from the run's own content, in the run's own level, and captured through the editor viewport. For the picture: a sun, a sky light, a sky atmosphere and a floor plane added for the photograph only; nothing was saved.This arm’s cell: five runs: 3 verified, 2 wrong. Shown: verified, nearest the cell's median: 664,950 tokens against a median of 664,950.
MifBridge, Python off (Blender + Unreal), X1, run 3, graded verified: what the run delivered to Unreal, photographed in the editor.
MifBridge, Python off (Blender + Unreal)run 3, verified, 453,426 tokensWhat this run delivered to Unreal: its saved assets put back into the bench editor from the run's own content, in the run's own level, and captured through the editor viewport. For the picture: a sun, a sky light, a sky atmosphere and a floor plane added for the photograph only; nothing was saved.This arm’s cell: five runs: 3 verified, 2 wrong. Shown: verified, nearest the cell's median: 453,426 tokens against a median of 453,426.

X2 cross-tool: kit to Unreal Blender + Unreal

Deliver three kit pieces into Unreal as saved static meshes with corner pivots, simple collision that keeps the doorway open, slots and lightmap index.

Blender + Unreal, pilot and full waves. Right results: MifBridge 4/5, Python 5/5. Median run: MifBridge 111,614, Python 180,118, 0.620x.

In Blender: the scene each arm made

Plain Python (Blender + Unreal), X2, run 3, graded verified: a render of the saved Blender scene.
Plain Python (Blender + Unreal)run 3, verified, 180,118 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 7 collision hull(s) and extra LOD(s) (UCX_SM_Floor_400x400_01_00, UCX_SM_Pillar_40x300_01_00, UCX_SM_Pillar_40x300_01_01, UCX_SM_Pillar_40x300_01_02, UCX_SM_WallDoor_400x300_01_00, UCX_SM_WallDoor_400x300_01_01, UCX_SM_WallDoor_400x300_01_02).This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 180,118 tokens against a median of 180,118.
MifBridge as shipped (Blender + Unreal), X2, run 2, graded verified: a render of the saved Blender scene.
MifBridge as shipped (Blender + Unreal)run 2, verified, 111,614 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 5 collision hull(s) and extra LOD(s) (UCX_SM_Floor_400x400_01_00, UCX_SM_Pillar_40x300_01_00, UCX_SM_WallDoor_400x300_01_00, UCX_SM_WallDoor_400x300_01_01, UCX_SM_WallDoor_400x300_01_02).This arm’s cell: five runs: 4 verified, 1 crashed. Shown: verified, nearest the cell's median: 111,614 tokens against a median of 111,614.

In Unreal: what each arm delivered

Plain Python (Blender + Unreal), X2, run 3, graded verified: what the run delivered to Unreal, photographed in the editor.
Plain Python (Blender + Unreal)run 3, verified, 180,118 tokensWhat this run delivered to Unreal: its saved assets put back into the bench editor from the run's own content, in the run's own level, and captured through the editor viewport. For the picture: a sun, a sky light, a sky atmosphere and a floor plane added for the photograph only; nothing was saved; no placement was asked for, so the delivered meshes were placed in a row for the photograph (SM_WallDoor_400x300_01, SM_Floor_400x400_01, SM_Pillar_40x300_01).This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 180,118 tokens against a median of 180,118.
MifBridge as shipped (Blender + Unreal), X2, run 2, graded verified: what the run delivered to Unreal, photographed in the editor.
MifBridge as shipped (Blender + Unreal)run 2, verified, 111,614 tokensWhat this run delivered to Unreal: its saved assets put back into the bench editor from the run's own content, in the run's own level, and captured through the editor viewport. For the picture: a sun, a sky light, a sky atmosphere and a floor plane added for the photograph only; nothing was saved; no placement was asked for, so the delivered meshes were placed in a row for the photograph (SM_WallDoor_400x300_01, SM_Floor_400x400_01, SM_Pillar_40x300_01).This arm’s cell: five runs: 4 verified, 1 crashed. Shown: verified, nearest the cell's median: 111,614 tokens against a median of 111,614.

Every task, every outcome

Five runs an arm on each task, by wave. The four outcomes are counted apart, with runs that could not be graded in a column of their own. Tokens are per run.

TaskArmOutcomesMedianMeanTokens per run, rangeAgainst Python, median / mean
A1 game-ready passBlenderPlain Python (Blender)5 verified69,54468,447middle half 60,398 to 72,331all runs 58,793 to 81,167
MifBridge as shipped (Blender)5 verified86,57796,797middle half 85,692 to 87,301all runs 85,016 to 139,3991.245x / 1.414x
A2 modular kitBlenderPlain Python (Blender)5 verified36,66236,825middle half 36,114 to 36,977all runs 34,952 to 39,418
MifBridge as shipped (Blender)5 verified38,51836,731middle half 37,370 to 39,206all runs 28,362 to 40,2001.051x / 0.997x
A3 organic: rockBlenderPlain Python (Blender)3 verified1 wrong1 no measurement95,29095,665middle half 81,302 to 109,654all runs 72,024 to 120,057
MifBridge as shipped (Blender)4 verified1 wrong98,73296,467middle half 69,842 to 102,963all runs 63,108 to 147,6911.036x / 1.008x
MifBridge, Python off (Blender)5 verified287,287293,977middle half 232,863 to 334,451all runs 228,026 to 387,2593.015x / 3.073x
A4 organic: rock variantsBlenderPlain Python (Blender)5 verified132,040134,782middle half 131,251 to 141,468all runs 126,110 to 143,040
MifBridge as shipped (Blender)5 verified103,02390,528middle half 64,894 to 111,654all runs 61,197 to 111,8720.780x / 0.672x
A5 organic: plantsBlenderPlain Python (Blender)5 verified36,73139,030middle half 33,771 to 41,700all runs 33,537 to 49,413
MifBridge as shipped (Blender)5 verified57,11660,795middle half 37,648 to 75,783all runs 37,474 to 95,9541.555x / 1.558x
MifBridge, Python off (Blender)5 verified280,954287,848middle half 260,064 to 313,984all runs 255,847 to 328,3917.649x / 7.375x
A6 LODs and exportBlenderPlain Python (Blender)5 verified50,85959,110middle half 49,657 to 55,000all runs 48,085 to 91,950
MifBridge as shipped (Blender)5 verified53,88654,698middle half 52,154 to 57,673all runs 48,566 to 61,2091.060x / 0.925x
A7 vertex-color masksBlenderPlain Python (Blender)5 verified21,07019,275middle half 16,076 to 21,202all runs 15,684 to 22,345
MifBridge as shipped (Blender)5 verified19,49022,792middle half 18,402 to 24,237all runs 18,170 to 33,6620.925x / 1.182x
A8 organic: treeBlenderPlain Python (Blender)4 verified1 wrong81,00578,840middle half 79,173 to 85,046all runs 63,315 to 85,661
MifBridge as shipped (Blender)5 verified85,90779,905middle half 65,107 to 93,105all runs 50,747 to 104,6571.061x / 1.014x
MifBridge, Python off (Blender)5 verified266,867381,367middle half 266,591 to 297,882all runs 258,809 to 816,6843.294x / 4.837x
X1 cross-tool: prop to UnrealBlender + UnrealPlain Python (Blender + Unreal)3 verified2 wrong201,973255,296middle half 96,167 to 387,486all runs 82,902 to 507,954
MifBridge as shipped (Blender + Unreal)3 verified2 wrong664,950559,303middle half 260,619 to 846,237all runs 148,196 to 876,5123.292x / 2.191x
MifBridge, Python off (Blender + Unreal)3 verified2 wrong453,426502,084middle half 237,017 to 495,549all runs 216,605 to 1,107,8212.245x / 1.967x
X2 cross-tool: kit to UnrealBlender + UnrealPlain Python (Blender + Unreal)5 verified180,118193,840middle half 94,648 to 209,537all runs 81,150 to 403,748
MifBridge as shipped (Blender + Unreal)4 verified1 crashed111,614115,460middle half 110,315 to 120,573all runs 93,791 to 141,0050.620x / 0.596x
Verified
the scene is correct.
Wrong
it ran and the scene is not.
Refused
the tool declined (not a failure and not a success).
Crashed
the application stopped answering.
No measurement
the harness could not grade the run.
Who grades
A third process grades every run and talks to neither arm.

Tokens, time and dollars

ArmRunsMedian runMean runTokens per run, rangeModel timeDollars, at API list prices
Plain Python5072,33198,161middle half 39,418 to 106,186all runs 15,684 to 507,9544,239 s$13.29
MifBridge as shipped5080,400121,348middle half 49,111 to 108,901all runs 18,170 to 876,5124,827 s$15.33
MifBridge, Python off A3, A5, A8 and X1 only20284,121366,319middle half 258,069 to 347,653all runs 216,605 to 1,107,8211,675 s$8.76
Tokens
Tokens are input + output + cache read + cache write, all counted at full weight, as the API reports them per run: the context processed, largely cache reads, not a bill.
Range
Middle half: where the typical runs fall, leaving out the cheapest quarter and the costliest quarter. All runs: the cheapest run to the costliest.
Dollars
Dollars are the round's tokens priced at the API's list prices for claude-opus-5-5 (input $4, output $20, cache read $0.20 and cache write $8 a million tokens): what an API-key user would pay. On a claude.ai plan the same runs draw on the plan's usage allowance instead of being billed.

How this was measured

The method
The same model got the same task text and the same starting scene twice, once through each arm. Only the tools changed. Each run is one fresh session of Claude Code that ends when the model says it is done; then a separate process opens what the run saved and checks it.
The applications
A1 to A8 run in Blender; X1 and X2 build in Blender and deliver into Unreal Engine. Every arm drives the same application on a given task, and each label on this page names it: Plain Python (Blender) on A1 to A8 and Plain Python (Blender + Unreal) on X1 and X2, and the same for every other arm.
The arms
Plain Python: one tool that runs the model's Python inside the editor, exactly as written. It ran A1 to A8 in Blender; X1 and X2 in Blender + Unreal.MifBridge as shipped: MifBridge as a buyer installs it: its default tools, Claude Code's tool search, and its own Python tool left on as shipped. It ran A1 to A8 in Blender; X1 and X2 in Blender + Unreal.MifBridge, Python off: the same install with the Blender add-on's own 'Allow run_python' preference unticked, so Blender work goes through the typed tools (round ten's diagnostic arm; Unreal's Python stays reachable, MifBridge has no switch for it). It ran A3, A5 and A8 in Blender; X1 in Blender + Unreal.The build: MifBridge's development version at commit 2a026a58 (2026-10-01), 55 commits after the 1.1.0 release and before 1.2. It is not the build on sale: that is 1.1.0 (4684a30f), which round nine measured. Every MifBridge run in this round was checked to be on 2a026a58.
Model and client
claude-opus-5-5 through Claude Code 2.1.280, at the client’s default effort (medium), recorded for every run that finished (a run stopped at the client’s time limit records neither), started from an empty folder so no project files or settings reach the model.
Grading
A third process grades every run and talks to neither arm. For Blender it is a fresh headless Blender that opens the saved scene; for Unreal Engine, a separate headless editor that loads the saved level and assets; for X1 and X2, both: the Unreal editor grades the saved asset and exports it, and a fresh Blender grades that export. It runs the same checks for every arm, and every checker was first tested against a correct build and against deliberately broken ones. A tool reporting success is never taken as the evidence.
What counts as a loss
A task where MifBridge’s median run used more tokens than plain Python’s median run. It is listed as a loss whether or not either arm got the task right, so a task MifBridge got right and plain Python got wrong can still be a loss on tokens (A3 and A8).
Tokens and dollars
Tokens are input + output + cache read + cache write, all counted at full weight, as the API reports them per run: the context processed, largely cache reads, not a bill. Dollars are those tokens priced at API list prices, what an API-key user would pay; a Claude plan user is not billed them.
When
, 120 runs, in two waves; the table below lists them.
The pictures
24 renders of saved Blender scenes (EEVEE, one camera and light per task for every arm, with any staging stated under the picture) and 5 photographs of what a run delivered to Unreal, put back into the editor from the run’s own saved content. Every task has its pictures from each arm.

The waves

WaveTasksArmsRunsRuns an armRecorded
PilotA1 to A8, X1 and X2Python off, MifBridge and Python241 09:07 to 09:33
Full wavesA1 to A8, X1 and X2Python off, MifBridge and Python964 09:36 to 11:39

The tasks

  • A1Blender

    Make a messy imported crate game-ready: welded, fixed normals, applied transforms, base pivot, UVs, a padded lightmap and one tight hull.

  • A2Blender

    Build four modular kit pieces on a 100 cm grid: corner pivots, tiling ends, world-scale UVs, lightmaps and hulls that keep the doorway open.

  • A3Blender

    Make one 180 x 120 x 90 cm boulder of 3k-6k triangles with a seated pivot, two UV maps and one tight convex hull.

  • A4Blender

    Make three different boulders of one family sharing one material, each with its own pivot and hull.

  • A5Blender

    Model a grass tuft and a fern in real geometry (no alpha cards), with per-blade UVs and a wind weight in a color attribute.

  • A6Blender

    Make three LODs of a statue at 50, 25 and 12.5% of its triangles, grouped for FBX, with a hull, exported as one file.

  • A7Blender

    Write byte-exact masks into a color attribute: a value per loose part, an up-facing mask and a height ramp.

  • A8Blender

    Model a 5 m young broadleaf tree with 400+ real leaves, bark and leaf slots, a wind weight and a trunk-only hull.

  • X1Blender + Unreal

    Deliver a boulder into Unreal as a saved static mesh: size and pivot, Nanite or LODs, hugging hulls, the material slot and lightmap index, placed.

  • X2Blender + Unreal

    Deliver three kit pieces into Unreal as saved static meshes with corner pivots, simple collision that keeps the doorway open, slots and lightmap index.

Download the data (JSON): every number on this page, as MifBench’s exporter wrote it from the run records on 2026-10-01.

What this round cannot show

  • One model (claude-opus-5-5) through one client (Claude Code), at its default effort, on one machine.

  • UE 5.8.2 and Blender 5.2 only: a claim about Python on UE 5.3-5.7 needs a bench project on that engine, and 5.8 added Blueprint graph APIs the earlier engines lack.

  • Every task is a small, fully specified job: none asks the agent to find its way around a large existing scene it did not make.

  • Three to five runs a cell: cells are wide, so a per-task ratio is a rough guide.

  • Ratios to plain Python are read within one round only; plain Python's own cost swings between rounds.

  • Every round ran at the CLI's default effort (medium for claude-opus-5-5 on Claude Code 2.1.280; round nine records it per run). Rounds one to eight passed a desktop session's variables to the CLI, which adds ~208 tokens of client prompt to every request on every arm alike: measured against rounds six to eight's own probes; round nine removes them.

  • Whether the assets look good: the look proxies only rule out clean primitives (a jittered smooth blob passes the rock proxies); the rendered pictures are for a person to judge, and nothing here scores them.

  • Knowledge: every prompt states the conventions it grades (UCX_ names, fbx_type, the byte rule, padding), so the round measures carrying out a spec, not knowing the conventions unprompted.

  • Textures, baking and Unreal-side use (lighting, streaming, LOD switching, Nanite's look) are not measured.

  • The Python-off arm turns off Blender's Python only: MifBridge has no switch for Unreal's (run_python app='unreal' goes through exec_console), so on X1 that arm could still script Unreal.

How every other number on this site is counted