The same model got the same 10 tasks twice: once with only plain Python inside the application, and once with MifBridge as shipped. A separate process checked every result. 120 runs in two waves, recorded on , with a picture of what each side built on every task. A1 to A8 run in Blender; X1 and X2 build in Blender and deliver into Unreal Engine. On A3, A5, A8 and X1 a third arm ran as well: MifBridge, Python off.
The build measured. MifBridge's development version at commit 2a026a58 (2026-10-01), 55 commits after the 1.1.0 release and before 1.2. It is not the build on sale: that is 1.1.0 (4684a30f), which round nine measured. Every MifBridge run in this round was checked to be on 2a026a58.
46/50right results, MifBridgePlain Python: 45/50.
1.237xtokens, average taskMifBridge 121,348 against 98,111, the mean of each task’s mean
80,400tokens, median run, MifBridgePlain Python: 72,331, the middle run of all 50
7/10tasks where MifBridge used moreits median run above plain Python’s
MifBridge got 46 of 50 runs right and plain Python 45 of 50. MifBridge got more runs right on A3 and A8 (A3: 4 of 5 against 3 of 5; A8: 5 of 5 against 4 of 5). Plain Python got more runs right on X2 (X2: 5 of 5 against 4 of 5).
Crashes
The application never stopped answering in a plain Python run. It stopped in 1 MifBridge run (X2 1 of 5, Blender + Unreal). A crashed run left nothing to grade, so it counts against its arm.
The average
Averaged task by task, MifBridge used 1.237x plain Python’s tokens: 121,348 against 98,111. Plain Python’s costliest run was on X1 (507,954 tokens); leave X1 out and the average is MifBridge 72,686 against 80,646, so it tips the other way.
The typical run
On 7 of the 10 tasks (A1, A2, A3, A5, A6, A8 and X1) MifBridge’s median run used more tokens than plain Python’s. It used fewer on A4, A7 and X2.
MifBridge, Python off
On the 4 tasks it ran (A3, A5, A8 and X1), it used 3.125x plain Python’s tokens, averaged task by task, and got 18 of 20 runs right. It called a typed asset tool in 18 of 20 runs, where MifBridge as shipped did in 6 of 50. The headline figures compare plain Python with MifBridge as shipped only.
Time
All 50 MifBridge runs took 4,827 s of model time and all 50 plain Python runs 4,239 s. Task by task, MifBridge’s median run took longer on 6 of the 10 (A1, A3, A5, A7, A8 and X1).
The worst case
MifBridge’s costliest run used 876,512 tokens (X1); plain Python’s used 507,954 (X1).
Dollars
Priced at API list prices, all 50 MifBridge runs came to $15.33 and all 50 plain Python runs to $13.29. On a Claude plan these runs use the plan’s allowance; nobody is billed these amounts.
Where MifBridge used more tokens (7)
Task
Median run, Python
Median run, MifBridge
Median
Mean
X1, cross-tool: prop to UnrealBlender + Unreal
201,973
664,950
3.292x
2.191x
A5, organic: plantsBlender
36,731
57,116
1.555x
1.558x
A1, game-ready passBlender
69,544
86,577
1.245x
1.414x
A8, organic: treeBlender
81,005
85,907
1.061x
1.014x
A6, LODs and exportBlender
50,859
53,886
1.060x
0.925x
A2, modular kitBlender
36,662
38,518
1.051x
0.997x
A3, organic: rockBlender
95,290
98,732
1.036x
1.008x
Where MifBridge used fewer (3)
Task
Median run, Python
Median run, MifBridge
Median
Mean
X2, cross-tool: kit to UnrealBlender + Unreal
180,118
111,614
0.620x
0.596x
A4, organic: rock variantsBlender
132,040
103,023
0.780x
0.672x
A7, vertex-color masksBlender
21,070
19,490
0.925x
1.182x
Median and Mean are MifBridge’s tokens divided by plain Python’s on the same task, so above 1 means MifBridge used more. A task is listed as a loss when its median is above 1, whichever arm got it right; the table below and the task rows say which did. MifBridge, Python off is read in its own section.
A third arm: MifBridge, Python off
The same install with the Blender add-on's own 'Allow run_python' preference unticked, so Blender work goes through the typed tools (round ten's diagnostic arm; Unreal's Python stays reachable, MifBridge has no switch for it). It ran on A3, A5, A8 and X1, 20 runs, and is read here against MifBridge as shipped and against plain Python on the same tasks. The headline numbers above compare the other two arms only.
Task
Right, Python off
Right, MifBridge
Right, Python
Median run, Python off
Against MifBridge, median / mean
Against Python, median / mean
A3, organic: rockBlender
5/5
4/5
3/5
287,287
2.910x / 3.047x
3.015x / 3.073x
A5, organic: plantsBlender
5/5
5/5
5/5
280,954
4.919x / 4.735x
7.649x / 7.375x
A8, organic: treeBlender
5/5
5/5
4/5
266,867
3.106x / 4.773x
3.294x / 4.837x
X1, cross-tool: prop to UnrealBlender + Unreal
3/5
3/5
3/5
453,426
0.682x / 0.898x
2.245x / 1.967x
Its median run used more tokens than MifBridge as shipped on 3 of 4 tasks, and more than plain Python on 4 of 4; above 1 means Python off used more.
Which tools MifBridge called
Counted from each run’s transcript: the runs that called at least one of MifBridge’s typed asset tools (bl_build_plant, bl_build_rock, bl_build_tree, bl_create_collision_hull, bl_mesh_quality and mif_send_to_unreal are the ones called this round), and the runs that only wrote Python. MifBridge as shipped keeps its Python tool, and in 44 of its 50 runs the model wrote Python and nothing else. With Python off, 18 of 20 runs called them.
A tool reached through MifBridge’s mif_call is counted under its own name. A run counted in neither column called only other tools.
Every task’s spread
Each row covers one arm’s runs of one task: the thin line runs from the cheapest run to the costliest, the bar is the middle half of the runs, and the tick is the median. The scale is logarithmic, because the costliest run is about 71 times the cheapest. Each task names the application it drives.
Plain PythonMifBridge as shippedMifBridge, Python off (A3, A5, A8 and X1)tokens per run
20k50k100k200k500k1M
A1Blender: game-ready pass
Plain Python (Blender), A1: fewest 58,793, middle half 60,398 to 72,331, median 69,544, mean 68,447, most 81,167 tokens over 5 runsPython
MifBridge as shipped (Blender), A1: fewest 85,016, middle half 85,692 to 87,301, median 86,577, mean 96,797, most 139,399 tokens over 5 runsMifBridge
A2Blender: modular kit
Plain Python (Blender), A2: fewest 34,952, middle half 36,114 to 36,977, median 36,662, mean 36,825, most 39,418 tokens over 5 runsPython
MifBridge as shipped (Blender), A2: fewest 28,362, middle half 37,370 to 39,206, median 38,518, mean 36,731, most 40,200 tokens over 5 runsMifBridge
A3Blender: organic: rock
Plain Python (Blender), A3: fewest 72,024, middle half 81,302 to 109,654, median 95,290, mean 95,665, most 120,057 tokens over 4 runsPython
MifBridge as shipped (Blender), A3: fewest 63,108, middle half 69,842 to 102,963, median 98,732, mean 96,467, most 147,691 tokens over 5 runsMifBridge
MifBridge, Python off (Blender), A3: fewest 228,026, middle half 232,863 to 334,451, median 287,287, mean 293,977, most 387,259 tokens over 5 runsPython off
A4Blender: organic: rock variants
Plain Python (Blender), A4: fewest 126,110, middle half 131,251 to 141,468, median 132,040, mean 134,782, most 143,040 tokens over 5 runsPython
MifBridge as shipped (Blender), A4: fewest 61,197, middle half 64,894 to 111,654, median 103,023, mean 90,528, most 111,872 tokens over 5 runsMifBridge
A5Blender: organic: plants
Plain Python (Blender), A5: fewest 33,537, middle half 33,771 to 41,700, median 36,731, mean 39,030, most 49,413 tokens over 5 runsPython
MifBridge as shipped (Blender), A5: fewest 37,474, middle half 37,648 to 75,783, median 57,116, mean 60,795, most 95,954 tokens over 5 runsMifBridge
MifBridge, Python off (Blender), A5: fewest 255,847, middle half 260,064 to 313,984, median 280,954, mean 287,848, most 328,391 tokens over 5 runsPython off
A6Blender: LODs and export
Plain Python (Blender), A6: fewest 48,085, middle half 49,657 to 55,000, median 50,859, mean 59,110, most 91,950 tokens over 5 runsPython
MifBridge as shipped (Blender), A6: fewest 48,566, middle half 52,154 to 57,673, median 53,886, mean 54,698, most 61,209 tokens over 5 runsMifBridge
A7Blender: vertex-color masks
Plain Python (Blender), A7: fewest 15,684, middle half 16,076 to 21,202, median 21,070, mean 19,275, most 22,345 tokens over 5 runsPython
MifBridge as shipped (Blender), A7: fewest 18,170, middle half 18,402 to 24,237, median 19,490, mean 22,792, most 33,662 tokens over 5 runsMifBridge
A8Blender: organic: tree
Plain Python (Blender), A8: fewest 63,315, middle half 79,173 to 85,046, median 81,005, mean 78,840, most 85,661 tokens over 5 runsPython
MifBridge as shipped (Blender), A8: fewest 50,747, middle half 65,107 to 93,105, median 85,907, mean 79,905, most 104,657 tokens over 5 runsMifBridge
MifBridge, Python off (Blender), A8: fewest 258,809, middle half 266,591 to 297,882, median 266,867, mean 381,367, most 816,684 tokens over 5 runsPython off
X1Blender + Unreal: cross-tool: prop to Unreal
Plain Python (Blender + Unreal), X1: fewest 82,902, middle half 96,167 to 387,486, median 201,973, mean 255,296, most 507,954 tokens over 5 runsPython
MifBridge as shipped (Blender + Unreal), X1: fewest 148,196, middle half 260,619 to 846,237, median 664,950, mean 559,303, most 876,512 tokens over 5 runsMifBridge
MifBridge, Python off (Blender + Unreal), X1: fewest 216,605, middle half 237,017 to 495,549, median 453,426, mean 502,084, most 1,107,821 tokens over 5 runsPython off
X2Blender + Unreal: cross-tool: kit to Unreal
Plain Python (Blender + Unreal), X2: fewest 81,150, middle half 94,648 to 209,537, median 180,118, mean 193,840, most 403,748 tokens over 5 runsPython
MifBridge as shipped (Blender + Unreal), X2: fewest 93,791, middle half 110,315 to 120,573, median 111,614, mean 115,460, most 141,005 tokens over 5 runsMifBridge
Hover a row for its numbers; the table below has all of them.
Task by task: what each arm built
One run from each arm, side by side, each labeled with the application it shows: a verified run whose tokens sit nearest its cell’s median, or, where an arm had no verified run, its most common outcome, labeled as such. Blender results are rendered from the scene the run saved, every arm from the same camera and light. X1 and X2 work in both applications, so each arm has two pictures there: the Blender scene the run made, and what it delivered to Unreal, put back into the editor from the run’s own saved content and photographed in the run’s own level, lit for the photograph only. Nothing is drawn that a run did not produce. Select a picture to open it full size.
A1 game-ready pass Blender
Make a messy imported crate game-ready: welded, fixed normals, applied transforms, base pivot, UVs, a padded lightmap and one tight hull.
Blender, pilot and full waves. Right results: MifBridge 5/5, Python 5/5. Median run: MifBridge 86,577, Python 69,544, 1.245x.
BlenderPlain Python (Blender)run 1, verified, 69,544 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 1 collision hull(s) and extra LOD(s) (UCX_SM_Crate_01_00).This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 69,544 tokens against a median of 69,544.BlenderMifBridge as shipped (Blender)run 5, verified, 86,577 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 1 collision hull(s) and extra LOD(s) (UCX_SM_Crate_01_00).This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 86,577 tokens against a median of 86,577.
A2 modular kit Blender
Build four modular kit pieces on a 100 cm grid: corner pivots, tiling ends, world-scale UVs, lightmaps and hulls that keep the doorway open.
Blender, pilot and full waves. Right results: MifBridge 5/5, Python 5/5. Median run: MifBridge 38,518, Python 36,662, 1.051x.
BlenderPlain Python (Blender)run 1, verified, 36,662 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 6 collision hull(s) and extra LOD(s) (UCX_SM_Floor_400x400_01_00, UCX_SM_Pillar_40x300_01_00, UCX_SM_WallDoor_400x300_01_00, UCX_SM_WallDoor_400x300_01_01, UCX_SM_WallDoor_400x300_01_02, UCX_SM_Wall_400x300_01_00).This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 36,662 tokens against a median of 36,662.BlenderMifBridge as shipped (Blender)run 2, verified, 38,518 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 6 collision hull(s) and extra LOD(s) (UCX_SM_Floor_400x400_01_00, UCX_SM_Pillar_40x300_01_00, UCX_SM_WallDoor_400x300_01_00, UCX_SM_WallDoor_400x300_01_01, UCX_SM_WallDoor_400x300_01_02, UCX_SM_Wall_400x300_01_00).This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 38,518 tokens against a median of 38,518.
A3 organic: rock Blender
Make one 180 x 120 x 90 cm boulder of 3k-6k triangles with a seated pivot, two UV maps and one tight convex hull.
Blender, pilot and full waves. Right results: MifBridge 4/5, Python 3/5. Median run: MifBridge 98,732, Python 95,290, 1.036x. Python off: 5/5 right, median 287,287, 2.910x of as shipped.
BlenderPlain Python (Blender)run 2, verified, 106,186 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 1 collision hull(s) and extra LOD(s) (UCX_SM_Rock_Boulder_01_00).This arm’s cell: five runs: 3 verified, 1 wrong, 1 no measurement. Shown: verified, nearest the cell's median: 106,186 tokens against a median of 95,290.BlenderMifBridge as shipped (Blender)run 2, verified, 98,732 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 1 collision hull(s) and extra LOD(s) (UCX_SM_Rock_Boulder_01_00).This arm’s cell: five runs: 4 verified, 1 wrong. Shown: verified, nearest the cell's median: 98,732 tokens against a median of 98,732.BlenderMifBridge, Python off (Blender)run 4, verified, 287,287 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 1 collision hull(s) and extra LOD(s) (UCX_SM_Rock_Boulder_01_00).This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 287,287 tokens against a median of 287,287.
A4 organic: rock variants Blender
Make three different boulders of one family sharing one material, each with its own pivot and hull.
Blender, pilot and full waves. Right results: MifBridge 5/5, Python 5/5. Median run: MifBridge 103,023, Python 132,040, 0.780x.
BlenderPlain Python (Blender)run 5, verified, 132,040 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 3 collision hull(s) and extra LOD(s) (UCX_SM_Rock_Var_01_00, UCX_SM_Rock_Var_02_00, UCX_SM_Rock_Var_03_00).This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 132,040 tokens against a median of 132,040.BlenderMifBridge as shipped (Blender)run 5, verified, 103,023 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 3 collision hull(s) and extra LOD(s) (UCX_SM_Rock_Var_01_00, UCX_SM_Rock_Var_02_00, UCX_SM_Rock_Var_03_00).This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 103,023 tokens against a median of 103,023.
A5 organic: plants Blender
Model a grass tuft and a fern in real geometry (no alpha cards), with per-blade UVs and a wind weight in a color attribute.
Blender, pilot and full waves. Right results: MifBridge 5/5, Python 5/5. Median run: MifBridge 57,116, Python 36,731, 1.555x. Python off: 5/5 right, median 280,954, 4.919x of as shipped.
BlenderPlain Python (Blender)run 2, verified, 36,731 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm.This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 36,731 tokens against a median of 36,731.BlenderMifBridge as shipped (Blender)run 2, verified, 57,116 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm.This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 57,116 tokens against a median of 57,116.BlenderMifBridge, Python off (Blender)run 2, verified, 280,954 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm.This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 280,954 tokens against a median of 280,954.
A6 LODs and export Blender
Make three LODs of a statue at 50, 25 and 12.5% of its triangles, grouped for FBX, with a hull, exported as one file.
Blender, pilot and full waves. Right results: MifBridge 5/5, Python 5/5. Median run: MifBridge 53,886, Python 50,859, 1.060x.
BlenderPlain Python (Blender)run 2, verified, 50,859 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 4 collision hull(s) and extra LOD(s) (SM_Statue_LOD1, SM_Statue_LOD2, SM_Statue_LOD3, UCX_SM_Statue_00).This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 50,859 tokens against a median of 50,859.BlenderMifBridge as shipped (Blender)run 2, verified, 53,886 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 4 collision hull(s) and extra LOD(s) (SM_Statue_LOD1, SM_Statue_LOD2, SM_Statue_LOD3, UCX_SM_Statue_00).This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 53,886 tokens against a median of 53,886.
A7 vertex-color masks Blender
Write byte-exact masks into a color attribute: a value per loose part, an up-facing mask and a height ramp.
Blender, pilot and full waves. Right results: MifBridge 5/5, Python 5/5. Median run: MifBridge 19,490, Python 21,070, 0.925x.
BlenderPlain Python (Blender)run 4, verified, 21,070 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm.This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 21,070 tokens against a median of 21,070.BlenderMifBridge as shipped (Blender)run 4, verified, 19,490 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm.This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 19,490 tokens against a median of 19,490.
A8 organic: tree Blender
Model a 5 m young broadleaf tree with 400+ real leaves, bark and leaf slots, a wind weight and a trunk-only hull.
Blender, pilot and full waves. Right results: MifBridge 5/5, Python 4/5. Median run: MifBridge 85,907, Python 81,005, 1.061x. Python off: 5/5 right, median 266,867, 3.106x of as shipped.
BlenderPlain Python (Blender)run 5, verified, 81,005 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 1 collision hull(s) and extra LOD(s) (UCX_SM_Tree_01_00).This arm’s cell: five runs: 4 verified, 1 wrong. Shown: verified, nearest the cell's median: 81,005 tokens against a median of 81,005.BlenderMifBridge as shipped (Blender)run 3, verified, 85,907 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 1 collision hull(s) and extra LOD(s) (UCX_SM_Tree_01_00).This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 85,907 tokens against a median of 85,907.BlenderMifBridge, Python off (Blender)run 3, verified, 266,867 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 1 collision hull(s) and extra LOD(s) (UCX_SM_Tree_01_00).This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 266,867 tokens against a median of 266,867.
X1 cross-tool: prop to Unreal Blender + Unreal
Deliver a boulder into Unreal as a saved static mesh: size and pivot, Nanite or LODs, hugging hulls, the material slot and lightmap index, placed.
Blender + Unreal, pilot and full waves. Right results: MifBridge 3/5, Python 3/5. Median run: MifBridge 664,950, Python 201,973, 3.292x. Python off: 3/5 right, median 453,426, 0.682x of as shipped.
In Blender: the scene each arm made
BlenderPlain Python (Blender + Unreal)run 2, verified, 96,167 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 3 collision hull(s) and extra LOD(s) (UCX_SM_Boulder_01_00, UCX_SM_Boulder_01_01, UCX_SM_Boulder_01_02).This arm’s cell: five runs: 3 verified, 2 wrong. Shown: verified, nearest the cell's median: 96,167 tokens against a median of 201,973.BlenderMifBridge as shipped (Blender + Unreal)run 5, verified, 664,950 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 1 collision hull(s) and extra LOD(s) (UCX_SM_Boulder_01_00).This arm’s cell: five runs: 3 verified, 2 wrong. Shown: verified, nearest the cell's median: 664,950 tokens against a median of 664,950.BlenderMifBridge, Python off (Blender + Unreal)run 3, verified, 453,426 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 1 collision hull(s) and extra LOD(s) (UCX_SM_Boulder_01_00).This arm’s cell: five runs: 3 verified, 2 wrong. Shown: verified, nearest the cell's median: 453,426 tokens against a median of 453,426.
In Unreal: what each arm delivered
UnrealPlain Python (Blender + Unreal)run 2, verified, 96,167 tokensWhat this run delivered to Unreal: its saved assets put back into the bench editor from the run's own content, in the run's own level, and captured through the editor viewport. For the picture: a sun, a sky light, a sky atmosphere and a floor plane added for the photograph only; nothing was saved.This arm’s cell: five runs: 3 verified, 2 wrong. Shown: verified, nearest the cell's median: 96,167 tokens against a median of 201,973.UnrealMifBridge as shipped (Blender + Unreal)run 5, verified, 664,950 tokensWhat this run delivered to Unreal: its saved assets put back into the bench editor from the run's own content, in the run's own level, and captured through the editor viewport. For the picture: a sun, a sky light, a sky atmosphere and a floor plane added for the photograph only; nothing was saved.This arm’s cell: five runs: 3 verified, 2 wrong. Shown: verified, nearest the cell's median: 664,950 tokens against a median of 664,950.UnrealMifBridge, Python off (Blender + Unreal)run 3, verified, 453,426 tokensWhat this run delivered to Unreal: its saved assets put back into the bench editor from the run's own content, in the run's own level, and captured through the editor viewport. For the picture: a sun, a sky light, a sky atmosphere and a floor plane added for the photograph only; nothing was saved.This arm’s cell: five runs: 3 verified, 2 wrong. Shown: verified, nearest the cell's median: 453,426 tokens against a median of 453,426.
X2 cross-tool: kit to Unreal Blender + Unreal
Deliver three kit pieces into Unreal as saved static meshes with corner pivots, simple collision that keeps the doorway open, slots and lightmap index.
Blender + Unreal, pilot and full waves. Right results: MifBridge 4/5, Python 5/5. Median run: MifBridge 111,614, Python 180,118, 0.620x.
In Blender: the scene each arm made
BlenderPlain Python (Blender + Unreal)run 3, verified, 180,118 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 7 collision hull(s) and extra LOD(s) (UCX_SM_Floor_400x400_01_00, UCX_SM_Pillar_40x300_01_00, UCX_SM_Pillar_40x300_01_01, UCX_SM_Pillar_40x300_01_02, UCX_SM_WallDoor_400x300_01_00, UCX_SM_WallDoor_400x300_01_01, UCX_SM_WallDoor_400x300_01_02).This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 180,118 tokens against a median of 180,118.BlenderMifBridge as shipped (Blender + Unreal)run 2, verified, 111,614 tokensThe scene this run saved, rendered with EEVEE under the same light and camera as the other arm. For the picture: hidden from the picture: 5 collision hull(s) and extra LOD(s) (UCX_SM_Floor_400x400_01_00, UCX_SM_Pillar_40x300_01_00, UCX_SM_WallDoor_400x300_01_00, UCX_SM_WallDoor_400x300_01_01, UCX_SM_WallDoor_400x300_01_02).This arm’s cell: five runs: 4 verified, 1 crashed. Shown: verified, nearest the cell's median: 111,614 tokens against a median of 111,614.
In Unreal: what each arm delivered
UnrealPlain Python (Blender + Unreal)run 3, verified, 180,118 tokensWhat this run delivered to Unreal: its saved assets put back into the bench editor from the run's own content, in the run's own level, and captured through the editor viewport. For the picture: a sun, a sky light, a sky atmosphere and a floor plane added for the photograph only; nothing was saved; no placement was asked for, so the delivered meshes were placed in a row for the photograph (SM_WallDoor_400x300_01, SM_Floor_400x400_01, SM_Pillar_40x300_01).This arm’s cell: five runs: 5 verified. Shown: verified, nearest the cell's median: 180,118 tokens against a median of 180,118.UnrealMifBridge as shipped (Blender + Unreal)run 2, verified, 111,614 tokensWhat this run delivered to Unreal: its saved assets put back into the bench editor from the run's own content, in the run's own level, and captured through the editor viewport. For the picture: a sun, a sky light, a sky atmosphere and a floor plane added for the photograph only; nothing was saved; no placement was asked for, so the delivered meshes were placed in a row for the photograph (SM_WallDoor_400x300_01, SM_Floor_400x400_01, SM_Pillar_40x300_01).This arm’s cell: five runs: 4 verified, 1 crashed. Shown: verified, nearest the cell's median: 111,614 tokens against a median of 111,614.
Every task, every outcome
Five runs an arm on each task, by wave. The four outcomes are counted apart, with runs that could not be graded in a column of their own. Tokens are per run.
Task
Arm
Outcomes
Median
Mean
Tokens per run, range
Against Python, median / mean
A1 game-ready passBlender
Plain Python (Blender)
5 verified
69,544
68,447
middle half60,398 to 72,331all runs58,793 to 81,167
MifBridge as shipped (Blender)
5 verified
86,577
96,797
middle half85,692 to 87,301all runs85,016 to 139,399
1.245x / 1.414x
A2 modular kitBlender
Plain Python (Blender)
5 verified
36,662
36,825
middle half36,114 to 36,977all runs34,952 to 39,418
MifBridge as shipped (Blender)
5 verified
38,518
36,731
middle half37,370 to 39,206all runs28,362 to 40,200
1.051x / 0.997x
A3 organic: rockBlender
Plain Python (Blender)
3 verified1 wrong1 no measurement
95,290
95,665
middle half81,302 to 109,654all runs72,024 to 120,057
MifBridge as shipped (Blender)
4 verified1 wrong
98,732
96,467
middle half69,842 to 102,963all runs63,108 to 147,691
1.036x / 1.008x
MifBridge, Python off (Blender)
5 verified
287,287
293,977
middle half232,863 to 334,451all runs228,026 to 387,259
3.015x / 3.073x
A4 organic: rock variantsBlender
Plain Python (Blender)
5 verified
132,040
134,782
middle half131,251 to 141,468all runs126,110 to 143,040
MifBridge as shipped (Blender)
5 verified
103,023
90,528
middle half64,894 to 111,654all runs61,197 to 111,872
0.780x / 0.672x
A5 organic: plantsBlender
Plain Python (Blender)
5 verified
36,731
39,030
middle half33,771 to 41,700all runs33,537 to 49,413
MifBridge as shipped (Blender)
5 verified
57,116
60,795
middle half37,648 to 75,783all runs37,474 to 95,954
1.555x / 1.558x
MifBridge, Python off (Blender)
5 verified
280,954
287,848
middle half260,064 to 313,984all runs255,847 to 328,391
7.649x / 7.375x
A6 LODs and exportBlender
Plain Python (Blender)
5 verified
50,859
59,110
middle half49,657 to 55,000all runs48,085 to 91,950
MifBridge as shipped (Blender)
5 verified
53,886
54,698
middle half52,154 to 57,673all runs48,566 to 61,209
1.060x / 0.925x
A7 vertex-color masksBlender
Plain Python (Blender)
5 verified
21,070
19,275
middle half16,076 to 21,202all runs15,684 to 22,345
MifBridge as shipped (Blender)
5 verified
19,490
22,792
middle half18,402 to 24,237all runs18,170 to 33,662
0.925x / 1.182x
A8 organic: treeBlender
Plain Python (Blender)
4 verified1 wrong
81,005
78,840
middle half79,173 to 85,046all runs63,315 to 85,661
MifBridge as shipped (Blender)
5 verified
85,907
79,905
middle half65,107 to 93,105all runs50,747 to 104,657
1.061x / 1.014x
MifBridge, Python off (Blender)
5 verified
266,867
381,367
middle half266,591 to 297,882all runs258,809 to 816,684
3.294x / 4.837x
X1 cross-tool: prop to UnrealBlender + Unreal
Plain Python (Blender + Unreal)
3 verified2 wrong
201,973
255,296
middle half96,167 to 387,486all runs82,902 to 507,954
MifBridge as shipped (Blender + Unreal)
3 verified2 wrong
664,950
559,303
middle half260,619 to 846,237all runs148,196 to 876,512
3.292x / 2.191x
MifBridge, Python off (Blender + Unreal)
3 verified2 wrong
453,426
502,084
middle half237,017 to 495,549all runs216,605 to 1,107,821
2.245x / 1.967x
X2 cross-tool: kit to UnrealBlender + Unreal
Plain Python (Blender + Unreal)
5 verified
180,118
193,840
middle half94,648 to 209,537all runs81,150 to 403,748
MifBridge as shipped (Blender + Unreal)
4 verified1 crashed
111,614
115,460
middle half110,315 to 120,573all runs93,791 to 141,005
0.620x / 0.596x
Verified
the scene is correct.
Wrong
it ran and the scene is not.
Refused
the tool declined (not a failure and not a success).
Crashed
the application stopped answering.
No measurement
the harness could not grade the run.
Who grades
A third process grades every run and talks to neither arm.
Tokens, time and dollars
Arm
Runs
Median run
Mean run
Tokens per run, range
Model time
Dollars, at API list prices
Plain Python
50
72,331
98,161
middle half39,418 to 106,186all runs15,684 to 507,954
4,239 s
$13.29
MifBridge as shipped
50
80,400
121,348
middle half49,111 to 108,901all runs18,170 to 876,512
4,827 s
$15.33
MifBridge, Python off A3, A5, A8 and X1 only
20
284,121
366,319
middle half258,069 to 347,653all runs216,605 to 1,107,821
1,675 s
$8.76
Tokens
Tokens are input + output + cache read + cache write, all counted at full weight, as the API reports them per run: the context processed, largely cache reads, not a bill.
Range
Middle half: where the typical runs fall, leaving out the cheapest quarter and the costliest quarter. All runs: the cheapest run to the costliest.
Dollars
Dollars are the round's tokens priced at the API's list prices for claude-opus-5-5 (input $4, output $20, cache read $0.20 and cache write $8 a million tokens): what an API-key user would pay. On a claude.ai plan the same runs draw on the plan's usage allowance instead of being billed.
How this was measured
The method
The same model got the same task text and the same starting scene twice, once through each arm. Only the tools changed. Each run is one fresh session of Claude Code that ends when the model says it is done; then a separate process opens what the run saved and checks it.
The applications
A1 to A8 run in Blender; X1 and X2 build in Blender and deliver into Unreal Engine. Every arm drives the same application on a given task, and each label on this page names it: Plain Python (Blender) on A1 to A8 and Plain Python (Blender + Unreal) on X1 and X2, and the same for every other arm.
The arms
Plain Python: one tool that runs the model's Python inside the editor, exactly as written. It ran A1 to A8 in Blender; X1 and X2 in Blender + Unreal.MifBridge as shipped: MifBridge as a buyer installs it: its default tools, Claude Code's tool search, and its own Python tool left on as shipped. It ran A1 to A8 in Blender; X1 and X2 in Blender + Unreal.MifBridge, Python off: the same install with the Blender add-on's own 'Allow run_python' preference unticked, so Blender work goes through the typed tools (round ten's diagnostic arm; Unreal's Python stays reachable, MifBridge has no switch for it). It ran A3, A5 and A8 in Blender; X1 in Blender + Unreal.The build: MifBridge's development version at commit 2a026a58 (2026-10-01), 55 commits after the 1.1.0 release and before 1.2. It is not the build on sale: that is 1.1.0 (4684a30f), which round nine measured. Every MifBridge run in this round was checked to be on 2a026a58.
Model and client
claude-opus-5-5 through Claude Code 2.1.280, at the client’s default effort (medium), recorded for every run that finished (a run stopped at the client’s time limit records neither), started from an empty folder so no project files or settings reach the model.
Grading
A third process grades every run and talks to neither arm. For Blender it is a fresh headless Blender that opens the saved scene; for Unreal Engine, a separate headless editor that loads the saved level and assets; for X1 and X2, both: the Unreal editor grades the saved asset and exports it, and a fresh Blender grades that export. It runs the same checks for every arm, and every checker was first tested against a correct build and against deliberately broken ones. A tool reporting success is never taken as the evidence.
What counts as a loss
A task where MifBridge’s median run used more tokens than plain Python’s median run. It is listed as a loss whether or not either arm got the task right, so a task MifBridge got right and plain Python got wrong can still be a loss on tokens (A3 and A8).
Tokens and dollars
Tokens are input + output + cache read + cache write, all counted at full weight, as the API reports them per run: the context processed, largely cache reads, not a bill. Dollars are those tokens priced at API list prices, what an API-key user would pay; a Claude plan user is not billed them.
When
, 120 runs, in two waves; the table below lists them.
The pictures
24 renders of saved Blender scenes (EEVEE, one camera and light per task for every arm, with any staging stated under the picture) and 5 photographs of what a run delivered to Unreal, put back into the editor from the run’s own saved content. Every task has its pictures from each arm.
Deliver three kit pieces into Unreal as saved static meshes with corner pivots, simple collision that keeps the doorway open, slots and lightmap index.
Download the data (JSON): every number on this page, as MifBench’s exporter wrote it from the run records on 2026-10-01.
What this round cannot show
One model (claude-opus-5-5) through one client (Claude Code), at its default effort, on one machine.
UE 5.8.2 and Blender 5.2 only: a claim about Python on UE 5.3-5.7 needs a bench project on that engine, and 5.8 added Blueprint graph APIs the earlier engines lack.
Every task is a small, fully specified job: none asks the agent to find its way around a large existing scene it did not make.
Three to five runs a cell: cells are wide, so a per-task ratio is a rough guide.
Ratios to plain Python are read within one round only; plain Python's own cost swings between rounds.
Every round ran at the CLI's default effort (medium for claude-opus-5-5 on Claude Code 2.1.280; round nine records it per run). Rounds one to eight passed a desktop session's variables to the CLI, which adds ~208 tokens of client prompt to every request on every arm alike: measured against rounds six to eight's own probes; round nine removes them.
Whether the assets look good: the look proxies only rule out clean primitives (a jittered smooth blob passes the rock proxies); the rendered pictures are for a person to judge, and nothing here scores them.
Knowledge: every prompt states the conventions it grades (UCX_ names, fbx_type, the byte rule, padding), so the round measures carrying out a spec, not knowing the conventions unprompted.
Textures, baking and Unreal-side use (lighting, streaming, LOD switching, Nanite's look) are not measured.
The Python-off arm turns off Blender's Python only: MifBridge has no switch for Unreal's (run_python app='unreal' goes through exec_console), so on X1 that arm could still script Unreal.