TL;DR: Two local models, one RTX 4090 (24 GB), 30+ hidden-graded agentic runs through a real coding harness. The 27B on NInfer scored 64/64 on our daily-driver suite at ~150 t/s decode. Flash-Next on Strata scored 63/64 at ~63 t/s — but was the only one that got every hidden edge case in the hardest coding test. Speed and correctness-under-pressure pull in different directions, and the numbers show exactly how far.
Why another benchmark
We already knew both models could code. What we didn't know was how they behave as agents: calling tools, reading their own failures, driving a browser, clicking through a GUI, reading images, and editing real production workflow files — the things we actually use them for every day. Leaderboard numbers don't answer that, and neither does eyeballing a transcript.
So we built a harness that does. Every run goes through QuetzaCodetl (our fork of an agentic coding CLI) headless, with the exact environment each model's launcher uses, in a fresh working directory. When the agent says it's finished, a hidden grader the model never sees checks the result. Every grader was proven first against a reference solution (it must score 100%) and against a cheat (it must catch it). One model on the GPU at a time, same prompts for both.
| Contender | What it is | Engine |
|---|---|---|
| Flash-Next | Swift 1.5 Qwen3.8-Flash-Next — 177B MoE, 512 experts, IQ3_XXS, uncensored via a refusal-direction control vector | Strata (experts streamed from RAM, hot ones cached in VRAM) |
| 27B | Swift 1.5 Qwen3.8-27B Uncensored — dense, Q4/Q5, MTP speculative decoding + n-gram drafts | NInfer (purpose-built CUDA engine, 200K context) |
Round 1: write-it-yourself interpreters
Two "super hard" coding tasks. A JSON parser from scratch (no json, re or eval), graded on 40 valid and 46 invalid inputs with type-exact comparison. And a mini Scheme interpreter with closures, mutual recursion and recursion 3,000 deep — 25 spec cases the model sees, plus 35 hidden robustness cases it doesn't.
| Flash-Next | 27B | |
|---|---|---|
| JSON parser | 86/86 · green in 4m25s · 5 requests | 86/86 · green in 2m00s · 26 requests |
| Scheme — spec cases | 25/25 | 25/25 |
| Scheme — hidden cases | 35/35 | 32/35 |
| Scheme — time to green | 14m45s | 7m01s |
| Decode (median) | ~72 t/s | ~122 t/s |
The three hidden Scheme failures are the interesting part. The 27B built a trampolined evaluator — correct for tail calls, so (countdown 100000) passes — but applied built-ins by calling the continuation directly instead of returning it to the loop. Any non-tail recursion ((sum 3000)) quietly rebuilt Python's call stack and hit the recursion limit. Flash-Next took longer, landed more bugs on its first write, debugged them with small targeted probes, and ended up with a genuine CEK machine that handles all of it. On JSON it wrote the parser and the tests, ran them once, and was done.
Round 2: the daily-driver suite
Six tasks modelled on real work:
- Vision — 15 questions over 12 images: chart values, a table sum, shape counts, tracing code in a screenshot, an analog clock, rotated text, a flowchart, UI toggle states, a dense receipt.
- Browser — a login with CSRF, a cookie wall, a JS-paginated orders table, flagging exactly the right four orders among near-miss distractors ($500.00 exactly, $499.99, "Delay review"), and finding a reference number buried in a note.
- Desktop — driving a GTK settings form through a desktop-automation CLI on an isolated X display.
- ComfyUI read / modify / write — answering questions about real production video workflows, making five surgical edits (including inserting a LoRA node and keeping the link graph consistent), and writing a Flux-2 workflow from scratch validated against a live 4,454-node schema.
| Arm | Score | Tasks passed | Total time | Decode (median) |
|---|---|---|---|---|
| 27B · NInfer | 64/64 | 6/6 | 9m18s | 152 t/s |
| 27B · low thinking (4K budget) | 62/64 | 5/6 | 8m17s | 162 t/s |
| Flash-Next · Strata | 63/64 | 5/6 | 17m05s | 63 t/s |
| Conductor pipeline (below) | 32/32 | 3/3 run | 29m30s | — |
The 27B's run is about as clean as a benchmark gets: every check on every task, averaging 93 seconds a task. Low thinking saved almost no time and cost the hardest task (it forgot that Flux 2's negative has to be a zeroed-out copy of the positive). Flash-Next's one miss was leaving a checkbox ticked on the desktop form.
Throughput, measured per request
| 27B · NInfer | Flash-Next · Strata | |
|---|---|---|
| Decode, median (p10–p90) | 152 t/s (135–156) | 63 t/s (58–83) |
| Prefill, uncached tokens | ~2.7K t/s (4.4K peak) | ~1.1K t/s (2.5K peak) |
| Time to first token, median | 225 ms | 820 ms |
| Prompt-cache hit | 96–99% | 99% |
| Speculation | MTP ~52% + n-gram ~41% accepted | MTP; expert-cache hit ~91% |
Agentic sessions are prefill-heavy — every turn resends the whole conversation — so the prompt cache does the real work: 96–99% of prompt tokens never get recomputed on either engine. The 27B's advantage is everywhere else: 2.4× the decode, 2.5× the prefill, a quarter of the latency to first token.
Can the big model plan and the fast one execute?
With only one GPU, the two models can't run side by side — but they can take turns. We tested a conductor pipeline: Flash-Next investigates and writes a plan, the engine swaps to the 27B to execute it, Flash-Next swaps back in to review the result, and the 27B fixes anything flagged. Model swaps cost ~24 s for NInfer and ~60 s for Strata.
It worked — every task it ran passed, the plans were thorough (80–170 lines of verified commands), and on one ComfyUI task the reviewer caught a real mistake the executor made. But it was ~6× slower than the 27B alone, and on tasks the 27B already aces there's nothing left to improve. The pattern earns its keep on work where the fast model is subtly wrong — exactly the Scheme edge cases above — not as a default.
It also produced the best cautionary tale of the week. At that point Flash-Next's vision was silently broken (next section), so as reviewer it confidently told the 27B to change nine correct vision answers, "verified by reading each image directly". The 27B re-checked the images itself, kept its answers, and scored 15/15. A good executor has to be willing to disagree with a confident reviewer.
What the benchmark broke (and we fixed)
Running real agent loops at volume shakes out bugs no synthetic test finds. Four of them, all now fixed on our side:
- Strata was blind in agent sessions. Its Anthropic-API converter flattened tool results to text, silently dropping images — and an agent's file-read tool returns images inside tool results. Plain image messages worked, so it looked fine in a quick test. Flash-Next scored 0/15 on vision until we patched the converter; afterwards, 15/15.
- Strata's vision encoder ran out of VRAM on its first image. Raising the VRAM reserve from 700 MB to 1 GB fixed it, and raising the image-token budget from 1,024 to 6,144 took a full 3440×1440 screenshot from unreadable (0/3 on tiny text) to native resolution. The budget itself costs nothing, but keeping the encoder resident costs ~15% of the GPU expert cache and ~25–35% of decode speed, so we switch vision on only when we need it.
- NInfer wedged on multi-turn image sessions. Follow-up turns with images in cached history died with "retained materialization source is unavailable" and took the engine down until restart. The fix was capturing fewer cached states per turn (one more state slot, no automatic long anchors, fewer shared prefixes). Zero failures since.
- Our desktop-automation tool returned stale screenshots — newer
scrotrefuses to overwrite a file, so every "fresh" screen analysis silently reused the first one. One flag.
The verdict
The 27B is the daily driver. On the work we actually do — browser, desktop, vision, ComfyUI, tool-heavy coding — it scored everything, 2–2.5× faster, on a single 24 GB card with 200K context. Flash-Next is the specialist: slower, but the only model that got every hidden edge case right in the hardest coding test, and a thorough planner and reviewer. It's worth reaching for when being right in the corners matters more than turnaround.
The broader lesson is about measurement. Both models aced every test they could see. The differences only showed up in hidden cases, in per-request throughput, and in bugs that only appear when an agent does the same thing a hundred times. If you're choosing a local model for agent work, build the harness first — it takes an afternoon, and it tells you things no leaderboard will.
