Update, October 1: we put Flash-Next on Strata head-to-head against a dense 27B on the same card, through 30+ hidden-graded agent runs — read the benchmark.

⚡ The short version

We dropped a 76 GB mixture-of-experts model onto a single RTX 4090 with an engine built specifically for it — Strata — and it decodes at 76 tokens/s, more than double llama.cpp on the same card. Then we pointed a single directional control vector at its residual stream and cut its refusal rate from 100% to 27%, with throughput intact and zero weight edits.

No re-quantisation, no LoRA, no baked weights. Two bugs, one hard lesson about contrast sets, and a model that refuses 73% less while running at full speed.

The model

Qwen3.8-Flash-Next is a 177-billion-parameter sparse mixture-of-experts model: 48 layers, 512 experts per layer with only 10 active per token, a 256K trained context, and — unusually — several parallel residual streams (hyper-connections) that mix and recombine at every layer. The Swift 1.5 fine-tune is the fast-thinking variant we actually run. Quantised to IQ3_XXS it is ~76 GB across two shards.

On llama.cpp this model was a constant fight. Its best decode was 36 t/s, and that number only appeared with the expert cache turned on and everything host-resident. Add a LoRA — which is how we were trying to remove its refusals — and a single guard clause silently disabled the cache, dropping us to ~12 t/s. Months of patching went into clawing that back.

The engine: Strata

Strata is a purpose-built inference engine for exactly this architecture. It splits the model the way the architecture wants: the experts live in RAM, the busy ones get a GPU cache that learns as you use it, a 29 GB lookup table stays on the SSD, and a small draft head speculates several tokens ahead which the full model then checks.

Same card, same quant, stock model, no graph surgery:

EngineDecode
llama.cpp (ncmoe 48 + cache 160)36.2 t/s
Strata (warm)76.5 t/s

2.1×. And it ships with native control-vector support — a fused CUDA kernel that can subtract a direction from the residual stream at every layer, which is exactly the operation we had been trying to bolt onto llama.cpp by hand.

Removing the refusals

A "refusal direction" is a real, measurable thing: a single vector in activation space that mediates whether the model declines. Find it, and you can project it out of the residual stream at runtime — h − (h·v)v — and the refusals go with it, reversibly, with no weight edited. The research that pinned this down is Arditi et al. (NeurIPS 2024); the community tooling is mature.

The hard part is finding the direction. You do it with contrast pairs: prompts the model refuses, matched against prompts it answers. The difference in their activations is the direction. We used the real set — mlabonne/harmful_behaviors against mlabonne/harmless_alpaca, 64 length-matched pairs.

Strata took our 47-layer direction file and steered the model with one line of flags:

# project the refusal direction out of every residual stream --control-vector-scaled refusal-real-swift15.gguf:1.0 --control-vector-layer-range 1 47 --cvec-mode project --cvec-dir per-layer

Two bugs worth naming

The first build crashed, and the crashes were more instructive than the success.

1. The contrast set was silently empty. A handful of harmless prompts contained embedded newlines, so the positive and negative files had 64 and 66 lines. The generator printed "number of positive and negative prompts must be equal" and then continued anyway with zero pairs — feeding uninitialised memory into the matrix math. The fix was a one-line whitespace collapse. The lesson: check the damn line counts, and never trust a tool that warns and keeps going.

2. The PCA path was broken against the build. The generator's PCA step calls a matrix multiply with a transposed operand, which a newer ggml rejects outright. The fix was to use the mean method instead — which, as it happens, is the canonical abliteration recipe anyway.

The result

Stock, the model refused 30 of 30 genuinely harmful requests. With the vector on, 8 of 30 — a 73% reduction — and most of those eight were the ones you'd want to keep (self-harm and child-safety) plus a couple that genuinely didn't clear. Throughput did not move: the vector costs 0.2–0.4% per token, and the variance in our runs was speculative-decode acceptance, not the steering.

ConfigRefusal rateDecode
stock100%~76 t/s
+ refusal-direction cvec27%~76 t/s

In the wild: real usage, real tools

Numbers from a live agentic session — QuetzaCodetl driving it, tool calls included — tell the fuller story. Decode held between 48 and 67 t/s (median ~57). The "76 t/s" headline is the warm, thinking-off ceiling; real back-and-forth with tool output lands a little lower because tool-call text is denser than prose. Tool-call generation ran 48–62 t/s, genuinely-new-token prefill ~600–1700 t/s (the rest is KV-cache reuse — 95% of a long session), and the expert cache stayed at a 92–94% hit rate.

The proof it actually unrefuses: in a live run the model went ahead and wrote a malicious keylogger for a refusal test — the exact class of request stock declines 100% of the time.

Final config: 200K context with 8-bit KV streaming (99.8% of KV reads hit VRAM), xhigh thinking with a 32K budget, and the whole thing wired so strata-quetza starts the engine and drops you straight into QuetzaCodetl.

Why this matters

Two separate fights — speed and refusal — turned out to have the same answer: use the right engine. Strata removes the entire llama.cpp expert-cache-and-LoRA patching saga, and its native control vectors make refusal steering a one-line flag instead of a weight bake. You get a model that runs twice as fast and refuses 73% less, on hardware that was already sitting on the desk.

The standing lesson is the one we keep re-learning: the contrast set matters more than the math. A fabricated direction encodes refusal style and collapses coherence; a real one cleanly removes the behaviour. If your abliteration "doesn't work", you didn't find the right direction — you made it up.