TL;DR
Over five weeks we took Qwen3.8-27B on a single Tenstorrent Blackhole P150 (32 GB GDDR6) from a generic bring-up running at 7.7 tok/s to a native mixed BFP4/BFP8 build that decodes at ~50–55 tok/s single-stream (greedy, DFlash2 speculative decoding), prefills at ~1,200 tok/s, and serves the model's full 262,144-token native context on one card.
| Milestone | Single-stream decode | Notes |
|---|---|---|
| Generic baseline (Sep 7) | 7.7 tok/s | BFP4 MLP / BFP8 rest, no fast paths, MTP off |
| Qwen 9B optimization stack ported (Sep 8) | 15.6 tok/s → 23.2 tok/s with MTP-1 | Native GDN kernels, packed BF4/BF8 readers, L1-resident state |
| Native DFlash2 + dual-NoC (Sep 19) | 36–43 tok/s | Workload-dependent, HTTP |
Round-two fusion work, dflash5 (Sep 27) | 49.7 / 49.0 tok/s at 128 / 2K prompts | Bit-exact vs. reference runtime |
Current default, ctx5 (Oct 3) | 50.2 / 54.6 tok/s at 128 / 2K prompts | GPTQ checkpoint, prefix caching, speculative output identical to plain decoding |
The weights are public: Lottolabs/Qwen3.8-27B-TT-Mixed-BFP4-BFP8-P150.
All decode numbers are single-stream, greedy (temperature 0) and measured over the OpenAI-compatible HTTP streaming API unless stated otherwise. Speculative decoding speed depends on how predictable the text is, so different prompts give different numbers. Each table says which prompt it used.
Hardware and software
- Accelerator: 1× Tenstorrent P150 (Blackhole), 32 GB GDDR6, 512 GB/s rated, 11×10 = 110 usable worker cores, ~1.5 MB L1 SRAM per core, AI clock 1,350 MHz.
- Host: AMD Ryzen 9 9950X, ~90 GB RAM, Ubuntu 26.04, TT-KMD 2.11.0, firmware bundle 19.15.0.
- Runtime: TT-Metal/TTNN with the Tenstorrent vLLM plugin. Our own model code is a fork of the
qwen36Blackhole demo, extended for 27B geometry. - Model:
Qwen/Qwen3.8-27B, a dense hybrid with 64 layers: 48 Gated DeltaNet (GDN) + 16 full attention, hidden 5,120, MLP 17,408, a 248K vocabulary and one MTP layer.
How we got here
1. The card, and a 9B rehearsal (Aug 30 – Sep 7)
The P150 went into a PCIe 4.0 x16 slot on Aug 30. First light was Falcon3-7B (BFP8) at 19.9 tok/s through Forge/vLLM, which also drove adding P150 and the tt-metal backend to LocalMaxxing.
Qwen3.5-9B on Tenstorrent's native qwen36 runtime then decoded at 19.3 tok/s, with full-vocabulary logits copied to the host for every token. We used the 9B as a rehearsal: on-device argmax, packed BF4/BF8 weight readers, native fused GDN kernels, recurrent state kept in L1, traced decode, MTP and later DFlash. That took the 9B to 36.6 tok/s without MTP and 50.6 tok/s with MTP, released as Lottolabs/Qwen3.5-9B-TT-Mixed-BFP4-BFP8-P150. Everything afterwards was a port of that playbook.
2. Building the 27B quant (Sep 7)
We started from the original BF16 weights, not an existing GGUF/AWQ/GPTQ file, so we could choose precision per tensor before losing any information. Quantizing runs on the host, one tensor at a time, so one P150 is enough: the BF16 model never has to fit on the card.
| Component | Precision |
|---|---|
| MLP gate / up / down | BFP4 (TT block floating point: 16-value blocks with a shared exponent) |
| Attention and GDN projections, LM head | BFP8 |
| Embeddings, norms, small tensors | BF16 / lossless |
| MTP layer | stored losslessly |
| Activations | BF16 |
| GDN recurrent state | BF16 in L1 (FP32 in the original control) |
- Native checkpoint: 22.10 GB. The public package with runtime is ~27.6 GB.
- Dominant weight bytes read per token: 18.68 GB, including block-exponent overhead. "27B × 4 bits ≈ 13.5 GB" is wrong for this mix.
- Device DRAM after load: 20.17 GiB at short context.
- Integrity: 962/962 tensor round-trips pass. A checksummed equivalence proof shows the native checkpoint reproduces the quantized model's full-vocabulary logits bit-for-bit at 95 positions. Every runtime change since then has had to regenerate and pass that proof before serving.
First served baseline: 7.69 tok/s decode, with time-to-first-token (TTFT) of 0.61 s at 128 tokens and 4.31 s at 1,024.
3. Porting the 9B stack to 27B geometry (Sep 7–8)
My first attempt tuned the generic fallback and gained 1.2%. That was the wrong priority. The 9B kernels hard-coded 9B geometry, and 27B breaks several assumptions:
- 48 value heads instead of 32 (the MTP verifier hard-coded 32).
- Convolution state needed 144 cores, but the P150 has 110.
- At 2K-token prefill, two kernel batch limits (512 and 2,048) were silently exceeded, so execution fell back to slow generic paths.
After porting every fast path at 27B shapes (native GDN frontend/recurrence/output, packed BF4/BF8 projections, fused QKV/SDPA, sharded norms, traces, MTP-1 with fused 2-token verify), all 75.5 MB of BF16 GDN state fits in L1: 749,568 B/core with an 84 KB margin.
| Control | Optimized, MTP off | Optimized, MTP-1 | |
|---|---|---|---|
| Decode | 9.17 tok/s | 15.59 | 23.15 |
Prefill fixes. One hard-coded batch cap of 2,048 in the fused Horner solve silently fell back to a 30-iteration Python loop. Fixing it took the GDN layer from 186 ms to 88 ms. The packed weight reader re-streamed the whole weight once per 32 prompt rows; routing packed prefill through a 2D multicast matmul reads each weight slab once. Together: 128-token prefill 59 → 362 tok/s, 1K 301 → 888 tok/s, TTFT on a 55-token prompt 2.26 s → 0.41 s, with no extra DRAM. The 9B solution, a second unpacked weight copy, would have cost ~9 GB.
4. Measuring the ceiling, and lots of dead ends (Sep 8–18)
Batch-1 decode is limited by memory bandwidth:
18.68 GB ÷ 512 GB/s = 36.5 ms per token → at most 27.4 tok/s without speculation
We were at 63 ms/token, roughly 57% of rated bandwidth as useful weight reads. A three-way experiment on the biggest projection (MLP gate/up) settled what limits it:
| Gate/up mode | Latency |
|---|---|
| Actual projection | 253 µs |
| Read weights, skip arithmetic | 253 µs |
| Reuse weights already in L1, keep arithmetic | 132 µs |
Removing arithmetic changed nothing; removing weight reads nearly halved latency. Decode is weight-delivery-bound, not compute-bound.
What worked (all bit-exact):
- Explicit M2 GDN output-projection config: −2.3% MTP cycle time.
- NoC zero-fill in place of CPU tile clearing: −2.0%.
- Exact fused RoPE plus SwiGLU balanced across 110 cores: −1.9%.
- Attention layout fusion (+1.4%), fused M1 recurrence + output norm (+3.5%), bulk recurrent-state DMA (+4.9%): MTP-off decode went from 16.2 to 18.2 tok/s.
- MTP-3 chained verification: 32.95 tok/s pooled in native runs (+18% over MTP-1); MTP-4 lost.
- Dual-NoC weight reads: a validated microbenchmark reached 435 GB/s versus 374 GB/s on one NoC. In serving it gave +7% for MTP and +3–4% for DFlash.
What didn't (measured, then rejected):
- More circular-buffer slots or deeper outstanding reads: always slower.
- Reversing the weight/activation NoC pair: 15% slower.
- A streaming fused FFN: permanent no-go.
- Deeper in0 multicast buffers: monotonically slower.
- Several 0.01–0.3% micro-fusions: below noise.
- BFP4 LM head: +2.6% speed, but top-1 agreement fell to 94.9% and NLL rose by 0.047. Rejected.
- GDN input projections to BFP4: +7.8% speed, but failed fidelity limits (max KL 21). A calibrated 8-layer variant gave +0.4% and still failed.
5. DFlash2 on Blackhole (Sep 13–21)
We ported the full z-lab/Qwen3.8-27B-DFlash2 drafter natively: five draft layers, dynamic convolutions, top-16 candidates and the learned selector. It verifies blocks of 8 tokens on the P150, with exact rollback of GDN state and KV cache at every acceptance index.
| Workload (HTTP) | MTP-1 | DFlash B8 |
|---|---|---|
| Code | 26.7 | 38.8 |
| Reasoning prose | 26.7 | 43.5 |
| 3K-token retrieval | 23.0 | 35.9 |
A two-branch DFlash variant worked but was 4–17% slower. We also confirmed the vLLM HTTP path costs only 0.6% versus native execution on identical token streams. Lower LocalMaxxing numbers come from the prompt, not from serving overhead.
Publication: the quant went to Hugging Face on Sep 20, verified by an anonymous fresh download and relaunch. LocalMaxxing runs:
- MTP-1, official prompt, 256 tokens: 23.8 tok/s.
- DFlash, custom repetitive prompt, 128 tokens: 55.9 tok/s median (61.9 best). This is a peak number, not general chat speed.
6. Round two: making DFlash the fast path (Sep 26–27)
Lessons from our Qwen 35B work came back to 27B, with a rule: every change keeps verified logits bit-identical, and every build regenerates its proof from the original weights.
| Build | 128-token prompt | 2K prompt | What changed |
|---|---|---|---|
| Reference (MTP-1) | 27.7 | 27.9 | — |
mtp3 (MTP-2) | 35.2 | 34.0 | Fused verify attention, batched GDN commits, 65,536-id draft vocab, BFP8 draft weights |
dflash1 | 39.6 | 40.8 | MTP verify fusions ported to the 8-row DFlash verify; BFP8 drafter |
dflash2 | 43.2 | 42.5 | Native accept/argmax (2 launches instead of ~40), cheaper draft ops |
dflash3 | 47.4 | 46.5 | GDN verify batched across all 8 rows (999 → 860 µs/layer) |
dflash4 | 49.2 | 48.3 | Chunked activation multicast; dual-NoC down, GDN-out and attn-out |
dflash5 | 49.7 | 49.0 | Blackhole DRAM prefetcher, which upstream tt-metal disables on Blackhole |
DFlash emits ~3.3 tokens per verify cycle, and the cycle fell from ~117 ms to ~67 ms. The target model's 64 layers take ~56.6 ms per cycle against a ~55 ms weight-read floor at current read rates.
Prefill, same rules (all bit-exact):
dflash6: the 27B batched decay/normalization and convolution-filter kernels finally ran at 27B sizes.dflash7: trace replay for short prompts, which had been paying ~14,600 Python op dispatches per request.
| Before | dflash7 | |
|---|---|---|
| TTFT, 128-token prompt | 0.38 s | 0.21 s |
| TTFT, 2K prompt | 2.53 s | 1.69 s |
| LocalMaxxing official prompt, decode | 24.7 tok/s | 42.5 tok/s |
| LocalMaxxing official prompt, TTFT | 403 ms | 221 ms |
| Peak prefill (2,045-token prompt) | ~880 tok/s | 1,212 tok/s |
A fused GDN prefill kernel was rejected: it was not exact (perplexity 26.05 → 26.27) and clashed with the L1-resident decode state.
7. Native 262K context (Sep 27–28)
The fast path was capped at 8K. Lifting that meant re-validating every kernel built only for 8K, plus fitting the KV cache. At 262K the BFP8 KV cache alone is 9.1 GB: 27B stores 34.8 KB/token versus 10.9 KB for Qwen 35B-A3B.
- Up to 64K: BF16 KV cache, same numerics as the 8K build.
- Above 64K: BFP8 KV cache.
- Up to 128K: DFlash.
- At 262K: MTP-3, because the DFlash drafter no longer fits, and a 4-bit drafter collapsed acceptance.
- New verify attention: splits each head's KV read across cores. At 60K the cycle fell from 145 ms to 84 ms, with identical output.
| Context | Decode (current ctx5) | TTFT |
|---|---|---|
| 2K (8K setup) | 54.6 tok/s | 1.7 s |
| 64K (128K DFlash setup) | 43.1 tok/s | ~70 s |
| 128K (128K DFlash setup) | 36.4 tok/s | ~179 s |
| 262K (MTP-3) | 23.9 tok/s | ~510 s |
Passkeys planted at the start, middle and end are found at 32K, 64K, 128K and 262K, and over-length requests are rejected.
8. Quality, then a better quant (Sep 28 – Oct 3)
Quality checks on the first long-context build (ctx1):
| Check | Result |
|---|---|
| Perplexity vs. original BF16 model (Newton's Opticks, real code) | +1–4%, not growing with context (16K +4%, 32K +2%, code +1%) |
| BFP8 vs. BF16 KV cache, short context | +0.5% perplexity (chat/instructions), 97.1% top-1 agreement |
| BFP8 vs. BF16 KV cache, 16K–128K | within noise |
| gsm8k, 50 questions | 50/50 with either KV precision |
| LocalMaxxing gsm8k shard 1 | 99/101 (98.0%) |
| LocalMaxxing arc-challenge shard 1 | 283/289 (97.9%) |
A finer-grained comparison then showed what "bit-exact checkpoint" doesn't tell you: plain-rounding BFP4 had a real cost. Code perplexity was +4.9% against BF16, and chat KL was 0.180 against an all-BFP8 build. We built a BFP4-aware sequential GPTQ (error-compensated rounding) and re-quantized only the 192 MLP matrices, at the same size:
| KL vs. reference | Plain rounding | GPTQ (balanced calibration) |
|---|---|---|
| Chat | 0.180 | 0.070 |
| Code | 0.053 | 0.047 |
| Overall | 0.103 | 0.055 |
The GPTQ checkpoint is now the default and on Hugging Face. Our first GPTQ run accidentally calibrated on chat only: chat improved 2.7× while code got worse. Calibration data has to cover every domain you serve.
Two more exactness fixes landed by Oct 3:
- Prefix caching: GDN state is snapshotted at 2,048-token chunk boundaries, so the next turn of a 64K conversation starts in 3.8 s instead of 71 s (128K: 5.3 s instead of 172 s), with output identical to a full re-prefill.
- Exact speculation: verify rows previously rounded differently from single-token decode. Three causes were fixed: final-norm fusion, the attention core split, and a GDN K-block size. Greedy speculative output is now token-identical to plain decoding in every launch setup: 20/20 HTTP checks, up from 14/20.
Other interesting data
- Power (stock 1,350 MHz, DFlash decode): board input ~213 W average and 371 W peak; chip core rail ~68 W average and 163 W peak. The firmware's "150 W" limit covers only the core rail, and in this firmware's power mode setting it has no effect. Only clock caps take hold.
| Clock cap | Decode at 128 / 2K prompt | vs. stock | Board power avg / peak |
|---|---|---|---|
| 1,350 MHz (stock) | 50.2 / 53.0 | 100% | 213 / 371 W |
| 1,200 MHz | 43.8 / 47.6 | ~88% | 200 / 309 W |
| 1,000 MHz | 37.2 / 40.8 | ~76% | 191 / 265 W |
| 800 MHz | 30.3 / 32.6 | ~61% | 172 / 239 W |
Decode speed tracks clock almost linearly while power drops only 10–20%, so stock is also the most efficient setting.
- Thermal: after 5–12 minutes of sustained load the AI clock steps from 1,350 to 1,343 MHz. That ~0.5% dip is larger than many of the wins we were measuring, so every benchmark bracket was clock-checked and some candidates needed a 40-minute cooldown.
- Bandwidth actually achieved: gate/up ~455–458 GB/s, LM head ~445 GB/s, GDN mega ~421 GB/s, roughly 82–89% of the 512 GB/s rating for the best kernels. The whole model still runs well below that, because of non-matmul work and smaller projections.
- Temperature > 0: speculation is greedy-only, so sampled requests fall back to plain decode with host sampling, at ~9–11 tok/s versus ~50 greedy.
- Prefill buckets: prompts are padded to 128/256/512/1,024 tokens, then 2,048-token chunks. A 1,094-token prompt pays for 2,048 and scores 648 tok/s; a 2,045-token prompt scores 1,210 tok/s.
- Speculative acceptance governs everything. DFlash at a ~125 ms cycle on a 7K-token prompt accepted only 2.7 of 7 drafts and ran at 29 tok/s. On repetitive text it accepts 5–6 and runs at 55+.
- Concurrency: this build is single-sequence by design. 75.5 MB of L1 recurrent state leaves no room for a second sequence, and multiple clients queue in FIFO order. Our Qwen 35B work solved this with DRAM-resident per-request state; 27B has not been ported yet.
Lessons learned
- Count bytes per token first. For batch-1 decode, the ceiling is bytes read divided by bandwidth. Our real 18.68 GB/token, not a naive 13.5 GB, set every target, and the gate/up experiment showed arithmetic didn't matter.
- Port optimizations, don't flip flags. Moving from 9B to 27B broke head counts, core counts, L1 budgets and kernel batch limits. Several fast paths silently fell back to generic code until we checked that each one actually activated at 27B shapes.
- Make exactness a gate, not a hope. Equivalence proofs and bit-exact A/Bs caught a circular-buffer overwrite, unflushed NoC atomics, an SDPA core-split rounding change, KV DMA misalignment and more, each before it shipped.
- A bit-exact checkpoint is not a quality measurement. Proofs showed the checkpoint loaded exactly as intended, but plain-rounding BFP4 still cost ~5% perplexity on code. Measure the quant itself against BF16, and calibrate on every domain you serve.
- Benchmark hygiene matters as much as kernels. We hit thermal clock dips, a GiB-vs-GB unit bug that briefly "proved" bandwidth was exhausted, CPU contention from a background job (35B dropped from 123 to 77 tok/s), and unlike-for-unlike prompts. Compare identical token streams, bracket clocks, and state which prompt you used.
- Speculative decoding wins once verification is cheap. DFlash only beat MTP after its 8-row verify received the same fusions as MTP's. Drafters tolerate BFP8; BFP4 drafters collapsed acceptance every time we tried.
- Small exact wins compound. Successive rounds of mostly 1–10% changes took the same weights from 27.7 to ~50 tok/s.
- Be careful with firmware. Our second P150, meant for TP=2, was left unbootable by an interrupted 19.15.0 firmware update: ARC firmware not booting, board ID 0000, no telemetry. It needs SWD recovery or an RMA. TP=2 is on hold.
What's next
- 4-bit KV cache with rotation (in progress): 8-bit keys plus Hadamard-rotated, round-to-nearest BFP4 values, inspired by TurboQuant. Goals: free ~2 GB so the faster DFlash drafter fits at 262K, and cut long-context attention reads. It ships only if quality matches today's BFP8 KV.
- More bytes off the hot path: a GPTQ'd LM head and selective 4-bit for low-sensitivity BFP8 projections, held to the same KL/top-1/gsm8k bar.
- One plugin: consolidating the 9B, 27B, 35B and Gemma runtimes into a single
vllm_tt_plugin, and packaging the build as a one-command LocalMaxxing CLI recipe for P150 owners. - Verified LocalMaxxing runs: the endpoint needs to expose prompt/output hashes, raw engine timings and draft/accept counters so submissions earn the verified badge.
- Missing features: speculative sampling at temperature > 0,
prompt_logprobson TT (needed for HellaSwag scoring), and batched decode with DRAM-resident GDN state for multi-user serving. - Drafting ideas: an n-gram lookup structure alongside DFlash for repetitive text.
- TP=2: once the second card is recovered, the same model on two P150s halves weight reads per card. The ideal ceiling moves from ~21 to ~42 tok/s without speculation, before communication costs.
Reproduce
The Hugging Face repo includes a launch.py that downloads and checks the package against its SHA-256 manifest, loads the pinned runtime image, and serves an OpenAI-compatible API on one P150:
curl -fL -o launch.py \
https://huggingface.co/Lottolabs/Qwen3.8-27B-TT-Mixed-BFP4-BFP8-P150/resolve/main/launch.py
python3 launch.py --cache-root "$HOME/.cache/qwen27b-tt-native" --device-ownership-confirmed
The public package contains the GPTQ checkpoint and the native-context runtime. The prefix-caching and exact-speculation builds described above are local and will be published after the next release check.
Attached LocalMaxxing runs: MTP-1 official prompt (23.8 tok/s), DFlash peak custom prompt (55.9 tok/s), and DFlash 2,045-token prefill run (1,212 tok/s prefill). Eval shards: gsm8k shard 1 (cmul8gl3p0ihjlq0147enm8y6) and arc-challenge shard 1 (cmul9l8v90iknlq01uiolccyc).
