Начало работыТаблица лидеровDecode calculatorМоделиReportsОборудованиеБенчмаркиМаркетплейсАрендаProДокументация API
Язык
Actual Computer — Every computer, one endpoint

Reports

Community-authored local inference reports

Published field reportsnewest

Zeus 2× Arc Pro B70: llama.cpp SYCL FP16 tuning across Qwen3.8 and nine other model routes

Qwen3.8-27B-GGUF

On two Intel Arc Pro B70s, a b11190 SYCL FP16 build substantially improved prompt processing across 10 production routes. This report gives matched serving results, copy-heavy generation gains from n-gram drafting, rejected tuning variants, and measurement limits.

Q6_K_XL focus; Q8_0, Q4_K_M, UD-Q4_K_XL llama.cpp b11190, SYCL / Level Zero
@Captain-Tripps Sep 26, 2026 0

895 tok/s from a 27B dense model on one RTX 5080: DFlash2 + n-gram hybrid drafting on PQ2_0

Bonsai-2-27B-Ternary-CRACK-GGUF

How a 2.13-bpw ternary 27B, a self-speculative DFlash2 draft head, an n-gram drafting layer, and four small llama.cpp patches took a single RTX 5080 to 895 tok/s warm / 142 tok/s fresh - with verified runs, kernel-level analysis, and a full reproduction package.

PQ2_0 llama.cppRTX 5080 · 16 GB
7 benchmarks
@zotowata Sep 21, 2026 0

FreeToken vs vLLM on one RTX 3090: BF16 MoE RAM offload compared

Qwen3.6-35B-A3B

An apples-to-apples single-RTX-3090 comparison of FreeToken expert offload and vLLM CPU offload with Qwen3.6-35B-A3B BF16. FreeToken reached 42.78 tok/s established decode versus 7.01 tok/s for tuned vLLM, while vLLM delivered lower TTFT. Includes sequential, prefill, concurrency, memory, tuning, and MTP/DFlash compatibility findings.

BF16 FreeToken,vLLM
@Lottolabs Aug 25, 2026 0