{"rows":[{"id":"cmtc8lwnb000to601lv07j4ox","modelRevision":"86f0121a909e464fed5055b213ae5032c0066d0f","promptTokens":5,"outputTokens":64,"contextLength":512,"prefillTokens":null,"batchSize":256,"ttftMs":5.887,"tokSOut":306140,"tokSPrefill":217428.234375,"tokSTotal":297348.71875,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-28T00:52:30.984Z","notes":"Official llama-batched-bench b10665 on Vulkan with 1x RTX PRO 6000 Blackwell. Aggregate S_TG = (256 independent sequences × 64 generated tokens) / 0.053518 s = 306,140 tok/s. Ten independent process runs: median 306,140 tok/s, range 261,045–311,619 (first runs include cold start; warmed results 301,293–311,619). PP=5 tokens per sequence, TG=64, prompts not shared, N_KV=17,664, -b/-ub 512, all layers offloaded (-ngl 999); Vulkan automatically disabled unsupported Flash Attention. No speculative decoding. Exact HF revision 86f0121a909e464fed5055b213ae5032c0066d0f; source safetensors SHA-256 ac54a8dea35ed6ccac58300c78ac9bfd368b35cc3b6b18e0b2391a733800e601. F32 GGUF converted locally from those exact weights; converter used standard gpt-2 pre-tokenizer fallback because this custom BPE is unrecognized (affects generated text/tokenization, not throughput), matching the caveat on the existing CPU record. GGUF SHA-256 377ffe9c6ec79f0034bc3cc01ce37a8a85c26004b19c0548bdfbc6175f9a0a00. Ubuntu 24.04, NVIDIA driver 595.84. Output figure is transparent aggregate batch throughput, not single-stream.","engineFlags":{"commandSnippet":"llama-batched-bench -m stories-llama2-50k-f32.gguf -c 32768 -b 512 -ub 512 -ngl 999 -npp 5 -ntg 64 -npl 256 --output-format jsonl","tensorParallel":null,"gpuLayers":999,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"delphi-suite/stories-llama2-50k","displayName":"stories-llama2-50k","family":"Llama","params":0,"activeParams":null,"isMoE":false,"baseModel":null},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"RTX PRO 6000 Blackwell","gpuCount":1,"vramGb":96,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen 9 9950X3D","os":"Ubuntu 24.04","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"llama.cpp","engineVersion":"b10665","quantization":"F32","backend":"vulkan"},"user":{"id":"cmtc57rvs0000o601l1w3d4sg","username":"ProCreations","verified":false,"verifiedAt":null,"pro":false},"hardwareGroupLabel":"RTX PRO 6000 Blackwell","hardwareGroupKey":"DISCRETE_GPU:rtx pro 6000 blackwell","rank":1,"reactionCounts":{},"myEmoji":null},{"id":"cmsxu9zyi0ck7ms01v41wipnd","modelRevision":"main","promptTokens":512,"outputTokens":64,"contextLength":2048,"prefillTokens":null,"batchSize":1,"ttftMs":2.7,"tokSOut":36715.725095,"tokSPrefill":189913.419741,"tokSTotal":129756.4,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-17T23:02:34.315Z","notes":"CPU-only on a Ryzen 7 9800X3D (-ngl 0, GPU hidden): llama-bench b10470, -t 3 -lm mlock --poll 100 --prio 3, 512p/64n, 20 reps, best of N attempts. No public GGUF existed for this 50K-param model (1 layer, hidden_size 6), so it was converted locally to F32 with llama.cpp's convert_hf_to_gguf.py (custom tokenizer not recognized by the converter; labeled with the standard gpt-2 rule — affects text quality only, not benchmark speed). Model runs entirely from per-core L2 cache.","engineFlags":{"commandSnippet":"E:\\llama.cpp-b10470\\llama-bench.exe -m E:\\huggingface\\stories-llama2-50k.F32.gguf -ngl 0 -t 3 -lm mlock --poll 100 --prio 3 --delay 2 -p 512 -n 64 -r 20 -o json","tensorParallel":null,"gpuLayers":0,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"delphi-suite/stories-llama2-50k","displayName":"stories-llama2-50k","family":"Llama","params":0,"activeParams":null,"isMoE":false,"baseModel":null},"hardware":{"hwClass":"CPU_ONLY","gpuName":null,"gpuCount":1,"vramGb":null,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen 7 9800X3D","os":"windows","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"llama.cpp","engineVersion":null,"quantization":"F32","backend":"cuda"},"user":{"id":"cmsxeur0y0b7xms01r3olg42y","username":"CsGoat","verified":false,"verifiedAt":null,"pro":false},"hardwareGroupLabel":"AMD Ryzen 7 9800X3D","hardwareGroupKey":"CPU_ONLY:amd ryzen 7 9800x3d","rank":2,"reactionCounts":{},"myEmoji":null},{"id":"cmsxtay630cj6ms01d7jbpdi8","modelRevision":"main","promptTokens":512,"outputTokens":64,"contextLength":2048,"prefillTokens":null,"batchSize":1,"ttftMs":41.5,"tokSOut":16740.728684,"tokSPrefill":12338.485672,"tokSTotal":12709.8,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-17T22:35:19.036Z","notes":null,"engineFlags":{"commandSnippet":"E:\\llama.cpp-b10470\\llama-bench.exe -m E:\\huggingface\\hub\\models--ggml-org--models\\snapshots\\499bc8821c6b12b4e53c5bffcb21ec206f212d81\\tinyllamas\\stories260K.gguf -ngl 0 -t 3 -lm mlock --poll 100 --prio 3 --delay 2 -p 512 -n 64 -r 20 -o json","tensorParallel":null,"gpuLayers":0,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"ggml-org/models-moved","displayName":"models-moved","family":null,"params":null,"activeParams":null,"isMoE":false,"baseModel":null},"hardware":{"hwClass":"CPU_ONLY","gpuName":null,"gpuCount":1,"vramGb":null,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen 7 9800X3D","os":"windows","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"llama.cpp","engineVersion":null,"quantization":"F32","backend":"cuda"},"user":{"id":"cmsxeur0y0b7xms01r3olg42y","username":"CsGoat","verified":false,"verifiedAt":null,"pro":false},"hardwareGroupLabel":"AMD Ryzen 7 9800X3D","hardwareGroupKey":"CPU_ONLY:amd ryzen 7 9800x3d","rank":3,"reactionCounts":{},"myEmoji":null},{"id":"cmsxs32ai0chxms01j90izclk","modelRevision":"main","promptTokens":512,"outputTokens":64,"contextLength":2048,"prefillTokens":null,"batchSize":1,"ttftMs":70,"tokSOut":6319.713109,"tokSPrefill":7310.920391,"tokSTotal":7185.7,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-17T22:01:11.515Z","notes":null,"engineFlags":{"commandSnippet":"E:\\llama.cpp-b10470\\llama-bench.exe -m E:\\huggingface\\hub\\models--afrideva--Tinystories-gpt-0.1-3m-GGUF\\snapshots\\3cad9ac61b9fa06f2647b6b77d38640bfa1766aa\\tinystories-gpt-0.1-3m.Q4_K_M.gguf -ngl 0 -t 4 -lm mlock --poll 100 --prio 3 --delay 2 -p 512 -n 64 -r 20 -o json","tensorParallel":null,"gpuLayers":0,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"afrideva/Tinystories-gpt-0.1-3m-GGUF","displayName":"Tinystories-gpt-0.1-3m-GGUF","family":"Gpt","params":0,"activeParams":null,"isMoE":false,"baseModel":{"hfId":"segestic/Tinystories-gpt-0.1-3m","displayName":"Tinystories-gpt-0.1-3m","params":0,"activeParams":null,"isMoE":false}},"hardware":{"hwClass":"CPU_ONLY","gpuName":null,"gpuCount":1,"vramGb":null,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen 7 9800X3D","os":"windows","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"llama.cpp","engineVersion":null,"quantization":"Q4_K_M","backend":"cuda"},"user":{"id":"cmsxeur0y0b7xms01r3olg42y","username":"CsGoat","verified":false,"verifiedAt":null,"pro":false},"hardwareGroupLabel":"AMD Ryzen 7 9800X3D","hardwareGroupKey":"CPU_ONLY:amd ryzen 7 9800x3d","rank":4,"reactionCounts":{},"myEmoji":null},{"id":"cmsnp1rko00jco001mvymwmsb","modelRevision":"main","promptTokens":0,"outputTokens":64,"contextLength":1024,"prefillTokens":null,"batchSize":112,"ttftMs":120,"tokSOut":6001.4,"tokSPrefill":null,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-10T20:38:30.360Z","notes":"BATCHED AGGREGATE THROUGHPUT — not single-stream decode. DATA-PARALLEL across 4 cards; aggregate of 4 independent replicas. batch=112, concurrency=112, 100 input rows, 3 waves, max_tokens=64, ctx=1024, KV=q8, greedy; TTFT not comparable. Per-card median MODEL tok/s: d0:1495.3, d1:1504.4, d2:1500.1, d3:1501.6 (sum 6001.4); wall-clock aggregate 3889.4 tok/s (per-card wall d0:969.6, d1:975.7, d2:973.2, d3:970.9). MODEL tok/s is sum of per-card decode rates excluding inter-wave gap; wall includes all overhead. Do not compare DP aggregate to single-stream. engine: hipfire 0.3.0+b0bcc3f91506, backend rocm, gfx1201 x4 quant: MQ4R = MagnumQuant 4-bit Redline (graded mixed-tier: MQ4 attn/router/shared + graded MQ4 routed experts, Redline retained-PM4 dispatch), 4.16 bpw effective (18700048128 bytes *8 / 35.95B params) model file: qwen3.6-35b-a3b.mq4r sha256 4685c140c46b1a6f prompt: uniform short-serve jsonl md5 3d22208dcf539818ff2e8341f53474ee, 64 max_new method: median of 3 waves, fresh process per wave, temperature 0, bit-exact outputs (sha ec47dcfb9b0c54e0) evidence: sealed case a3b-mq4r-dp4-b112-gfx1201-r9700-4x, corpus manifest manifest-sha12-8f3c2e1a4b9d not measured: peak VRAM (not instrumented this campaign)","engineFlags":{"commandSnippet":"hipfire-batch-generate --model ~/.hipfire/models/qwen3.6-35b-a3b.mq4r --batch 112 --max-seq 1024 --max-new 64 --devices 0,1,2,3 --data-parallel --kv q8 --fresh-process-per-wave","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":"q8","attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"Qwen/Qwen3.6-35B-A3B","displayName":"Qwen3.6-35B-A3B","family":"Qwen","params":36,"activeParams":3,"isMoE":true,"baseModel":null},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"Radeon AI Pro R9700","gpuCount":4,"vramGb":32,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen Threadripper 9970X","os":"Ubuntu 26.04","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"hipfire","engineVersion":"0.3.0+b0bcc3f91506","quantization":"MQ4R","backend":"rocm"},"user":{"id":"cmoeye1gq0000le04ie2kqs58","username":"schuttdev","verified":false,"verifiedAt":null,"pro":true},"hardwareGroupLabel":"Radeon AI Pro R9700","hardwareGroupKey":"DISCRETE_GPU:radeon ai pro r9700","rank":5,"reactionCounts":{},"myEmoji":null},{"id":"cmsnp232a00kao001r4wh9920","modelRevision":"main","promptTokens":0,"outputTokens":64,"contextLength":1024,"prefillTokens":null,"batchSize":100,"ttftMs":120,"tokSOut":5993.3,"tokSPrefill":null,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-10T20:38:45.251Z","notes":"BATCHED AGGREGATE THROUGHPUT — not single-stream decode. DATA-PARALLEL across 4 cards; aggregate of 4 independent replicas. batch=100, concurrency=100, 100 input rows, 3 waves, max_tokens=64, ctx=1024, KV=q8, greedy; TTFT not comparable. Per-card median MODEL tok/s: d0:1493.8, d1:1500.3, d2:1499.3, d3:1499.9 (sum 5993.3); wall-clock aggregate 3885.8 tok/s (per-card wall d0:970.3, d1:973.6, d2:972.5, d3:969.4). MODEL tok/s is sum of per-card decode rates excluding inter-wave gap; wall includes all overhead. Do not compare DP aggregate to single-stream. engine: hipfire 0.3.0+b0bcc3f91506, backend rocm, gfx1201 x4 quant: MQ4R = MagnumQuant 4-bit Redline (graded mixed-tier: MQ4 attn/router/shared + graded MQ4 routed experts, Redline retained-PM4 dispatch), 4.16 bpw effective (18700048128 bytes *8 / 35.95B params) model file: qwen3.6-35b-a3b.mq4r sha256 4685c140c46b1a6f prompt: uniform short-serve jsonl md5 3d22208dcf539818ff2e8341f53474ee, 64 max_new method: median of 3 waves, fresh process per wave, temperature 0, bit-exact outputs (sha ec47dcfb9b0c54e0) evidence: sealed case a3b-mq4r-dp4-b100-gfx1201-r9700-4x, corpus manifest manifest-sha12-8f3c2e1a4b9d not measured: peak VRAM (not instrumented this campaign)","engineFlags":{"commandSnippet":"hipfire-batch-generate --model ~/.hipfire/models/qwen3.6-35b-a3b.mq4r --batch 100 --max-seq 1024 --max-new 64 --devices 0,1,2,3 --data-parallel --kv q8 --fresh-process-per-wave","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":"q8","attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"Qwen/Qwen3.6-35B-A3B","displayName":"Qwen3.6-35B-A3B","family":"Qwen","params":36,"activeParams":3,"isMoE":true,"baseModel":null},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"Radeon AI Pro R9700","gpuCount":4,"vramGb":32,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen Threadripper 9970X","os":"Ubuntu 26.04","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"hipfire","engineVersion":"0.3.0+b0bcc3f91506","quantization":"MQ4R","backend":"rocm"},"user":{"id":"cmoeye1gq0000le04ie2kqs58","username":"schuttdev","verified":false,"verifiedAt":null,"pro":true},"hardwareGroupLabel":"Radeon AI Pro R9700","hardwareGroupKey":"DISCRETE_GPU:radeon ai pro r9700","rank":6,"reactionCounts":{},"myEmoji":null},{"id":"cmsxs31hf0chsms0121zih53z","modelRevision":"main","promptTokens":512,"outputTokens":128,"contextLength":2048,"prefillTokens":null,"batchSize":1,"ttftMs":69.4,"tokSOut":5860.853353,"tokSPrefill":7373.26084,"tokSTotal":7011.4,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-17T22:01:10.467Z","notes":null,"engineFlags":{"commandSnippet":"E:\\llama.cpp-b10470\\llama-bench.exe -m E:\\huggingface\\hub\\models--afrideva--Tinystories-gpt-0.1-3m-GGUF\\snapshots\\3cad9ac61b9fa06f2647b6b77d38640bfa1766aa\\tinystories-gpt-0.1-3m.Q4_K_M.gguf -ngl 0 -t 4 -lm mlock --poll 100 --prio 3 --delay 2 -p 512 -n 128 -r 20 -o json","tensorParallel":null,"gpuLayers":0,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"afrideva/Tinystories-gpt-0.1-3m-GGUF","displayName":"Tinystories-gpt-0.1-3m-GGUF","family":"Gpt","params":0,"activeParams":null,"isMoE":false,"baseModel":{"hfId":"segestic/Tinystories-gpt-0.1-3m","displayName":"Tinystories-gpt-0.1-3m","params":0,"activeParams":null,"isMoE":false}},"hardware":{"hwClass":"CPU_ONLY","gpuName":null,"gpuCount":1,"vramGb":null,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen 7 9800X3D","os":"windows","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"llama.cpp","engineVersion":null,"quantization":"Q4_K_M","backend":"cuda"},"user":{"id":"cmsxeur0y0b7xms01r3olg42y","username":"CsGoat","verified":false,"verifiedAt":null,"pro":false},"hardwareGroupLabel":"AMD Ryzen 7 9800X3D","hardwareGroupKey":"CPU_ONLY:amd ryzen 7 9800x3d","rank":7,"reactionCounts":{},"myEmoji":null},{"id":"cmrio3ys401uvmj01lcom115h","modelRevision":"main","promptTokens":512,"outputTokens":128,"contextLength":2048,"prefillTokens":null,"batchSize":1,"ttftMs":null,"tokSOut":4980.696149,"tokSPrefill":7664.915099,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-07-13T03:33:40.181Z","notes":"CPU-only build (Vulkan/CUDA/HIP disabled). 8 physical-core threads. One explicit throwaway plus llama-bench internal warmup; two measured repetitions. Headless Ubuntu, BIOS Silent mode, Core Turbo Boost disabled. Decode samples tok/s: 4930.78, 5030.61. GGUF-reported parameters: 6963968.","engineFlags":{"commandSnippet":"llama-bench -m TinyStories-3M-Q4_K_M.gguf -t 8 -p 512 -n 128 -r 2 -o json","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"segestic/Tinystories-gpt-0.1-3m","displayName":"Tinystories-gpt-0.1-3m","family":"Gpt","params":0,"activeParams":null,"isMoE":false,"baseModel":null},"hardware":{"hwClass":"CPU_ONLY","gpuName":null,"gpuCount":1,"vramGb":null,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen 7 7840HS","os":"Ubuntu 24.04.4 LTS","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"llama.cpp","engineVersion":"commit 6b4dc21","quantization":"Q4_K_M","backend":"cpu"},"user":{"id":"cmoext1fr0000jl04skh3hjqf","username":"steveseguin","verified":false,"verifiedAt":null,"pro":false},"hardwareGroupLabel":"AMD Ryzen 7 7840HS","hardwareGroupKey":"CPU_ONLY:amd ryzen 7 7840hs","rank":8,"reactionCounts":{},"myEmoji":null},{"id":"cmrio8nyr01vumj010r04d1cl","modelRevision":"main","promptTokens":512,"outputTokens":128,"contextLength":2048,"prefillTokens":null,"batchSize":1,"ttftMs":null,"tokSOut":4979.352396,"tokSPrefill":7767.919137,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-07-13T03:37:19.443Z","notes":"CPU quantization/thread sweep. CPU-only build; 8 physical-core threads; llama-bench warmup and two measured repetitions. Headless; BIOS Silent mode; Core Turbo Boost disabled. Decode samples tok/s: 4972.35, 4986.36. GGUF-reported parameters: 6963968.","engineFlags":{"commandSnippet":"llama-bench -m TinyStories-3M-Q2_K.gguf -t 8 -p 512 -n 128 -r 2 -o json","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"segestic/Tinystories-gpt-0.1-3m","displayName":"Tinystories-gpt-0.1-3m","family":"Gpt","params":0,"activeParams":null,"isMoE":false,"baseModel":null},"hardware":{"hwClass":"CPU_ONLY","gpuName":null,"gpuCount":1,"vramGb":null,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen 7 7840HS","os":"Ubuntu 24.04.4 LTS","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"llama.cpp","engineVersion":"commit 6b4dc21","quantization":"Q2_K","backend":"cpu"},"user":{"id":"cmoext1fr0000jl04skh3hjqf","username":"steveseguin","verified":false,"verifiedAt":null,"pro":false},"hardwareGroupLabel":"AMD Ryzen 7 7840HS","hardwareGroupKey":"CPU_ONLY:amd ryzen 7 7840hs","rank":9,"reactionCounts":{},"myEmoji":null},{"id":"cmrio8ond01vzmj01rlxlvcc0","modelRevision":"main","promptTokens":512,"outputTokens":128,"contextLength":2048,"prefillTokens":null,"batchSize":1,"ttftMs":null,"tokSOut":4911.930039,"tokSPrefill":7848.04374,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-07-13T03:37:20.329Z","notes":"CPU quantization/thread sweep. CPU-only build; 8 physical-core threads; llama-bench warmup and two measured repetitions. Headless; BIOS Silent mode; Core Turbo Boost disabled. Decode samples tok/s: 4925.10, 4898.76. GGUF-reported parameters: 6963968.","engineFlags":{"commandSnippet":"llama-bench -m TinyStories-3M-Q8_0.gguf -t 8 -p 512 -n 128 -r 2 -o json","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"segestic/Tinystories-gpt-0.1-3m","displayName":"Tinystories-gpt-0.1-3m","family":"Gpt","params":0,"activeParams":null,"isMoE":false,"baseModel":null},"hardware":{"hwClass":"CPU_ONLY","gpuName":null,"gpuCount":1,"vramGb":null,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen 7 7840HS","os":"Ubuntu 24.04.4 LTS","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"llama.cpp","engineVersion":"commit 6b4dc21","quantization":"Q8_0","backend":"cpu"},"user":{"id":"cmoext1fr0000jl04skh3hjqf","username":"steveseguin","verified":false,"verifiedAt":null,"pro":false},"hardwareGroupLabel":"AMD Ryzen 7 7840HS","hardwareGroupKey":"CPU_ONLY:amd ryzen 7 7840hs","rank":10,"reactionCounts":{},"myEmoji":null},{"id":"cmsnp1spu00jho001f7bpvkac","modelRevision":"main","promptTokens":27,"outputTokens":128,"contextLength":4096,"prefillTokens":null,"batchSize":128,"ttftMs":644.1,"tokSOut":4816.072,"tokSPrefill":null,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-10T20:38:31.843Z","notes":"BATCHED AGGREGATE THROUGHPUT — not single-stream decode.\nbatch=128, concurrency=128, 128 requests/run, 3 runs, mode=fixed_wave, max_tokens=128, ctx=4096, KV=default (unspecified, hipfire fp16), greedy; TTFT is under load.\nengine: hipfire 0.3.0+2e0c4d39f29f, backend rocm, gfx1201\nquant: MQ4 = MagnumQuant 4-bit, group-256, FWHT-rotated, 5.68 bpw effective (163173488 bytes *8 / 229693184 params)\nmodel file: lfm2.5-230m.mq4 sha256 3b92b9fd27c68d7f\nprompt: longform.txt md5 e6f8ee7d66b934c49493347500f48083, 27 prompt tokens\nmethod: median of 3 runs, fresh process, temperature 0 (greedy)\ncaveats: batched-prefill-coalesced build; pre-coalesce fixed-wave baseline B128 is 1523.1 tok/s (case lfm2.5-230m-gfx1201-b128r128-fixed).\nevidence: sealed case lfm2.5-230m-gfx1201-b128r128-batched-prefill-fixed, host k9lin, corpus manifest 45b76818319f\nnot measured: peak VRAM (not instrumented this campaign)","engineFlags":{"commandSnippet":"scripts/lmx_continuous_batch.py --model /home/kaden/.hipfire/models/lfm2.5-230m.mq4 --prompt-file benchmarks/prompts/sweep/longform.txt --batch-size 128 --requests 128 --runs 3 --max-tokens 128 --max-seq 4096 --port 11543 --home-root /tmp/lmx-lfm230-b128-coalesce-home --log-dir /home/kaden/localmaxxing-benchmarks/raw/lfm230-batched-prefill-coalesce-b128 --out /home/kaden/localmaxxing-benchmarks/validation/lfm230-local-b128r128-batched-prefill-coalesce.json --cli target/release/hipfire --daemon target/release/examples/daemon --device 0 --timeout 600 --thinking off","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"LiquidAI/LFM2.5-230M","displayName":"LFM2.5-230M","family":null,"params":0,"activeParams":null,"isMoE":false,"baseModel":{"hfId":"LiquidAI/LFM2.5-230M-Base","displayName":"LFM2.5-230M-Base","params":0,"activeParams":null,"isMoE":false}},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"RX 9070 XT","gpuCount":1,"vramGb":16,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen 9 3900X","os":"Ubuntu 24.04","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"hipfire","engineVersion":"0.3.0+2e0c4d39f29f","quantization":"MQ4","backend":"rocm"},"user":{"id":"cmoeye1gq0000le04ie2kqs58","username":"schuttdev","verified":false,"verifiedAt":null,"pro":true},"hardwareGroupLabel":"RX 9070 XT","hardwareGroupKey":"DISCRETE_GPU:rx 9070 xt","rank":11,"reactionCounts":{},"myEmoji":null},{"id":"cmrio8ozf01w5mj01054z2ynf","modelRevision":"main","promptTokens":512,"outputTokens":128,"contextLength":2048,"prefillTokens":null,"batchSize":1,"ttftMs":null,"tokSOut":4747.562394,"tokSPrefill":7728.936959,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-07-13T03:37:20.763Z","notes":"CPU quantization/thread sweep. CPU-only build; 8 physical-core threads; llama-bench warmup and two measured repetitions. Headless; BIOS Silent mode; Core Turbo Boost disabled. Decode samples tok/s: 4747.24, 4747.88. GGUF-reported parameters: 6963968.","engineFlags":{"commandSnippet":"llama-bench -m TinyStories-3M-F16.gguf -t 8 -p 512 -n 128 -r 2 -o json","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"segestic/Tinystories-gpt-0.1-3m","displayName":"Tinystories-gpt-0.1-3m","family":"Gpt","params":0,"activeParams":null,"isMoE":false,"baseModel":null},"hardware":{"hwClass":"CPU_ONLY","gpuName":null,"gpuCount":1,"vramGb":null,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen 7 7840HS","os":"Ubuntu 24.04.4 LTS","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"llama.cpp","engineVersion":"commit 6b4dc21","quantization":"F16","backend":"cpu"},"user":{"id":"cmoext1fr0000jl04skh3hjqf","username":"steveseguin","verified":false,"verifiedAt":null,"pro":false},"hardwareGroupLabel":"AMD Ryzen 7 7840HS","hardwareGroupKey":"CPU_ONLY:amd ryzen 7 7840hs","rank":12,"reactionCounts":{},"myEmoji":null},{"id":"cmsxkcqvd0btgms0137r1bi9o","modelRevision":"main","promptTokens":512,"outputTokens":128,"contextLength":2048,"prefillTokens":null,"batchSize":1,"ttftMs":1.1,"tokSOut":4360.94,"tokSPrefill":471381.77,"tokSTotal":21026.6,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-17T18:24:46.345Z","notes":null,"engineFlags":{"commandSnippet":"E:\\llama.cpp-b10470\\llama-bench.exe -m E:\\huggingface\\hub\\models--ggml-org--models\\snapshots\\499bc8821c6b12b4e53c5bffcb21ec206f212d81\\tinyllamas\\stories260K.gguf -p 512 -n 128","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"ggml-org/models-moved","displayName":"models-moved","family":null,"params":null,"activeParams":null,"isMoE":false,"baseModel":null},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"RTX 5090","gpuCount":1,"vramGb":32,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"amd64","os":"windows","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"llama.cpp","engineVersion":null,"quantization":"F32","backend":"cuda"},"user":{"id":"cmsxeur0y0b7xms01r3olg42y","username":"CsGoat","verified":false,"verifiedAt":null,"pro":false},"hardwareGroupLabel":"RTX 5090","hardwareGroupKey":"DISCRETE_GPU:rtx 5090","rank":13,"reactionCounts":{},"myEmoji":null},{"id":"cmsnp1tv300jko001nfbt14ek","modelRevision":"main","promptTokens":27,"outputTokens":128,"contextLength":4096,"prefillTokens":null,"batchSize":128,"ttftMs":580.5,"tokSOut":4347.308,"tokSPrefill":null,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-10T20:38:33.327Z","notes":"BATCHED AGGREGATE THROUGHPUT — not single-stream decode.\nbatch=128, concurrency=128, 128 requests/run, 1 runs, mode=fixed_wave, max_tokens=128, ctx=4096, KV=default (unspecified, hipfire fp16), greedy; TTFT is under load.\nengine: hipfire 0.3.0+48802a005625, backend rocm, gfx1201\nquant: MQ4 = MagnumQuant 4-bit, group-256, FWHT-rotated, 5.68 bpw effective (163173488 bytes *8 / 229693184 params)\nmodel file: lfm2.5-230m.mq4 sha256 3b92b9fd27c68d7f\nprompt: longform.txt md5 e6f8ee7d66b934c49493347500f48083, 27 prompt tokens\nmethod: median of 1 runs, fresh process, temperature 0 (greedy)\ncaveats: R9700 230M opt-B128 single-card figure device d2 4347.3 tok/s; per-card d0/d1/d2/d3 = 4291/4310/4347/4309 tok/s (all similar).\nevidence: sealed case lfm25-230m-opt-b128-fixed-wave-20260809-gfx1201-r9700-d2, host hiptrx, corpus manifest 45b76818319f\nnot measured: peak VRAM (not instrumented this campaign)","engineFlags":{"commandSnippet":"scripts/lmx_continuous_batch.py --model /home/kaden/.hipfire/models/lfm2.5-230m.mq4 --prompt-file benchmarks/prompts/sweep/longform.txt --batch-size 128 --requests 128 --runs 3 --max-tokens 128 --max-seq 4096 --port 18822 --home-root /tmp/lmx_opt230_d2_home --log-dir /tmp/lmx_opt230_d2_logs --out /tmp/lmx_opt230_d2_b128.json --device 2","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"LiquidAI/LFM2.5-230M","displayName":"LFM2.5-230M","family":null,"params":0,"activeParams":null,"isMoE":false,"baseModel":{"hfId":"LiquidAI/LFM2.5-230M-Base","displayName":"LFM2.5-230M-Base","params":0,"activeParams":null,"isMoE":false}},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"Radeon AI Pro R9700","gpuCount":1,"vramGb":32,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen Threadripper 9970X","os":"Ubuntu 26.04","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"hipfire","engineVersion":"0.3.0+48802a005625","quantization":"MQ4","backend":"rocm"},"user":{"id":"cmoeye1gq0000le04ie2kqs58","username":"schuttdev","verified":false,"verifiedAt":null,"pro":true},"hardwareGroupLabel":"Radeon AI Pro R9700","hardwareGroupKey":"DISCRETE_GPU:radeon ai pro r9700","rank":14,"reactionCounts":{},"myEmoji":null},{"id":"cmsnp247500kfo001s4pg03pu","modelRevision":"main","promptTokens":0,"outputTokens":64,"contextLength":1024,"prefillTokens":null,"batchSize":64,"ttftMs":120,"tokSOut":4207.7,"tokSPrefill":null,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-10T20:38:46.721Z","notes":"BATCHED AGGREGATE THROUGHPUT — not single-stream decode. DATA-PARALLEL across 4 cards; aggregate of 4 independent replicas. batch=64, concurrency=64, 100 input rows, 3 waves, max_tokens=64, ctx=1024, KV=q8, greedy; TTFT not comparable. Per-card median MODEL tok/s: d0:1046.6, d1:1053.0, d2:1052.2, d3:1055.9 (sum 4207.7); wall-clock aggregate 3058.6 tok/s (per-card wall d0:761.8, d1:766.0, d2:765.4, d3:765.4). MODEL tok/s is sum of per-card decode rates excluding inter-wave gap; wall includes all overhead. Do not compare DP aggregate to single-stream. engine: hipfire 0.3.0+b0bcc3f91506, backend rocm, gfx1201 x4 quant: MQ4R = MagnumQuant 4-bit Redline (graded mixed-tier: MQ4 attn/router/shared + graded MQ4 routed experts, Redline retained-PM4 dispatch), 4.16 bpw effective (18700048128 bytes *8 / 35.95B params) model file: qwen3.6-35b-a3b.mq4r sha256 4685c140c46b1a6f prompt: uniform short-serve jsonl md5 3d22208dcf539818ff2e8341f53474ee, 64 max_new method: median of 3 waves, fresh process per wave, temperature 0, bit-exact outputs (sha ec47dcfb9b0c54e0) evidence: sealed case a3b-mq4r-dp4-b64-gfx1201-r9700-4x, corpus manifest manifest-sha12-8f3c2e1a4b9d not measured: peak VRAM (not instrumented this campaign)","engineFlags":{"commandSnippet":"hipfire-batch-generate --model ~/.hipfire/models/qwen3.6-35b-a3b.mq4r --batch 64 --max-seq 1024 --max-new 64 --devices 0,1,2,3 --data-parallel --kv q8 --fresh-process-per-wave","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":"q8","attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"Qwen/Qwen3.6-35B-A3B","displayName":"Qwen3.6-35B-A3B","family":"Qwen","params":36,"activeParams":3,"isMoE":true,"baseModel":null},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"Radeon AI Pro R9700","gpuCount":4,"vramGb":32,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen Threadripper 9970X","os":"Ubuntu 26.04","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"hipfire","engineVersion":"0.3.0+b0bcc3f91506","quantization":"MQ4R","backend":"rocm"},"user":{"id":"cmoeye1gq0000le04ie2kqs58","username":"schuttdev","verified":false,"verifiedAt":null,"pro":true},"hardwareGroupLabel":"Radeon AI Pro R9700","hardwareGroupKey":"DISCRETE_GPU:radeon ai pro r9700","rank":15,"reactionCounts":{},"myEmoji":null},{"id":"cmsnp1v0f00jno001z3qt1vjn","modelRevision":"main","promptTokens":27,"outputTokens":128,"contextLength":4096,"prefillTokens":null,"batchSize":128,"ttftMs":659.7,"tokSOut":3669.847,"tokSPrefill":null,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-10T20:38:34.815Z","notes":"BATCHED AGGREGATE THROUGHPUT — not single-stream decode.\nbatch=128, concurrency=128, 128 requests/run, 3 runs, mode=fixed_wave, max_tokens=128, ctx=4096, KV=default (unspecified, hipfire fp16), greedy; TTFT is under load.\nengine: hipfire 0.3.0+2e0c4d39f29f, backend rocm, gfx1201\nquant: MQ4 = MagnumQuant 4-bit, group-256, FWHT-rotated, 5.18 bpw effective (229474032 bytes *8 / 354483968 params)\nmodel file: lfm2.5-350m.mq4 sha256 4885d1cecbad59d7\nprompt: longform.txt md5 e6f8ee7d66b934c49493347500f48083, 27 prompt tokens\nmethod: median of 3 runs, fresh process, temperature 0 (greedy)\ncaveats: batched-prefill-coalesced build; pre-coalesce baseline B128 is 1237.7 tok/s (case lfm2.5-350m-gfx1201-b128r128-fixed).\nevidence: sealed case lfm2.5-350m-gfx1201-b128r128-batched-prefill-fixed, host k9lin, corpus manifest 45b76818319f\nnot measured: peak VRAM (not instrumented this campaign)","engineFlags":{"commandSnippet":"scripts/lmx_continuous_batch.py --model /home/kaden/.hipfire/models/lfm2.5-350m.mq4 --prompt-file benchmarks/prompts/sweep/longform.txt --batch-size 128 --requests 128 --runs 3 --max-tokens 128 --max-seq 4096 --port 11541 --home-root /tmp/lmx-lfm350-b128-coalesce-home --log-dir /home/kaden/localmaxxing-benchmarks/raw/lfm350-batched-prefill-coalesce-b128 --out /home/kaden/localmaxxing-benchmarks/validation/lfm350-local-b128r128-batched-prefill-coalesce.json --cli target/release/hipfire --daemon target/release/examples/daemon --device 0 --timeout 600 --thinking off","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"LiquidAI/LFM2.5-350M","displayName":"LFM2.5-350M","family":null,"params":0,"activeParams":null,"isMoE":false,"baseModel":{"hfId":"LiquidAI/LFM2.5-350M-Base","displayName":"LFM2.5-350M-Base","params":0,"activeParams":null,"isMoE":false}},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"RX 9070 XT","gpuCount":1,"vramGb":16,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen 9 3900X","os":"Ubuntu 24.04","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"hipfire","engineVersion":"0.3.0+2e0c4d39f29f","quantization":"MQ4","backend":"rocm"},"user":{"id":"cmoeye1gq0000le04ie2kqs58","username":"schuttdev","verified":false,"verifiedAt":null,"pro":true},"hardwareGroupLabel":"RX 9070 XT","hardwareGroupKey":"DISCRETE_GPU:rx 9070 xt","rank":16,"reactionCounts":{},"myEmoji":null},{"id":"cmrio408801vhmj01zdwdqzob","modelRevision":"main","promptTokens":512,"outputTokens":128,"contextLength":2048,"prefillTokens":null,"batchSize":1,"ttftMs":null,"tokSOut":3569.083997,"tokSPrefill":74947.634851,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-07-13T03:33:42.057Z","notes":"CPU-only build (Vulkan/CUDA/HIP disabled). 8 physical-core threads. One explicit throwaway plus llama-bench internal warmup; two measured repetitions. Headless Ubuntu, BIOS Silent mode, Core Turbo Boost disabled. Decode samples tok/s: 3570.080, 3568.090. GGUF-reported parameters: 33883264.","engineFlags":{"commandSnippet":"llama-bench -m TinyStories-1M-Q4_K_M.gguf -t 8 -p 512 -n 128 -r 2 -o json","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"ivnle/tinystories-lay4-hs128-hd2-1M","displayName":"tinystories-lay4-hs128-hd2-1M","family":"Llama","params":0,"activeParams":null,"isMoE":false,"baseModel":null},"hardware":{"hwClass":"CPU_ONLY","gpuName":null,"gpuCount":1,"vramGb":null,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen 7 7840HS","os":"Ubuntu 24.04.4 LTS","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"llama.cpp","engineVersion":"commit 6b4dc21","quantization":"Q4_K_M","backend":"cpu"},"user":{"id":"cmoext1fr0000jl04skh3hjqf","username":"steveseguin","verified":false,"verifiedAt":null,"pro":false},"hardwareGroupLabel":"AMD Ryzen 7 7840HS","hardwareGroupKey":"CPU_ONLY:amd ryzen 7 7840hs","rank":17,"reactionCounts":{},"myEmoji":null},{"id":"cmsnp25c400kko001bfn1g0ch","modelRevision":"main","promptTokens":27,"outputTokens":128,"contextLength":4096,"prefillTokens":null,"batchSize":128,"ttftMs":604.6,"tokSOut":3283.31,"tokSPrefill":null,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-10T20:38:48.196Z","notes":"BATCHED AGGREGATE THROUGHPUT — not single-stream decode.\nbatch=128, concurrency=128, 128 requests/run, 1 runs, mode=fixed_wave, max_tokens=128, ctx=4096, KV=default (unspecified, hipfire fp16), greedy; TTFT is under load.\nengine: hipfire 0.3.0+48802a005625, backend rocm, gfx1201\nquant: MQ4 = MagnumQuant 4-bit, group-256, FWHT-rotated, 5.18 bpw effective (229474032 bytes *8 / 354483968 params)\nmodel file: lfm2.5-350m.mq4 sha256 4885d1cecbad59d7\nprompt: longform.txt md5 e6f8ee7d66b934c49493347500f48083, 27 prompt tokens\nmethod: median of 1 runs, fresh process, temperature 0 (greedy)\ncaveats: R9700 350M opt-B128 single-card figure device d0 3283.3 tok/s; per-card spread d0/d1/d2/d3 = 3283/3189/3247/1609 tok/s (d3 outlier 1609). Do not use 4-card median without noting outlier.\nevidence: sealed case lfm25-350m-opt-b128-fixed-wave-20260809-gfx1201-r9700-d0, host hiptrx, corpus manifest 45b76818319f\nnot measured: peak VRAM (not instrumented this campaign)","engineFlags":{"commandSnippet":"scripts/lmx_continuous_batch.py --model /home/kaden/.hipfire/models/lfm2.5-350m.mq4 --prompt-file benchmarks/prompts/sweep/longform.txt --batch-size 128 --requests 128 --runs 3 --max-tokens 128 --max-seq 4096 --port 18830 --home-root /tmp/lmx_opt350_d0_home --log-dir /tmp/lmx_opt350_d0_logs --out /tmp/lmx_opt350_d0_b128.json --device 0","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"LiquidAI/LFM2.5-350M","displayName":"LFM2.5-350M","family":null,"params":0,"activeParams":null,"isMoE":false,"baseModel":{"hfId":"LiquidAI/LFM2.5-350M-Base","displayName":"LFM2.5-350M-Base","params":0,"activeParams":null,"isMoE":false}},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"Radeon AI Pro R9700","gpuCount":1,"vramGb":32,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen Threadripper 9970X","os":"Ubuntu 26.04","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"hipfire","engineVersion":"0.3.0+48802a005625","quantization":"MQ4","backend":"rocm"},"user":{"id":"cmoeye1gq0000le04ie2kqs58","username":"schuttdev","verified":false,"verifiedAt":null,"pro":true},"hardwareGroupLabel":"Radeon AI Pro R9700","hardwareGroupKey":"DISCRETE_GPU:radeon ai pro r9700","rank":18,"reactionCounts":{},"myEmoji":null},{"id":"cmokhzrb40005ld0493vy3ltd","modelRevision":"main","promptTokens":45000,"outputTokens":3000,"contextLength":262144,"prefillTokens":null,"batchSize":10,"ttftMs":143,"tokSOut":2665.14,"tokSPrefill":null,"tokSTotal":55846.85,"peakVramGb":133,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-04-29T20:18:51.521Z","notes":"Requires DeepGEMM installation after uv pip install vllm:\ngit clone --recursive https://github.com/deepseek-ai/DeepGEMM.git\ncd DeepGEMM\nDG_FORCE_BUILD=1 uv pip install --force-reinstall --no-build-isolation .\n\nBenchmark script:\nhttps://github.com/keennay/gpu-cluster-setup/blob/master/benchmark.py\n\nBenchmark Script Command:\npython benchmark.py --concurrency 10 --num-prompts 100 --label <label> --model <model>\n\nTTFT Results:\nMedian: 143ms, Mean: 340ms","engineFlags":{"commandSnippet":"vllm serve Qwen/Qwen3.5-0.8B-Base --served-model-name qwen3 --trust-remote-code --tensor-parallel-size 1 --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder --gpu-memory-utilization 0.95 --speculative-config '{\"method\":\"qwen3_next_mtp\",\"num_speculative_tokens\":2}'  --no-enable-prefix-caching --host 0.0.0.0 --port 8000","tensorParallel":1,"gpuLayers":4,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"Qwen/Qwen3.5-0.8B-Base","displayName":"Qwen3.5-0.8B-Base","family":"Qwen","params":1,"activeParams":null,"isMoE":false,"baseModel":null},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"H200 NVL","gpuCount":1,"vramGb":141,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD EPYC 9175F","os":"Ubuntu 24.04","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"vllm","engineVersion":"0.20.1rc1.dev74+gfaab18955","quantization":"BF16","backend":"cuda"},"user":{"id":"cmoeptucv000ajr042kk6r3i3","username":"keennay","verified":false,"verifiedAt":null,"pro":false},"hardwareGroupLabel":"H200 NVL","hardwareGroupKey":"DISCRETE_GPU:h200 nvl","rank":19,"reactionCounts":{},"myEmoji":null},{"id":"cmrk3vsz102zzmj01p4vu2ouo","modelRevision":"main","promptTokens":50592,"outputTokens":24576,"contextLength":32768,"prefillTokens":null,"batchSize":96,"ttftMs":38.65,"tokSOut":2527.68,"tokSPrefill":null,"tokSTotal":7731.146,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-07-14T03:42:59.437Z","notes":"Rented JarvisLabs RTX PRO 6000 Blackwell 96GB (cloud), CUDA 13 build sm_120. Median of 3 runs, temp 0, cold prefill, prompt ~50592 tok. Head-to-head vs my 3x R9700 rig, same GGUF quants.","engineFlags":{"commandSnippet":"vllm serve cyankiwi/GLM-4.5-Air-AWQ-4bit --max-model-len 32768 --max-num-seqs 96 --enable-chunked-prefill --enable-prefix-caching --gpu-memory-utilization 0.90","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"zai-org/GLM-4.5-Air","displayName":"GLM-4.5-Air","family":null,"params":110,"activeParams":null,"isMoE":true,"baseModel":null},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"RTX PRO 6000 Blackwell","gpuCount":1,"vramGb":96,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":null,"os":null,"isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"vllm","engineVersion":"vllm-0.25.0","quantization":"AWQ-4bit","backend":"cuda"},"user":{"id":"cmorcd2jg0008l804hfkput08","username":"1337Hero","verified":false,"verifiedAt":null,"pro":false},"hardwareGroupLabel":"RTX PRO 6000 Blackwell","hardwareGroupKey":"DISCRETE_GPU:rtx pro 6000 blackwell","rank":20,"reactionCounts":{},"myEmoji":null},{"id":"cmrk3vsnw02zumj01rfnmu5x4","modelRevision":"main","promptTokens":33728,"outputTokens":16384,"contextLength":32768,"prefillTokens":null,"batchSize":64,"ttftMs":40.37,"tokSOut":2037.835,"tokSPrefill":null,"tokSTotal":6232.908,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-07-14T03:42:59.036Z","notes":"Rented JarvisLabs RTX PRO 6000 Blackwell 96GB (cloud), CUDA 13 build sm_120. Median of 3 runs, temp 0, cold prefill, prompt ~33728 tok. Head-to-head vs my 3x R9700 rig, same GGUF quants.","engineFlags":{"commandSnippet":"vllm serve cyankiwi/GLM-4.5-Air-AWQ-4bit --max-model-len 32768 --max-num-seqs 96 --enable-chunked-prefill --enable-prefix-caching --gpu-memory-utilization 0.90","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"zai-org/GLM-4.5-Air","displayName":"GLM-4.5-Air","family":null,"params":110,"activeParams":null,"isMoE":true,"baseModel":null},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"RTX PRO 6000 Blackwell","gpuCount":1,"vramGb":96,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":null,"os":null,"isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"vllm","engineVersion":"vllm-0.25.0","quantization":"AWQ-4bit","backend":"cuda"},"user":{"id":"cmorcd2jg0008l804hfkput08","username":"1337Hero","verified":false,"verifiedAt":null,"pro":false},"hardwareGroupLabel":"RTX PRO 6000 Blackwell","hardwareGroupKey":"DISCRETE_GPU:rtx pro 6000 blackwell","rank":21,"reactionCounts":{},"myEmoji":null},{"id":"cmsnp26gr00kno0012jshzkpm","modelRevision":"main","promptTokens":27,"outputTokens":128,"contextLength":4096,"prefillTokens":null,"batchSize":256,"ttftMs":1260,"tokSOut":1876.917,"tokSPrefill":null,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-10T20:38:49.660Z","notes":"BATCHED AGGREGATE THROUGHPUT — not single-stream decode.\nbatch=256, concurrency=256, 256 requests/run, 3 runs, mode=fixed_wave, max_tokens=128, ctx=4096, KV=default (unspecified, hipfire fp16), greedy; TTFT is under load.\nengine: hipfire 0.3.0+2e0c4d39f29f, backend rocm, gfx1201\nquant: MQ4 = MagnumQuant 4-bit, group-256, FWHT-rotated, 5.18 bpw effective (229474032 bytes *8 / 354483968 params)\nmodel file: lfm2.5-350m.mq4 sha256 4885d1cecbad59d7\nprompt: longform.txt md5 e6f8ee7d66b934c49493347500f48083, 27 prompt tokens\nmethod: median of 3 runs, fresh process, temperature 0 (greedy)\ncaveats: batched-prefill-coalesced build; pre-coalesce baseline B256 is 1258.5 tok/s (case lfm2.5-350m-gfx1201-b256r256-fixed).\nevidence: sealed case lfm2.5-350m-gfx1201-b256r256-batched-prefill-fixed, host k9lin, corpus manifest 45b76818319f\nnot measured: peak VRAM (not instrumented this campaign)","engineFlags":{"commandSnippet":"scripts/lmx_continuous_batch.py --model /home/kaden/.hipfire/models/lfm2.5-350m.mq4 --prompt-file benchmarks/prompts/sweep/longform.txt --batch-size 256 --requests 256 --runs 3 --max-tokens 128 --max-seq 4096 --port 11542 --home-root /tmp/lmx-lfm350-b256-coalesce-home --log-dir /home/kaden/localmaxxing-benchmarks/raw/lfm350-batched-prefill-coalesce-b256 --out /home/kaden/localmaxxing-benchmarks/validation/lfm350-local-b256r256-batched-prefill-coalesce.json --cli target/release/hipfire --daemon target/release/examples/daemon --device 0 --timeout 600 --thinking off","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"LiquidAI/LFM2.5-350M","displayName":"LFM2.5-350M","family":null,"params":0,"activeParams":null,"isMoE":false,"baseModel":{"hfId":"LiquidAI/LFM2.5-350M-Base","displayName":"LFM2.5-350M-Base","params":0,"activeParams":null,"isMoE":false}},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"RX 9070 XT","gpuCount":1,"vramGb":16,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen 9 3900X","os":"Ubuntu 24.04","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"hipfire","engineVersion":"0.3.0+2e0c4d39f29f","quantization":"MQ4","backend":"rocm"},"user":{"id":"cmoeye1gq0000le04ie2kqs58","username":"schuttdev","verified":false,"verifiedAt":null,"pro":true},"hardwareGroupLabel":"RX 9070 XT","hardwareGroupKey":"DISCRETE_GPU:rx 9070 xt","rank":22,"reactionCounts":{},"myEmoji":null},{"id":"cmsnp27le00kqo0015qt1pjdp","modelRevision":"main","promptTokens":27,"outputTokens":128,"contextLength":4096,"prefillTokens":null,"batchSize":8,"ttftMs":80.25,"tokSOut":1704.611,"tokSPrefill":null,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-10T20:38:51.122Z","notes":"BATCHED AGGREGATE THROUGHPUT — not single-stream decode.\nbatch=8, concurrency=8, 16 requests/run, 3 runs, mode=continuous, max_tokens=128, ctx=4096, KV=default (unspecified, hipfire fp16), greedy; TTFT is under load.\nengine: hipfire 0.3.0+b0bcc3f91506, backend rocm, gfx1201\nquant: MQ4 = MagnumQuant 4-bit, group-256, FWHT-rotated, 5.68 bpw effective (163173488 bytes *8 / 229693184 params)\nmodel file: lfm2.5-230m.mq4 sha256 3b92b9fd27c68d7f\nprompt: longform.txt md5 e6f8ee7d66b934c49493347500f48083, 27 prompt tokens\nmethod: median of 3 runs, fresh process, temperature 0 (greedy)\ncaveats: FINAL continuous numbers shown; earlier continuous_refill sweep measured lower (B4R8 737.9 vs 1044.8, B8R16 1023.2 vs 1704.6).\nevidence: sealed case lfm2.5-230m-gfx1201-b8r16-FINAL-continuous, host k9lin, corpus manifest 45b76818319f\nnot measured: peak VRAM (not instrumented this campaign)","engineFlags":{"commandSnippet":"scripts/lmx_continuous_batch.py --model /home/kaden/.hipfire/models/lfm2.5-230m.mq4 --prompt-file benchmarks/prompts/sweep/longform.txt --batch-size 8 --requests 16 --runs 3 --max-tokens 128 --max-seq 4096 --port 11605 --home-root /home/kaden/lmx-final-runs/work/FINAL/lfm2.5-230m/continuous/home-b8r16 --log-dir /home/kaden/lmx-final-runs/logs/FINAL/lfm2.5-230m/continuous/b8r16 --out /home/kaden/lmx-final-runs/lfm230-FINAL-b8r16-gfx1201-continuous.json --device 0","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"LiquidAI/LFM2.5-230M","displayName":"LFM2.5-230M","family":null,"params":0,"activeParams":null,"isMoE":false,"baseModel":{"hfId":"LiquidAI/LFM2.5-230M-Base","displayName":"LFM2.5-230M-Base","params":0,"activeParams":null,"isMoE":false}},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"RX 9070 XT","gpuCount":1,"vramGb":16,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen 9 3900X","os":"Ubuntu 24.04","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"hipfire","engineVersion":"0.3.0+b0bcc3f91506","quantization":"MQ4","backend":"rocm"},"user":{"id":"cmoeye1gq0000le04ie2kqs58","username":"schuttdev","verified":false,"verifiedAt":null,"pro":true},"hardwareGroupLabel":"RX 9070 XT","hardwareGroupKey":"DISCRETE_GPU:rx 9070 xt","rank":23,"reactionCounts":{},"myEmoji":null},{"id":"cmsnp28q500kuo001070dwcev","modelRevision":"main","promptTokens":27,"outputTokens":128,"contextLength":4096,"prefillTokens":null,"batchSize":256,"ttftMs":2537.9,"tokSOut":1597.704,"tokSPrefill":null,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-10T20:38:52.589Z","notes":"BATCHED AGGREGATE THROUGHPUT — not single-stream decode.\nbatch=256, concurrency=256, 256 requests/run, 3 runs, mode=fixed_wave, max_tokens=128, ctx=4096, KV=default (unspecified, hipfire fp16), greedy; TTFT is under load.\nengine: hipfire 0.3.0+439233d92959, backend rocm, gfx1201\nquant: MQ4 = MagnumQuant 4-bit, group-256, FWHT-rotated, 5.68 bpw effective (163173488 bytes *8 / 229693184 params)\nmodel file: lfm2.5-230m.mq4 sha256 3b92b9fd27c68d7f\nprompt: longform.txt md5 e6f8ee7d66b934c49493347500f48083, 27 prompt tokens\nmethod: median of 3 runs, fresh process, temperature 0 (greedy)\nevidence: sealed case lfm25-230m-b256-fixed-wave-20260809-gfx1201-r9700, host hiptrx, corpus manifest 45b76818319f\nnot measured: peak VRAM (not instrumented this campaign)","engineFlags":{"commandSnippet":"scripts/lmx_continuous_batch.py --model /home/kaden/.hipfire/models/lfm2.5-230m.mq4 --prompt-file benchmarks/prompts/sweep/longform.txt --batch-size 256 --requests 256 --runs 3 --max-tokens 128 --max-seq 2048 --port 18816 --home-root /tmp/lmx_cb_home6 --log-dir /tmp/lmx_cb_logs6 --out /tmp/lmx_cb230_b256.json --device 0","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"LiquidAI/LFM2.5-230M","displayName":"LFM2.5-230M","family":null,"params":0,"activeParams":null,"isMoE":false,"baseModel":{"hfId":"LiquidAI/LFM2.5-230M-Base","displayName":"LFM2.5-230M-Base","params":0,"activeParams":null,"isMoE":false}},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"Radeon AI Pro R9700","gpuCount":1,"vramGb":32,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen Threadripper 9970X","os":"Ubuntu 26.04","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"hipfire","engineVersion":"0.3.0+439233d92959","quantization":"MQ4","backend":"rocm"},"user":{"id":"cmoeye1gq0000le04ie2kqs58","username":"schuttdev","verified":false,"verifiedAt":null,"pro":true},"hardwareGroupLabel":"Radeon AI Pro R9700","hardwareGroupKey":"DISCRETE_GPU:radeon ai pro r9700","rank":24,"reactionCounts":{},"myEmoji":null},{"id":"cmsnp29vb00kxo00147qwfmi0","modelRevision":"main","promptTokens":27,"outputTokens":128,"contextLength":4096,"prefillTokens":null,"batchSize":128,"ttftMs":2067.8,"tokSOut":1585.588,"tokSPrefill":null,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-10T20:38:54.071Z","notes":"BATCHED AGGREGATE THROUGHPUT — not single-stream decode.\nbatch=128, concurrency=128, 128 requests/run, 3 runs, mode=fixed_wave, max_tokens=128, ctx=4096, KV=default (unspecified, hipfire fp16), greedy; TTFT is under load.\nengine: hipfire 0.3.0+439233d92959, backend rocm, gfx1201\nquant: MQ4 = MagnumQuant 4-bit, group-256, FWHT-rotated, 5.68 bpw effective (163173488 bytes *8 / 229693184 params)\nmodel file: lfm2.5-230m.mq4 sha256 3b92b9fd27c68d7f\nprompt: longform.txt md5 e6f8ee7d66b934c49493347500f48083, 27 prompt tokens\nmethod: median of 3 runs, fresh process, temperature 0 (greedy)\ncaveats: pre-coalesce fixed-wave B128 1585.6 tok/s; opt batched-prefill-coalesced B128 per-card up to 4347 tok/s (case lfm25-230m-opt-b128...-d2).\nevidence: sealed case lfm25-230m-b128-fixed-wave-20260809-gfx1201-r9700, host hiptrx, corpus manifest 45b76818319f\nnot measured: peak VRAM (not instrumented this campaign)","engineFlags":{"commandSnippet":"scripts/lmx_continuous_batch.py --model /home/kaden/.hipfire/models/lfm2.5-230m.mq4 --prompt-file benchmarks/prompts/sweep/longform.txt --batch-size 128 --requests 128 --runs 3 --max-tokens 128 --max-seq 2048 --port 18812 --home-root /tmp/lmx_cb_home2 --log-dir /tmp/lmx_cb_logs2 --out /tmp/lmx_cb230_b128.json --device 0","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"LiquidAI/LFM2.5-230M","displayName":"LFM2.5-230M","family":null,"params":0,"activeParams":null,"isMoE":false,"baseModel":{"hfId":"LiquidAI/LFM2.5-230M-Base","displayName":"LFM2.5-230M-Base","params":0,"activeParams":null,"isMoE":false}},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"Radeon AI Pro R9700","gpuCount":1,"vramGb":32,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen Threadripper 9970X","os":"Ubuntu 26.04","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"hipfire","engineVersion":"0.3.0+439233d92959","quantization":"MQ4","backend":"rocm"},"user":{"id":"cmoeye1gq0000le04ie2kqs58","username":"schuttdev","verified":false,"verifiedAt":null,"pro":true},"hardwareGroupLabel":"Radeon AI Pro R9700","hardwareGroupKey":"DISCRETE_GPU:radeon ai pro r9700","rank":25,"reactionCounts":{},"myEmoji":null},{"id":"cmsnp2b0g00l1o001ekfx6umo","modelRevision":"main","promptTokens":27,"outputTokens":128,"contextLength":4096,"prefillTokens":null,"batchSize":128,"ttftMs":2231.9,"tokSOut":1523.117,"tokSPrefill":null,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-10T20:38:55.551Z","notes":"BATCHED AGGREGATE THROUGHPUT — not single-stream decode.\nbatch=128, concurrency=128, 128 requests/run, 3 runs, mode=fixed_wave, max_tokens=128, ctx=4096, KV=default (unspecified, hipfire fp16), greedy; TTFT is under load.\nengine: hipfire 0.3.0+b973024a973f, backend rocm, gfx1201\nquant: MQ4 = MagnumQuant 4-bit, group-256, FWHT-rotated, 5.68 bpw effective (163173488 bytes *8 / 229693184 params)\nmodel file: lfm2.5-230m.mq4 sha256 3b92b9fd27c68d7f\nprompt: longform.txt md5 e6f8ee7d66b934c49493347500f48083, 27 prompt tokens\nmethod: median of 3 runs, fresh process, temperature 0 (greedy)\ncaveats: pre-coalesce baseline (no batched prefill coalescing); coalesced build achieves 4816.1 tok/s (case lfm2.5-230m-gfx1201-b128r128-batched-prefill-fixed).\nevidence: sealed case lfm2.5-230m-gfx1201-b128r128-fixed, host k9lin, corpus manifest 45b76818319f\nnot measured: peak VRAM (not instrumented this campaign)","engineFlags":{"commandSnippet":"scripts/lmx_continuous_batch.py --model /home/kaden/.hipfire/models/lfm2.5-230m.mq4 --prompt-file benchmarks/prompts/sweep/longform.txt --batch-size 128 --requests 128 --runs 3 --max-tokens 128 --max-seq 4096 --port 11530 --home-root /home/kaden/localmaxxing-benchmarks/raw/lfm230-local-b128r128/home --log-dir /home/kaden/localmaxxing-benchmarks/raw/lfm230-local-b128r128/logs --out /home/kaden/localmaxxing-benchmarks/validation/lfm230-local-b128r128.json --cli /home/kaden/hipfire-lmx-lfm/target/release/hipfire --daemon /home/kaden/hipfire-lmx-lfm/target/release/examples/daemon --device 0","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"LiquidAI/LFM2.5-230M","displayName":"LFM2.5-230M","family":null,"params":0,"activeParams":null,"isMoE":false,"baseModel":{"hfId":"LiquidAI/LFM2.5-230M-Base","displayName":"LFM2.5-230M-Base","params":0,"activeParams":null,"isMoE":false}},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"RX 9070 XT","gpuCount":1,"vramGb":16,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen 9 3900X","os":"Ubuntu 24.04","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"hipfire","engineVersion":"0.3.0+b973024a973f","quantization":"MQ4","backend":"rocm"},"user":{"id":"cmoeye1gq0000le04ie2kqs58","username":"schuttdev","verified":false,"verifiedAt":null,"pro":true},"hardwareGroupLabel":"RX 9070 XT","hardwareGroupKey":"DISCRETE_GPU:rx 9070 xt","rank":26,"reactionCounts":{},"myEmoji":null},{"id":"cmsnp1w7300jqo001njoxttpa","modelRevision":"main","promptTokens":27,"outputTokens":128,"contextLength":4096,"prefillTokens":null,"batchSize":8,"ttftMs":87.2,"tokSOut":1470.768,"tokSPrefill":null,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-10T20:38:36.352Z","notes":"BATCHED AGGREGATE THROUGHPUT — not single-stream decode.\nbatch=8, concurrency=8, 16 requests/run, 3 runs, mode=continuous_refill, max_tokens=128, ctx=4096, KV=default (unspecified, hipfire fp16), greedy; TTFT is under load.\nengine: hipfire 0.3.0+b0bcc3f91506, backend rocm, gfx1100\nquant: MQ4 = MagnumQuant 4-bit, group-256, FWHT-rotated, 5.18 bpw effective (229474032 bytes *8 / 354483968 params)\nmodel file: lfm2.5-350m.mq4 sha256 4885d1cecbad59d7\nprompt: longform.txt md5 e6f8ee7d66b934c49493347500f48083, 27 prompt tokens\nmethod: median of 3 runs, fresh process, temperature 0 (greedy)\nevidence: sealed case lfm2.5-350m-gfx1100-hip0-b8r16, host hipx, corpus manifest 45b76818319f\nnot measured: peak VRAM (not instrumented this campaign)","engineFlags":{"commandSnippet":"scripts/lmx_continuous_batch.py --model /home/kaden/.hipfire/models/lfm2.5-350m.mq4 --prompt-file /home/kaden/hipfire-lmx-final-26cf8e148/benchmarks/prompts/sweep/longform.txt --batch-size 8 --requests 16 --runs 3 --max-tokens 128 --max-seq 4096 --port 11554 --home-root /home/kaden/lmx-final-runs/homes/FINAL-gfx1100-350-batch-b8r16-home --log-dir /home/kaden/lmx-final-runs/logs/FINAL/lfm2.5-350m/gfx1100-batch-b8r16 --out /home/kaden/lmx-final-runs/lfm2.5-350m-gfx1100-hip0-batch-b8r16.json --device 0 --timeout 600","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"LiquidAI/LFM2.5-350M","displayName":"LFM2.5-350M","family":null,"params":0,"activeParams":null,"isMoE":false,"baseModel":{"hfId":"LiquidAI/LFM2.5-350M-Base","displayName":"LFM2.5-350M-Base","params":0,"activeParams":null,"isMoE":false}},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"RX 7900 XTX","gpuCount":1,"vramGb":24,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen AI Max+ 395","os":"Ubuntu 26.04","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"hipfire","engineVersion":"0.3.0+b0bcc3f91506","quantization":"MQ4","backend":"rocm"},"user":{"id":"cmoeye1gq0000le04ie2kqs58","username":"schuttdev","verified":false,"verifiedAt":null,"pro":true},"hardwareGroupLabel":"RX 7900 XTX","hardwareGroupKey":"DISCRETE_GPU:rx 7900 xtx","rank":27,"reactionCounts":{},"myEmoji":null},{"id":"cmsnp2c5d00l5o0011vx4xq75","modelRevision":"main","promptTokens":27,"outputTokens":128,"contextLength":4096,"prefillTokens":null,"batchSize":64,"ttftMs":1055.7,"tokSOut":1456.962,"tokSPrefill":null,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-10T20:38:57.026Z","notes":"BATCHED AGGREGATE THROUGHPUT — not single-stream decode.\nbatch=64, concurrency=64, 64 requests/run, 3 runs, mode=fixed_wave, max_tokens=128, ctx=4096, KV=default (unspecified, hipfire fp16), greedy; TTFT is under load.\nengine: hipfire 0.3.0+1bec7dc3d1b5, backend rocm, gfx1201\nquant: MQ4 = MagnumQuant 4-bit, group-256, FWHT-rotated, 5.68 bpw effective (163173488 bytes *8 / 229693184 params)\nmodel file: lfm2.5-230m.mq4 sha256 3b92b9fd27c68d7f\nprompt: longform.txt md5 e6f8ee7d66b934c49493347500f48083, 27 prompt tokens\nmethod: median of 3 runs, fresh process, temperature 0 (greedy)\nevidence: sealed case lfm2.5-230m-gfx1201-b64r64-fixed, host k9lin, corpus manifest 45b76818319f\nnot measured: peak VRAM (not instrumented this campaign)","engineFlags":{"commandSnippet":"scripts/lmx_continuous_batch.py --model /home/kaden/.hipfire/models/lfm2.5-230m.mq4 --prompt-file benchmarks/prompts/sweep/longform.txt --batch-size 64 --requests 64 --runs 3 --max-tokens 128 --max-seq 4096 --port 11529 --home-root /home/kaden/localmaxxing-benchmarks/raw/lfm230-local-b64r64/home --log-dir /home/kaden/localmaxxing-benchmarks/raw/lfm230-local-b64r64/logs --out /home/kaden/localmaxxing-benchmarks/validation/lfm230-local-b64r64.json --cli /home/kaden/hipfire-lmx-lfm/target/release/hipfire --daemon /home/kaden/hipfire-lmx-lfm/target/release/examples/daemon --device 0","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"LiquidAI/LFM2.5-230M","displayName":"LFM2.5-230M","family":null,"params":0,"activeParams":null,"isMoE":false,"baseModel":{"hfId":"LiquidAI/LFM2.5-230M-Base","displayName":"LFM2.5-230M-Base","params":0,"activeParams":null,"isMoE":false}},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"RX 9070 XT","gpuCount":1,"vramGb":16,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen 9 3900X","os":"Ubuntu 24.04","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"hipfire","engineVersion":"0.3.0+1bec7dc3d1b5","quantization":"MQ4","backend":"rocm"},"user":{"id":"cmoeye1gq0000le04ie2kqs58","username":"schuttdev","verified":false,"verifiedAt":null,"pro":true},"hardwareGroupLabel":"RX 9070 XT","hardwareGroupKey":"DISCRETE_GPU:rx 9070 xt","rank":28,"reactionCounts":{},"myEmoji":null},{"id":"cmp7fqvp7006jo401h82kkyfb","modelRevision":"main","promptTokens":34998,"outputTokens":16384,"contextLength":32768,"prefillTokens":null,"batchSize":64,"ttftMs":270.64,"tokSOut":1405.394,"tokSPrefill":null,"tokSTotal":4407.469,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-05-15T21:34:40.123Z","notes":"Qwen2.5-7B BF16 vLLM ROCm source build. TP=1, max-seqs=128, batch-tokens=32768, client concurrency=64, chunked prefill + prefix caching. 1× R9700.","engineFlags":{"commandSnippet":"/home/mikekey/.venvs/vllm/bin/vllm serve /home/mikekey/models/hf/Qwen2.5-7B   --tensor-parallel-size 1   --dtype bfloat16   --max-model-len 32768   --max-num-seqs 128   --max-num-batched-tokens 32768   --enable-chunked-prefill   --enable-prefix-caching   --gpu-memory-utilization 0.90   --port 8000   --served-model-name qwen2.5-7b   --trust-remote-code","tensorParallel":1,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"Qwen/Qwen2.5-7B","displayName":"Qwen2.5-7B","family":"Qwen","params":8,"activeParams":null,"isMoE":false,"baseModel":null},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"Radeon AI Pro R9700","gpuCount":1,"vramGb":32,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":null,"os":"Arch Linux","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"vllm","engineVersion":"0.20.2+rocm723","quantization":"bf16","backend":"rocm"},"user":{"id":"cmorcd2jg0008l804hfkput08","username":"1337Hero","verified":false,"verifiedAt":null,"pro":false},"hardwareGroupLabel":"Radeon AI Pro R9700","hardwareGroupKey":"DISCRETE_GPU:radeon ai pro r9700","rank":29,"reactionCounts":{},"myEmoji":null},{"id":"cmp7f7rgi006co4019h144d9a","modelRevision":"main","promptTokens":34998,"outputTokens":16384,"contextLength":32768,"prefillTokens":null,"batchSize":64,"ttftMs":275.98,"tokSOut":1399.209,"tokSPrefill":null,"tokSTotal":4388.071,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-05-15T21:19:48.162Z","notes":"Qwen2.5-7B BF16 vLLM ROCm source build. TP=1, max-seqs=128, batch-tokens=32768, client concurrency=64, chunked prefill + prefix caching. 1× R9700.","engineFlags":{"commandSnippet":"/home/mikekey/.venvs/vllm/bin/vllm serve /home/mikekey/models/hf/Qwen2.5-7B   --tensor-parallel-size 1   --dtype bfloat16   --max-model-len 32768   --max-num-seqs 128   --max-num-batched-tokens 32768   --enable-chunked-prefill   --enable-prefix-caching   --gpu-memory-utilization 0.90   --port 8000   --served-model-name qwen2.5-7b   --trust-remote-code","tensorParallel":1,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"Qwen/Qwen2.5-7B","displayName":"Qwen2.5-7B","family":"Qwen","params":8,"activeParams":null,"isMoE":false,"baseModel":null},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"Radeon AI Pro R9700","gpuCount":1,"vramGb":32,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":null,"os":"Arch Linux","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"vllm","engineVersion":"0.20.2+rocm723","quantization":"bf16","backend":"rocm"},"user":{"id":"cmorcd2jg0008l804hfkput08","username":"1337Hero","verified":false,"verifiedAt":null,"pro":false},"hardwareGroupLabel":"Radeon AI Pro R9700","hardwareGroupKey":"DISCRETE_GPU:radeon ai pro r9700","rank":30,"reactionCounts":{},"myEmoji":null},{"id":"cmsnp2d9x00l9o001zb1nsuww","modelRevision":"main","promptTokens":27,"outputTokens":128,"contextLength":4096,"prefillTokens":null,"batchSize":32,"ttftMs":407.2,"tokSOut":1376.722,"tokSPrefill":null,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-10T20:38:58.485Z","notes":"BATCHED AGGREGATE THROUGHPUT — not single-stream decode.\nbatch=32, concurrency=32, 64 requests/run, 3 runs, mode=continuous_refill, max_tokens=128, ctx=4096, KV=default (unspecified, hipfire fp16), greedy; TTFT is under load.\nengine: hipfire 0.3.0+5f11453f892f, backend rocm, gfx1201\nquant: MQ4 = MagnumQuant 4-bit, group-256, FWHT-rotated, 5.68 bpw effective (163173488 bytes *8 / 229693184 params)\nmodel file: lfm2.5-230m.mq4 sha256 3b92b9fd27c68d7f\nprompt: longform.txt md5 e6f8ee7d66b934c49493347500f48083, 27 prompt tokens\nmethod: median of 3 runs, fresh process, temperature 0 (greedy)\nevidence: sealed case lfm2.5-230m-gfx1201-b32r64-continuous, host k9lin, corpus manifest 45b76818319f\nnot measured: peak VRAM (not instrumented this campaign)","engineFlags":{"commandSnippet":"scripts/lmx_continuous_batch.py --model /home/kaden/.hipfire/models/lfm2.5-230m.mq4 --prompt-file benchmarks/prompts/sweep/longform.txt --batch-size 32 --requests 64 --runs 3 --max-tokens 128 --max-seq 4096 --port 11525 --home-root /home/kaden/localmaxxing-benchmarks/raw/lfm230-local-b32r64/home --log-dir /home/kaden/localmaxxing-benchmarks/raw/lfm230-local-b32r64/logs --out /home/kaden/localmaxxing-benchmarks/validation/lfm230-local-b32r64.json --cli /home/kaden/hipfire-lmx-lfm/target/release/hipfire --daemon /home/kaden/hipfire-lmx-lfm/target/release/examples/daemon --device 0","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"LiquidAI/LFM2.5-230M","displayName":"LFM2.5-230M","family":null,"params":0,"activeParams":null,"isMoE":false,"baseModel":{"hfId":"LiquidAI/LFM2.5-230M-Base","displayName":"LFM2.5-230M-Base","params":0,"activeParams":null,"isMoE":false}},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"RX 9070 XT","gpuCount":1,"vramGb":16,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen 9 3900X","os":"Ubuntu 24.04","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"hipfire","engineVersion":"0.3.0+5f11453f892f","quantization":"MQ4","backend":"rocm"},"user":{"id":"cmoeye1gq0000le04ie2kqs58","username":"schuttdev","verified":false,"verifiedAt":null,"pro":true},"hardwareGroupLabel":"RX 9070 XT","hardwareGroupKey":"DISCRETE_GPU:rx 9070 xt","rank":31,"reactionCounts":{},"myEmoji":null},{"id":"cmqz6p4gl025toe01l70kn28a","modelRevision":"main","promptTokens":34,"outputTokens":838,"contextLength":8192,"prefillTokens":null,"batchSize":1,"ttftMs":260,"tokSOut":1369.3,"tokSPrefill":130.8,"tokSTotal":983.3,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-06-29T12:18:36.885Z","notes":"`lmx benchmark run vllm   --mode remote   --base-url http://localhost:8000   --hf-id RedHatAI/diffusiongemma-26B-A4B-it-FP8-dynamic   --served-model RedHatAI/diffusiongemma-26B-A4B-it-FP8-dynamic   --quantization fp8   --max-tokens 4096`","engineFlags":{"commandSnippet":"sudo docker run --gpus all   --privileged --ipc=host -p 8000:8000   -v ~/.cache/huggingface:/root/.cache/huggingface   vllm/vllm-openai:gemma RedHatAI/diffusiongemma-26B-A4B-it-FP8-dynamic   --max-num-seqs 4   --tensor-parallel-size 1","tensorParallel":1,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":"FA4","flashAttn":true,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"RedHatAI/diffusiongemma-26B-A4B-it-FP8-dynamic","displayName":"diffusiongemma-26B-A4B-it-FP8-dynamic","family":"Gemma","params":26,"activeParams":4,"isMoE":true,"baseModel":{"hfId":"google/diffusiongemma-26B-A4B-it","displayName":"diffusiongemma-26B-A4B-it","params":26,"activeParams":4,"isMoE":true}},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"H100","gpuCount":1,"vramGb":80,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"Intel(R) Xeon(R) Platinum 8481C","os":"Ubuntu 22.04.5 LTS","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"vllm","engineVersion":"0.24.0 (dev)","quantization":"FP8","backend":"cuda"},"user":{"id":"cmqz33l6e0231oe010lk3smio","username":"joaogante","verified":false,"verifiedAt":null,"pro":false},"hardwareGroupLabel":"H100","hardwareGroupKey":"DISCRETE_GPU:h100","rank":32,"reactionCounts":{},"myEmoji":null},{"id":"cmrz3la0h064ro401xaf55ehg","modelRevision":"main","promptTokens":512,"outputTokens":512,"contextLength":1024,"prefillTokens":null,"batchSize":1,"ttftMs":52.86484128860717,"tokSOut":1298.196691,"tokSPrefill":9828.284892,"tokSTotal":2293.45581167259,"peakVramGb":1.3173828125,"gpuPowerWatts":[61.8],"totalPowerWatts":61.8,"hardwareCost":null,"createdAt":"2026-07-24T15:31:20.945Z","notes":"512 prompt / 512 generation because this GPT-2 checkpoint has a 1024-token maximum context. Two throwaway runs preceded three measured configurations. Winner selected by minimum measured TTFT. Local file: /mnt/hdd/lmstudio_models/tiny-throughput-bench/gguf/tinystories-3m/tinystories-gpt-0.1-3m.Q8_0.gguf","engineFlags":{"commandSnippet":"/home/dario/llama.cpp/build/bin/llama-bench -m /mnt/hdd/lmstudio_models/tiny-throughput-bench/gguf/tinystories-3m/tinystories-gpt-0.1-3m.Q8_0.gguf -p 512 -n 512 -b 2048 -ub 512 -t 16 -ngl 99 -fa on -r 2 -o json --no-warmup","tensorParallel":1,"gpuLayers":99,"kvCacheDtype":"f16","attentionBackend":"flash_attn","flashAttn":true,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"segestic/Tinystories-gpt-0.1-3m","displayName":"Tinystories-gpt-0.1-3m","family":"Gpt","params":0,"activeParams":null,"isMoE":false,"baseModel":null},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"RTX 3060","gpuCount":1,"vramGb":12,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":null,"os":null,"isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"llama.cpp","engineVersion":"b9775 / be4a6a63e","quantization":"Q8_0","backend":"cuda"},"user":{"id":"cmqjl8xsp00mjnv01uq4z2m8m","username":"darioooooo0o","verified":false,"verifiedAt":null,"pro":false},"hardwareGroupLabel":"RTX 3060","hardwareGroupKey":"DISCRETE_GPU:rtx 3060","rank":33,"reactionCounts":{},"myEmoji":null},{"id":"cmsnp2eef00lco001ryud4uiu","modelRevision":"main","promptTokens":27,"outputTokens":128,"contextLength":4096,"prefillTokens":null,"batchSize":256,"ttftMs":2844.1,"tokSOut":1291.866,"tokSPrefill":null,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-10T20:38:59.943Z","notes":"BATCHED AGGREGATE THROUGHPUT — not single-stream decode.\nbatch=256, concurrency=256, 256 requests/run, 3 runs, mode=fixed_wave, max_tokens=128, ctx=4096, KV=default (unspecified, hipfire fp16), greedy; TTFT is under load.\nengine: hipfire 0.3.0+439233d92959, backend rocm, gfx1201\nquant: MQ4 = MagnumQuant 4-bit, group-256, FWHT-rotated, 5.18 bpw effective (229474032 bytes *8 / 354483968 params)\nmodel file: lfm2.5-350m.mq4 sha256 4885d1cecbad59d7\nprompt: longform.txt md5 e6f8ee7d66b934c49493347500f48083, 27 prompt tokens\nmethod: median of 3 runs, fresh process, temperature 0 (greedy)\nevidence: sealed case lfm25-350m-b256-fixed-wave-20260809-gfx1201-r9700, host hiptrx, corpus manifest 45b76818319f\nnot measured: peak VRAM (not instrumented this campaign)","engineFlags":{"commandSnippet":"scripts/lmx_continuous_batch.py --model /home/kaden/.hipfire/models/lfm2.5-350m.mq4 --prompt-file benchmarks/prompts/sweep/longform.txt --batch-size 256 --requests 256 --runs 3 --max-tokens 128 --max-seq 2048 --port 18815 --home-root /tmp/lmx_cb_home5 --log-dir /tmp/lmx_cb_logs5 --out /tmp/lmx_cb350_b256.json --device 0","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"LiquidAI/LFM2.5-350M","displayName":"LFM2.5-350M","family":null,"params":0,"activeParams":null,"isMoE":false,"baseModel":{"hfId":"LiquidAI/LFM2.5-350M-Base","displayName":"LFM2.5-350M-Base","params":0,"activeParams":null,"isMoE":false}},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"Radeon AI Pro R9700","gpuCount":1,"vramGb":32,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen Threadripper 9970X","os":"Ubuntu 26.04","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"hipfire","engineVersion":"0.3.0+439233d92959","quantization":"MQ4","backend":"rocm"},"user":{"id":"cmoeye1gq0000le04ie2kqs58","username":"schuttdev","verified":false,"verifiedAt":null,"pro":true},"hardwareGroupLabel":"Radeon AI Pro R9700","hardwareGroupKey":"DISCRETE_GPU:radeon ai pro r9700","rank":34,"reactionCounts":{},"myEmoji":null},{"id":"cmrk3vsd802zpmj01go7y8e39","modelRevision":"main","promptTokens":16864,"outputTokens":8192,"contextLength":32768,"prefillTokens":null,"batchSize":32,"ttftMs":34.02,"tokSOut":1288.275,"tokSPrefill":null,"tokSTotal":3940.309,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-07-14T03:42:58.652Z","notes":"Rented JarvisLabs RTX PRO 6000 Blackwell 96GB (cloud), CUDA 13 build sm_120. Median of 3 runs, temp 0, cold prefill, prompt ~16864 tok. Head-to-head vs my 3x R9700 rig, same GGUF quants.","engineFlags":{"commandSnippet":"vllm serve cyankiwi/GLM-4.5-Air-AWQ-4bit --max-model-len 32768 --max-num-seqs 96 --enable-chunked-prefill --enable-prefix-caching --gpu-memory-utilization 0.90","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"zai-org/GLM-4.5-Air","displayName":"GLM-4.5-Air","family":null,"params":110,"activeParams":null,"isMoE":true,"baseModel":null},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"RTX PRO 6000 Blackwell","gpuCount":1,"vramGb":96,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":null,"os":null,"isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"vllm","engineVersion":"vllm-0.25.0","quantization":"AWQ-4bit","backend":"cuda"},"user":{"id":"cmorcd2jg0008l804hfkput08","username":"1337Hero","verified":false,"verifiedAt":null,"pro":false},"hardwareGroupLabel":"RTX PRO 6000 Blackwell","hardwareGroupKey":"DISCRETE_GPU:rtx pro 6000 blackwell","rank":35,"reactionCounts":{},"myEmoji":null},{"id":"cmsnp2fk000lfo001s09siyds","modelRevision":"main","promptTokens":27,"outputTokens":128,"contextLength":4096,"prefillTokens":null,"batchSize":128,"ttftMs":2246,"tokSOut":1278.939,"tokSPrefill":null,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-10T20:39:01.440Z","notes":"BATCHED AGGREGATE THROUGHPUT — not single-stream decode.\nbatch=128, concurrency=128, 128 requests/run, 3 runs, mode=fixed_wave, max_tokens=128, ctx=4096, KV=default (unspecified, hipfire fp16), greedy; TTFT is under load.\nengine: hipfire 0.3.0+439233d92959, backend rocm, gfx1201\nquant: MQ4 = MagnumQuant 4-bit, group-256, FWHT-rotated, 5.18 bpw effective (229474032 bytes *8 / 354483968 params)\nmodel file: lfm2.5-350m.mq4 sha256 4885d1cecbad59d7\nprompt: longform.txt md5 e6f8ee7d66b934c49493347500f48083, 27 prompt tokens\nmethod: median of 3 runs, fresh process, temperature 0 (greedy)\ncaveats: pre-coalesce fixed-wave B128 1278.9 tok/s; opt batched-prefill-coalesced B128 per-card d0 3283.3 tok/s (spread includes d3 outlier 1609).\nevidence: sealed case lfm25-350m-b128-fixed-wave-20260809-gfx1201-r9700, host hiptrx, corpus manifest 45b76818319f\nnot measured: peak VRAM (not instrumented this campaign)","engineFlags":{"commandSnippet":"scripts/lmx_continuous_batch.py --model /home/kaden/.hipfire/models/lfm2.5-350m.mq4 --prompt-file benchmarks/prompts/sweep/longform.txt --batch-size 128 --requests 128 --runs 3 --max-tokens 128 --max-seq 2048 --port 18814 --home-root /tmp/lmx_cb_home4 --log-dir /tmp/lmx_cb_logs4 --out /tmp/lmx_cb350_b128.json --device 0","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"LiquidAI/LFM2.5-350M","displayName":"LFM2.5-350M","family":null,"params":0,"activeParams":null,"isMoE":false,"baseModel":{"hfId":"LiquidAI/LFM2.5-350M-Base","displayName":"LFM2.5-350M-Base","params":0,"activeParams":null,"isMoE":false}},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"Radeon AI Pro R9700","gpuCount":1,"vramGb":32,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen Threadripper 9970X","os":"Ubuntu 26.04","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"hipfire","engineVersion":"0.3.0+439233d92959","quantization":"MQ4","backend":"rocm"},"user":{"id":"cmoeye1gq0000le04ie2kqs58","username":"schuttdev","verified":false,"verifiedAt":null,"pro":true},"hardwareGroupLabel":"Radeon AI Pro R9700","hardwareGroupKey":"DISCRETE_GPU:radeon ai pro r9700","rank":36,"reactionCounts":{},"myEmoji":null},{"id":"cmsnp2gp300ljo001b1fr238f","modelRevision":"main","promptTokens":27,"outputTokens":128,"contextLength":4096,"prefillTokens":null,"batchSize":256,"ttftMs":4927.9,"tokSOut":1258.467,"tokSPrefill":null,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-10T20:39:02.919Z","notes":"BATCHED AGGREGATE THROUGHPUT — not single-stream decode.\nbatch=256, concurrency=256, 256 requests/run, 3 runs, mode=fixed_wave, max_tokens=128, ctx=4096, KV=default (unspecified, hipfire fp16), greedy; TTFT is under load.\nengine: hipfire 0.3.0+2ceddeb4f38d, backend rocm, gfx1201\nquant: MQ4 = MagnumQuant 4-bit, group-256, FWHT-rotated, 5.18 bpw effective (229474032 bytes *8 / 354483968 params)\nmodel file: lfm2.5-350m.mq4 sha256 4885d1cecbad59d7\nprompt: longform.txt md5 e6f8ee7d66b934c49493347500f48083, 27 prompt tokens\nmethod: median of 3 runs, fresh process, temperature 0 (greedy)\ncaveats: pre-coalesce baseline; coalesced build achieves 1876.9 tok/s (case lfm2.5-350m-gfx1201-b256r256-batched-prefill-fixed).\nevidence: sealed case lfm2.5-350m-gfx1201-b256r256-fixed, host k9lin, corpus manifest 45b76818319f\nnot measured: peak VRAM (not instrumented this campaign)","engineFlags":{"commandSnippet":"scripts/lmx_continuous_batch.py --model /home/kaden/.hipfire/models/lfm2.5-350m.mq4 --prompt-file benchmarks/prompts/sweep/longform.txt --batch-size 256 --requests 256 --runs 3 --max-tokens 128 --max-seq 4096 --port 11533 --home-root /home/kaden/localmaxxing-benchmarks/raw/lfm350-local-b256r256/home --log-dir /home/kaden/localmaxxing-benchmarks/raw/lfm350-local-b256r256/logs --out /home/kaden/localmaxxing-benchmarks/validation/lfm350-local-b256r256.json --cli /home/kaden/hipfire-lmx-lfm/target/release/hipfire --daemon /home/kaden/hipfire-lmx-lfm/target/release/examples/daemon --device 0","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"LiquidAI/LFM2.5-350M","displayName":"LFM2.5-350M","family":null,"params":0,"activeParams":null,"isMoE":false,"baseModel":{"hfId":"LiquidAI/LFM2.5-350M-Base","displayName":"LFM2.5-350M-Base","params":0,"activeParams":null,"isMoE":false}},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"RX 9070 XT","gpuCount":1,"vramGb":16,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen 9 3900X","os":"Ubuntu 24.04","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"hipfire","engineVersion":"0.3.0+2ceddeb4f38d","quantization":"MQ4","backend":"rocm"},"user":{"id":"cmoeye1gq0000le04ie2kqs58","username":"schuttdev","verified":false,"verifiedAt":null,"pro":true},"hardwareGroupLabel":"RX 9070 XT","hardwareGroupKey":"DISCRETE_GPU:rx 9070 xt","rank":37,"reactionCounts":{},"myEmoji":null},{"id":"cmsnp2hup00lmo001z7xdlo2y","modelRevision":"main","promptTokens":27,"outputTokens":128,"contextLength":4096,"prefillTokens":null,"batchSize":16,"ttftMs":302.9,"tokSOut":1246.758,"tokSPrefill":null,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-10T20:39:04.417Z","notes":"BATCHED AGGREGATE THROUGHPUT — not single-stream decode.\nbatch=16, concurrency=16, 32 requests/run, 3 runs, mode=continuous_refill, max_tokens=128, ctx=4096, KV=default (unspecified, hipfire fp16), greedy; TTFT is under load.\nengine: hipfire 0.3.0+5f11453f892f, backend rocm, gfx1201\nquant: MQ4 = MagnumQuant 4-bit, group-256, FWHT-rotated, 5.68 bpw effective (163173488 bytes *8 / 229693184 params)\nmodel file: lfm2.5-230m.mq4 sha256 3b92b9fd27c68d7f\nprompt: longform.txt md5 e6f8ee7d66b934c49493347500f48083, 27 prompt tokens\nmethod: median of 3 runs, fresh process, temperature 0 (greedy)\nevidence: sealed case lfm2.5-230m-gfx1201-b16r32-continuous, host k9lin, corpus manifest 45b76818319f\nnot measured: peak VRAM (not instrumented this campaign)","engineFlags":{"commandSnippet":"scripts/lmx_continuous_batch.py --model /home/kaden/.hipfire/models/lfm2.5-230m.mq4 --prompt-file benchmarks/prompts/sweep/longform.txt --batch-size 16 --requests 32 --runs 3 --max-tokens 128 --max-seq 4096 --port 11524 --home-root /home/kaden/localmaxxing-benchmarks/raw/lfm230-local-b16r32/home --log-dir /home/kaden/localmaxxing-benchmarks/raw/lfm230-local-b16r32/logs --out /home/kaden/localmaxxing-benchmarks/validation/lfm230-local-b16r32.json --cli /home/kaden/hipfire-lmx-lfm/target/release/hipfire --daemon /home/kaden/hipfire-lmx-lfm/target/release/examples/daemon --device 0","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"LiquidAI/LFM2.5-230M","displayName":"LFM2.5-230M","family":null,"params":0,"activeParams":null,"isMoE":false,"baseModel":{"hfId":"LiquidAI/LFM2.5-230M-Base","displayName":"LFM2.5-230M-Base","params":0,"activeParams":null,"isMoE":false}},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"RX 9070 XT","gpuCount":1,"vramGb":16,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen 9 3900X","os":"Ubuntu 24.04","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"hipfire","engineVersion":"0.3.0+5f11453f892f","quantization":"MQ4","backend":"rocm"},"user":{"id":"cmoeye1gq0000le04ie2kqs58","username":"schuttdev","verified":false,"verifiedAt":null,"pro":true},"hardwareGroupLabel":"RX 9070 XT","hardwareGroupKey":"DISCRETE_GPU:rx 9070 xt","rank":38,"reactionCounts":{},"myEmoji":null},{"id":"cmsnp2j0300lqo001fnw6dwfy","modelRevision":"main","promptTokens":27,"outputTokens":128,"contextLength":4096,"prefillTokens":null,"batchSize":128,"ttftMs":2456.3,"tokSOut":1237.749,"tokSPrefill":null,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-10T20:39:05.908Z","notes":"BATCHED AGGREGATE THROUGHPUT — not single-stream decode.\nbatch=128, concurrency=128, 128 requests/run, 3 runs, mode=fixed_wave, max_tokens=128, ctx=4096, KV=default (unspecified, hipfire fp16), greedy; TTFT is under load.\nengine: hipfire 0.3.0+e734d6c8561d, backend rocm, gfx1201\nquant: MQ4 = MagnumQuant 4-bit, group-256, FWHT-rotated, 5.18 bpw effective (229474032 bytes *8 / 354483968 params)\nmodel file: lfm2.5-350m.mq4 sha256 4885d1cecbad59d7\nprompt: longform.txt md5 e6f8ee7d66b934c49493347500f48083, 27 prompt tokens\nmethod: median of 3 runs, fresh process, temperature 0 (greedy)\ncaveats: pre-coalesce baseline; coalesced build achieves 3669.8 tok/s (case lfm2.5-350m-gfx1201-b128r128-batched-prefill-fixed).\nevidence: sealed case lfm2.5-350m-gfx1201-b128r128-fixed, host k9lin, corpus manifest 45b76818319f\nnot measured: peak VRAM (not instrumented this campaign)","engineFlags":{"commandSnippet":"scripts/lmx_continuous_batch.py --model /home/kaden/.hipfire/models/lfm2.5-350m.mq4 --prompt-file benchmarks/prompts/sweep/longform.txt --batch-size 128 --requests 128 --runs 3 --max-tokens 128 --max-seq 4096 --port 11532 --home-root /home/kaden/localmaxxing-benchmarks/raw/lfm350-local-b128r128/home --log-dir /home/kaden/localmaxxing-benchmarks/raw/lfm350-local-b128r128/logs --out /home/kaden/localmaxxing-benchmarks/validation/lfm350-local-b128r128.json --cli /home/kaden/hipfire-lmx-lfm/target/release/hipfire --daemon /home/kaden/hipfire-lmx-lfm/target/release/examples/daemon --device 0","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"LiquidAI/LFM2.5-350M","displayName":"LFM2.5-350M","family":null,"params":0,"activeParams":null,"isMoE":false,"baseModel":{"hfId":"LiquidAI/LFM2.5-350M-Base","displayName":"LFM2.5-350M-Base","params":0,"activeParams":null,"isMoE":false}},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"RX 9070 XT","gpuCount":1,"vramGb":16,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen 9 3900X","os":"Ubuntu 24.04","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"hipfire","engineVersion":"0.3.0+e734d6c8561d","quantization":"MQ4","backend":"rocm"},"user":{"id":"cmoeye1gq0000le04ie2kqs58","username":"schuttdev","verified":false,"verifiedAt":null,"pro":true},"hardwareGroupLabel":"RX 9070 XT","hardwareGroupKey":"DISCRETE_GPU:rx 9070 xt","rank":39,"reactionCounts":{},"myEmoji":null},{"id":"cmsnp1xcx00jto001lp2i7emk","modelRevision":"main","promptTokens":27,"outputTokens":128,"contextLength":4096,"prefillTokens":null,"batchSize":8,"ttftMs":114.4,"tokSOut":1181.636,"tokSPrefill":null,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-10T20:38:37.857Z","notes":"BATCHED AGGREGATE THROUGHPUT — not single-stream decode.\nbatch=8, concurrency=8, 16 requests/run, 3 runs, mode=continuous_refill, max_tokens=128, ctx=4096, KV=default (unspecified, hipfire fp16), greedy; TTFT is under load.\nengine: hipfire 0.3.0+b0bcc3f91506, backend rocm, gfx1151\nquant: MQ4 = MagnumQuant 4-bit, group-256, FWHT-rotated, 5.18 bpw effective (229474032 bytes *8 / 354483968 params)\nmodel file: lfm2.5-350m.mq4 sha256 4885d1cecbad59d7\nprompt: longform.txt md5 e6f8ee7d66b934c49493347500f48083, 27 prompt tokens\nmethod: median of 3 runs, fresh process, temperature 0 (greedy)\nevidence: sealed case lfm2.5-350m-gfx1151-b8r16-refill, host hipx, corpus manifest 45b76818319f\nnot measured: peak VRAM (not instrumented this campaign)","engineFlags":{"commandSnippet":"scripts/lmx_continuous_batch.py --model /home/kaden/.hipfire/models/lfm2.5-350m.mq4 --prompt-file benchmarks/prompts/sweep/longform.txt --batch-size 8 --requests 16 --runs 3 --max-tokens 128 --max-seq 4096 --port 18164 --home-root /home/kaden/lmx-final-runs/homes/hipx-1-batch-b8r16 --log-dir /home/kaden/lmx-final-runs/logs/FINAL/lfm350/gfx1151-continuous-b8r16 --out /home/kaden/lmx-final-runs/hipx-1-lfm350-b8r16-b8r16.json --device 1 --cli target/release/hipfire --daemon target/release/examples/daemon","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"LiquidAI/LFM2.5-350M","displayName":"LFM2.5-350M","family":null,"params":0,"activeParams":null,"isMoE":false,"baseModel":{"hfId":"LiquidAI/LFM2.5-350M-Base","displayName":"LFM2.5-350M-Base","params":0,"activeParams":null,"isMoE":false}},"hardware":{"hwClass":"UNIFIED","gpuName":null,"gpuCount":1,"vramGb":null,"chipVendor":"AMD","chipFamily":"Ryzen AI Max","chipVariant":"Ryzen AI Max 395+","unifiedMemoryGb":103,"cpu":"AMD Ryzen AI Max+ 395","os":"Ubuntu 26.04","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"hipfire","engineVersion":"0.3.0+b0bcc3f91506","quantization":"MQ4","backend":"rocm"},"user":{"id":"cmoeye1gq0000le04ie2kqs58","username":"schuttdev","verified":false,"verifiedAt":null,"pro":true},"hardwareGroupLabel":"Ryzen AI Max 395","hardwareGroupKey":"UNIFIED:ryzen ai max 395","rank":40,"reactionCounts":{},"myEmoji":null},{"id":"cmsiwwqmt00a9qm010iekvi3u","modelRevision":"main","promptTokens":0,"outputTokens":0,"contextLength":16384,"prefillTokens":null,"batchSize":64,"ttftMs":null,"tokSOut":1139.8,"tokSPrefill":8715,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-07T12:19:41.909Z","notes":"CONCURRENCY MAX: 64 concurrent users. Aggregate 1967.7 t/s wall-agg, generation 1139.8 t/s, diverse 512-token prompts, median TPOT 56ms. No-MTP native int4 v4 (MTP+concurrency blocked by GDN kernel). max-num-seqs 64. 165W. Multi-user aggregate throughput, not single-stream decode. 2026-08-07.","engineFlags":{"commandSnippet":"vllm serve /model --quantization gptq --dtype float16 --max-model-len 16384 --max-num-seqs 64 --max-num-batched-tokens 8192 --served-model-name Qwen3.6-35B-A3B-MTP-Preserved-GPTQ-Int4 --language-model-only","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":true,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"Qwen/Qwen3.6-35B-A3B","displayName":"Qwen3.6-35B-A3B","family":"Qwen","params":36,"activeParams":3,"isMoE":true,"baseModel":null},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"Intel Arc Pro B70","gpuCount":1,"vramGb":32,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen 7 7840HS w/ Radeon 780M Graphics","os":"Windows 11 Pro 10.0.26100 build 26100","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"vllm","engineVersion":"0.21.1.dev18+gec426505f XPU","quantization":"GPTQ-Int4","backend":"xpu"},"user":{"id":"cmr2iy4c506cupi01ltzdpbkf","username":"SergiioB","verified":false,"verifiedAt":null,"pro":false},"hardwareGroupLabel":"Intel Arc Pro B70","hardwareGroupKey":"DISCRETE_GPU:intel arc pro b70","rank":41,"reactionCounts":{"fire":1},"myEmoji":null},{"id":"cmsnsxbrr02fbo0012bo2krxv","modelRevision":"main","promptTokens":38,"outputTokens":128,"contextLength":4096,"prefillTokens":null,"batchSize":64,"ttftMs":628.5,"tokSOut":1115.74,"tokSPrefill":null,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-10T22:27:01.719Z","notes":"BATCHED AGGREGATE THROUGHPUT — not single-stream decode. batch=64, concurrency=64, 64 requests/run, 3 runs, continuous batching, max_tokens=128, KV=q8, greedy; TTFT measured under load.\nsingle R9700 (1 of 4 in the box; other cards ran unrelated jobs and were unused).\nengine: hipfire 0.3.0+b0bcc3f91506, backend rocm, gfx1201\nquant: MQ4R = MagnumQuant 4-bit Redline (MQ4 attention/router/shared + graded MQ4 routed experts; '.mq4r' also selects Redline retained-PM4 dispatch), 4.16 bpw effective.\nWORKLOAD (read before comparing): prompt merge_sort_thinking_off.txt md5 253c7ac50857fe6d0e10fb0d2c5e35c0, only 38 prompt tokens, 128 generated. The contextLength 4096 on this row is the max_seq ALLOCATION, not the prompt size. The Arc Pro B70 row on this page (1139.8 tok/s, concurrency 64) reports contextLength 16384, but that is likewise its --max-model-len; its own notes state diverse 512-token prompts. Its per-request prefill work is therefore ~13x ours, so this row is NOT like-for-like and should not be read as matching or beating it on equal workload.\nmethod: median of 3 runs; samples 1113.68 / 1115.74 / 1131.49 tok/s; per-lane 17.43 tok/s; median latency 7182.37 ms. All outputs byte-identical and coherent.\ncapacity: 64 lanes is the ceiling at max_seq 4096 on 32 GB; B96 and B128 both fail with 'hipMalloc: out of memory / continuous batch allocation failed'.\nevidence: 20260810T221534Z-a3b-mq4r-concurrency-ladder\nnot measured: peak VRAM.","engineFlags":{"commandSnippet":"python3 scripts/lmx_continuous_batch.py --model qwen3.6-35b-a3b.mq4r --prompt-file benchmarks/prompts/merge_sort_thinking_off.txt --batch-size 64 --requests 64 --runs 3 --max-tokens 128 --max-seq 4096 --device 1","tensorParallel":1,"gpuLayers":null,"kvCacheDtype":"q8","attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"Qwen/Qwen3.6-35B-A3B","displayName":"Qwen3.6-35B-A3B","family":"Qwen","params":36,"activeParams":3,"isMoE":true,"baseModel":null},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"Radeon AI Pro R9700","gpuCount":1,"vramGb":32,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen Threadripper 9970X","os":"Ubuntu 26.04","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"hipfire","engineVersion":"0.3.0+b0bcc3f91506","quantization":"MQ4R","backend":"rocm"},"user":{"id":"cmoeye1gq0000le04ie2kqs58","username":"schuttdev","verified":false,"verifiedAt":null,"pro":true},"hardwareGroupLabel":"Radeon AI Pro R9700","hardwareGroupKey":"DISCRETE_GPU:radeon ai pro r9700","rank":42,"reactionCounts":{},"myEmoji":null},{"id":"cmsxkikfy0btrms01s07oty3q","modelRevision":"main","promptTokens":512,"outputTokens":128,"contextLength":2048,"prefillTokens":null,"batchSize":1,"ttftMs":6.8,"tokSOut":1102.14,"tokSPrefill":75749.04,"tokSTotal":5207.6,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-17T18:29:17.950Z","notes":null,"engineFlags":{"commandSnippet":"E:\\llama.cpp-b10470\\llama-bench.exe -m E:\\huggingface\\hub\\models--unsloth--SmolLM2-135M-Instruct-GGUF\\snapshots\\9e6855bc4be717fca1ef21360a1db4b29d5c559a\\SmolLM2-135M-Instruct-Q2_K.gguf -p 512 -n 128 -fa on --prio 3 -r 10 --delay 2","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":true,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"unsloth/SmolLM2-135M-Instruct-GGUF","displayName":"SmolLM2-135M-Instruct-GGUF","family":"Llama","params":0,"activeParams":null,"isMoE":false,"baseModel":{"hfId":"HuggingFaceTB/SmolLM2-135M-Instruct","displayName":"SmolLM2-135M-Instruct","params":0,"activeParams":null,"isMoE":false}},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"RTX 5090","gpuCount":1,"vramGb":32,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"amd64","os":"windows","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"llama.cpp","engineVersion":null,"quantization":"Q2_K","backend":"cuda"},"user":{"id":"cmsxeur0y0b7xms01r3olg42y","username":"CsGoat","verified":false,"verifiedAt":null,"pro":false},"hardwareGroupLabel":"RTX 5090","hardwareGroupKey":"DISCRETE_GPU:rtx 5090","rank":43,"reactionCounts":{},"myEmoji":null},{"id":"cmrkrj5fl04twmj01a5s9onpv","modelRevision":"main","promptTokens":230,"outputTokens":128,"contextLength":512,"prefillTokens":null,"batchSize":64,"ttftMs":null,"tokSOut":1090.49,"tokSPrefill":1953.37,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-07-14T14:44:59.841Z","notes":"Warm steady-state CPU benchmark on SmolLM2-135M-Instruct Q2_K: 64 simultaneous HTTP requests, 64 server slots sharing eight physical CPU threads (CPUs 0-7), ubatch 1024, strict CPU and batch affinity. Two 64-request decode batches measured 1075.73 and 1105.25 aggregate tok/s; reported midpoint 1090.49 tok/s. Best prefill companion setting remains 32 slots/ubatch 1024/F16 K/V at 1953.37 aggregate prompt tok/s over 7360 prompt tokens. One warm-up batch was discarded. q8_0 K/V cache for decode; no GPU offload or SMT.","engineFlags":{"commandSnippet":"taskset -c 0-7 llama-server -m SmolLM2-135M-Instruct-Q2_K.gguf -c 512 -np 64 -t 8 -tb 8 -b 8192 -ub 1024 --poll 100 --cpu-range 0-7 --cpu-strict 1 --cpu-range-batch 0-7 --cpu-strict-batch 1 -ctk q8_0 -ctv q8_0 -fa auto","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":"q8_0","attentionBackend":null,"flashAttn":true,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"HuggingFaceTB/SmolLM2-135M-Instruct","displayName":"SmolLM2-135M-Instruct","family":"Llama","params":0,"activeParams":null,"isMoE":false,"baseModel":{"hfId":"HuggingFaceTB/SmolLM2-135M","displayName":"SmolLM2-135M","params":0,"activeParams":null,"isMoE":false}},"hardware":{"hwClass":"CPU_ONLY","gpuName":null,"gpuCount":1,"vramGb":null,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen 7 7840HS","os":"Ubuntu 24.04.4 LTS","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"llama.cpp","engineVersion":"commit 6b4dc21, AVX2 -Ofast Zen4-tuned no-repack server","quantization":"Q2_K","backend":"cpu"},"user":{"id":"cmoext1fr0000jl04skh3hjqf","username":"steveseguin","verified":false,"verifiedAt":null,"pro":false},"hardwareGroupLabel":"AMD Ryzen 7 7840HS","hardwareGroupKey":"CPU_ONLY:amd ryzen 7 7840hs","rank":44,"reactionCounts":{},"myEmoji":null},{"id":"cmozt9exg0085lo01c1g3q2i7","modelRevision":"main","promptTokens":200,"outputTokens":564,"contextLength":65536,"prefillTokens":null,"batchSize":16,"ttftMs":null,"tokSOut":1048,"tokSPrefill":null,"tokSTotal":1048,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-05-10T13:30:50.452Z","notes":"Concurrency=16 aggregate throughput, sweet-spot config. 16 parallel chat completions, 8 distinct prompts cycled with per-request seeding to defeat prefix caching, temp=0.7, max_tokens=600 (avg actual=564). Wall 8.62s, total 9,029 tokens, no queueing (max-num-seqs=16), p95 ≈ p50, 0 KV preemptions. Continuous batching only — Eagle3 spec-decode tested and dropped (27% accept on NVFP4 target → -29% throughput). Per-stream ~69 t/s under load vs 128 t/s single-stream. KV cache FP8. EPYC 32c/64t host, 160 GB DDR5. Run 2026-05-09.","engineFlags":{"commandSnippet":"vllm serve /opt/models/gemma-4-26B-A4B-it-NVFP4 --served-model-name gemma-4-26b-a4b-it --host 0.0.0.0 --port 8000 --max-model-len 65536 --gpu-memory-utilization 0.90 --max-num-seqs 16","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"nvidia/Gemma-4-26B-A4B-NVFP4","displayName":"Gemma-4-26B-A4B-NVFP4","family":"Gemma","params":14,"activeParams":null,"isMoE":false,"baseModel":{"hfId":"google/gemma-4-26B-A4B-it","displayName":"gemma-4-26B-A4B-it","params":27,"activeParams":4,"isMoE":true}},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"RTX PRO 6000 Blackwell","gpuCount":1,"vramGb":96,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":null,"os":null,"isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"vllm","engineVersion":"0.20.1","quantization":"NVFP4 (modelopt_fp4 W4A4)","backend":null},"user":{"id":"cmowc9fer00qdp101b007eug5","username":"murdarch","verified":false,"verifiedAt":null,"pro":false},"hardwareGroupLabel":"RTX PRO 6000 Blackwell","hardwareGroupKey":"DISCRETE_GPU:rtx pro 6000 blackwell","rank":45,"reactionCounts":{},"myEmoji":null},{"id":"cmsnp2k4r00lto001bsasg3ok","modelRevision":"main","promptTokens":27,"outputTokens":128,"contextLength":4096,"prefillTokens":null,"batchSize":4,"ttftMs":59.5,"tokSOut":1044.768,"tokSPrefill":null,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-10T20:39:07.371Z","notes":"BATCHED AGGREGATE THROUGHPUT — not single-stream decode.\nbatch=4, concurrency=4, 8 requests/run, 3 runs, mode=continuous, max_tokens=128, ctx=4096, KV=default (unspecified, hipfire fp16), greedy; TTFT is under load.\nengine: hipfire 0.3.0+b0bcc3f91506, backend rocm, gfx1201\nquant: MQ4 = MagnumQuant 4-bit, group-256, FWHT-rotated, 5.68 bpw effective (163173488 bytes *8 / 229693184 params)\nmodel file: lfm2.5-230m.mq4 sha256 3b92b9fd27c68d7f\nprompt: longform.txt md5 e6f8ee7d66b934c49493347500f48083, 27 prompt tokens\nmethod: median of 3 runs, fresh process, temperature 0 (greedy)\ncaveats: FINAL continuous numbers shown; earlier continuous_refill sweep measured lower (B4R8 737.9 vs 1044.8, B8R16 1023.2 vs 1704.6).\nevidence: sealed case lfm2.5-230m-gfx1201-b4r8-FINAL-continuous, host k9lin, corpus manifest 45b76818319f\nnot measured: peak VRAM (not instrumented this campaign)","engineFlags":{"commandSnippet":"scripts/lmx_continuous_batch.py --model /home/kaden/.hipfire/models/lfm2.5-230m.mq4 --prompt-file benchmarks/prompts/sweep/longform.txt --batch-size 4 --requests 8 --runs 3 --max-tokens 128 --max-seq 4096 --port 11604 --home-root /home/kaden/lmx-final-runs/work/FINAL/lfm2.5-230m/continuous/home-b4r8 --log-dir /home/kaden/lmx-final-runs/logs/FINAL/lfm2.5-230m/continuous/b4r8 --out /home/kaden/lmx-final-runs/lfm230-FINAL-b4r8-gfx1201-continuous.json --device 0","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"LiquidAI/LFM2.5-230M","displayName":"LFM2.5-230M","family":null,"params":0,"activeParams":null,"isMoE":false,"baseModel":{"hfId":"LiquidAI/LFM2.5-230M-Base","displayName":"LFM2.5-230M-Base","params":0,"activeParams":null,"isMoE":false}},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"RX 9070 XT","gpuCount":1,"vramGb":16,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen 9 3900X","os":"Ubuntu 24.04","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"hipfire","engineVersion":"0.3.0+b0bcc3f91506","quantization":"MQ4","backend":"rocm"},"user":{"id":"cmoeye1gq0000le04ie2kqs58","username":"schuttdev","verified":false,"verifiedAt":null,"pro":true},"hardwareGroupLabel":"RX 9070 XT","hardwareGroupKey":"DISCRETE_GPU:rx 9070 xt","rank":46,"reactionCounts":{},"myEmoji":null},{"id":"cmsnp2l9o00lwo001th6itv48","modelRevision":"main","promptTokens":27,"outputTokens":128,"contextLength":4096,"prefillTokens":null,"batchSize":4,"ttftMs":63.4,"tokSOut":1016.965,"tokSPrefill":null,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-10T20:39:08.844Z","notes":"BATCHED AGGREGATE THROUGHPUT — not single-stream decode.\nbatch=4, concurrency=4, 4 requests/run, 3 runs, mode=continuous, max_tokens=128, ctx=4096, KV=default (unspecified, hipfire fp16), greedy; TTFT is under load.\nengine: hipfire 0.3.0+b0bcc3f91506, backend rocm, gfx1201\nquant: MQ4 = MagnumQuant 4-bit, group-256, FWHT-rotated, 5.68 bpw effective (163173488 bytes *8 / 229693184 params)\nmodel file: lfm2.5-230m.mq4 sha256 3b92b9fd27c68d7f\nprompt: longform.txt md5 e6f8ee7d66b934c49493347500f48083, 27 prompt tokens\nmethod: median of 3 runs, fresh process, temperature 0 (greedy)\nevidence: sealed case lfm2.5-230m-gfx1201-b4r4-FINAL-continuous, host k9lin, corpus manifest 45b76818319f\nnot measured: peak VRAM (not instrumented this campaign)","engineFlags":{"commandSnippet":"scripts/lmx_continuous_batch.py --model /home/kaden/.hipfire/models/lfm2.5-230m.mq4 --prompt-file benchmarks/prompts/sweep/longform.txt --batch-size 4 --requests 4 --runs 3 --max-tokens 128 --max-seq 4096 --port 11603 --home-root /home/kaden/lmx-final-runs/work/FINAL/lfm2.5-230m/continuous/home-b4r4 --log-dir /home/kaden/lmx-final-runs/logs/FINAL/lfm2.5-230m/continuous/b4r4 --out /home/kaden/lmx-final-runs/lfm230-FINAL-b4r4-gfx1201-continuous.json --device 0","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"LiquidAI/LFM2.5-230M","displayName":"LFM2.5-230M","family":null,"params":0,"activeParams":null,"isMoE":false,"baseModel":{"hfId":"LiquidAI/LFM2.5-230M-Base","displayName":"LFM2.5-230M-Base","params":0,"activeParams":null,"isMoE":false}},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"RX 9070 XT","gpuCount":1,"vramGb":16,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen 9 3900X","os":"Ubuntu 24.04","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"hipfire","engineVersion":"0.3.0+b0bcc3f91506","quantization":"MQ4","backend":"rocm"},"user":{"id":"cmoeye1gq0000le04ie2kqs58","username":"schuttdev","verified":false,"verifiedAt":null,"pro":true},"hardwareGroupLabel":"RX 9070 XT","hardwareGroupKey":"DISCRETE_GPU:rx 9070 xt","rank":47,"reactionCounts":{},"myEmoji":null},{"id":"cmsxk9ynu0btams017h90ri7w","modelRevision":"main","promptTokens":512,"outputTokens":128,"contextLength":2048,"prefillTokens":null,"batchSize":1,"ttftMs":7.5,"tokSOut":1003.09,"tokSPrefill":68228.98,"tokSTotal":4736.9,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-17T18:22:36.474Z","notes":null,"engineFlags":{"commandSnippet":"E:\\llama.cpp-b10470\\llama-bench.exe -m E:\\huggingface\\hub\\models--unsloth--SmolLM2-135M-Instruct-GGUF\\snapshots\\9e6855bc4be717fca1ef21360a1db4b29d5c559a\\SmolLM2-135M-Instruct-Q4_K_M.gguf -p 512 -n 128","tensorParallel":null,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"unsloth/SmolLM2-135M-Instruct-GGUF","displayName":"SmolLM2-135M-Instruct-GGUF","family":"Llama","params":0,"activeParams":null,"isMoE":false,"baseModel":{"hfId":"HuggingFaceTB/SmolLM2-135M-Instruct","displayName":"SmolLM2-135M-Instruct","params":0,"activeParams":null,"isMoE":false}},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"RTX 5090","gpuCount":1,"vramGb":32,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"amd64","os":"windows","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"llama.cpp","engineVersion":null,"quantization":"Q4_K_M","backend":"cuda"},"user":{"id":"cmsxeur0y0b7xms01r3olg42y","username":"CsGoat","verified":false,"verifiedAt":null,"pro":false},"hardwareGroupLabel":"RTX 5090","hardwareGroupKey":"DISCRETE_GPU:rtx 5090","rank":48,"reactionCounts":{},"myEmoji":null},{"id":"cmot9gh4r000bib04esjf3r77","modelRevision":"main","promptTokens":589,"outputTokens":256,"contextLength":32768,"prefillTokens":null,"batchSize":16,"ttftMs":175.4,"tokSOut":991.1,"tokSPrefill":null,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-05-05T23:29:50.523Z","notes":"gpt-oss-20b MXFP4 concurrent throughput at batch=16 (matches --max-num-seqs 16). Aggregate output tok/s, best of 3 runs (991.1 / 986.8 / 990.6 -- variance under 0.5%). Per-request decode 64.8 tok/s. TTFT 175ms. This is vLLM's strength zone vs llama.cpp single-stream: ~6x throughput when serving multiple concurrent users. batch=1 single-stream was 48 tok/s (separate submission) -- same hardware, llama.cpp Q8_0 single-stream was 160 tok/s.","engineFlags":{"commandSnippet":"LD_LIBRARY_PATH=~/.local/lib/rccl-7.1.1:$LD_LIBRARY_PATH HIP_VISIBLE_DEVICES=0,1 VLLM_TARGET_DEVICE=rocm VLLM_ROCM_USE_AITER=0 FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE vllm serve openai/gpt-oss-20b --tensor-parallel-size 2 --dtype bfloat16 --max-model-len 32768 --max-num-seqs 16 --max-num-batched-tokens 4096 --enable-chunked-prefill --enable-prefix-caching --gpu-memory-utilization 0.90 --moe-backend triton --reasoning-parser openai_gptoss --tool-call-parser openai","tensorParallel":2,"gpuLayers":null,"kvCacheDtype":null,"attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"openai/gpt-oss-20b","displayName":"gpt-oss-20b","family":"Gpt","params":21,"activeParams":null,"isMoE":true,"baseModel":null},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"Radeon AI Pro R9700","gpuCount":2,"vramGb":64,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen 9 5950X","os":"Arch Linux","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"vllm","engineVersion":"0.20.1+rocm721","quantization":"MXFP4_MOE","backend":"rocm"},"user":{"id":"cmorcd2jg0008l804hfkput08","username":"1337Hero","verified":false,"verifiedAt":null,"pro":false},"hardwareGroupLabel":"Radeon AI Pro R9700","hardwareGroupKey":"DISCRETE_GPU:radeon ai pro r9700","rank":49,"reactionCounts":{},"myEmoji":null},{"id":"cmsnsxd8w02fgo001otjutgqo","modelRevision":"main","promptTokens":38,"outputTokens":128,"contextLength":4096,"prefillTokens":null,"batchSize":32,"ttftMs":519.15,"tokSOut":969.37,"tokSPrefill":null,"tokSTotal":null,"peakVramGb":null,"gpuPowerWatts":[],"totalPowerWatts":null,"hardwareCost":null,"createdAt":"2026-08-10T22:27:03.632Z","notes":"BATCHED AGGREGATE THROUGHPUT — not single-stream decode. batch=32, concurrency=32, 32 requests/run, 3 runs, continuous batching, max_tokens=128, KV=q8, greedy; TTFT measured under load.\nsingle R9700 (1 of 4 in the box; other cards ran unrelated jobs and were unused).\nengine: hipfire 0.3.0+b0bcc3f91506, backend rocm, gfx1201\nquant: MQ4R = MagnumQuant 4-bit Redline (MQ4 attention/router/shared + graded MQ4 routed experts; '.mq4r' also selects Redline retained-PM4 dispatch), 4.16 bpw effective.\nWORKLOAD (read before comparing): prompt merge_sort_thinking_off.txt md5 253c7ac50857fe6d0e10fb0d2c5e35c0, only 38 prompt tokens, 128 generated. The contextLength 4096 on this row is the max_seq ALLOCATION, not the prompt size. The Arc Pro B70 row on this page (1139.8 tok/s, concurrency 64) reports contextLength 16384, but that is likewise its --max-model-len; its own notes state diverse 512-token prompts. Its per-request prefill work is therefore ~13x ours, so this row is NOT like-for-like and should not be read as matching or beating it on equal workload.\nmethod: median of 3 runs; samples 968.61 / 969.37 / 971.82 tok/s; per-lane 30.29 tok/s; median latency 4171.65 ms. All outputs byte-identical and coherent.\ncapacity: 64 lanes is the ceiling at max_seq 4096 on 32 GB; B96 and B128 both fail with 'hipMalloc: out of memory / continuous batch allocation failed'.\nevidence: 20260810T221534Z-a3b-mq4r-concurrency-ladder\nnot measured: peak VRAM.","engineFlags":{"commandSnippet":"python3 scripts/lmx_continuous_batch.py --model qwen3.6-35b-a3b.mq4r --prompt-file benchmarks/prompts/merge_sort_thinking_off.txt --batch-size 32 --requests 32 --runs 3 --max-tokens 128 --max-seq 4096 --device 1","tensorParallel":1,"gpuLayers":null,"kvCacheDtype":"q8","attentionBackend":null,"flashAttn":null,"specDecoding":false,"mtpEnabled":false},"model":{"hfId":"Qwen/Qwen3.6-35B-A3B","displayName":"Qwen3.6-35B-A3B","family":"Qwen","params":36,"activeParams":3,"isMoE":true,"baseModel":null},"hardware":{"hwClass":"DISCRETE_GPU","gpuName":"Radeon AI Pro R9700","gpuCount":1,"vramGb":32,"chipVendor":null,"chipFamily":null,"chipVariant":null,"unifiedMemoryGb":null,"cpu":"AMD Ryzen Threadripper 9970X","os":"Ubuntu 26.04","isHeterogeneousGpu":false,"gpuSlots":[]},"engine":{"engineName":"hipfire","engineVersion":"0.3.0+b0bcc3f91506","quantization":"MQ4R","backend":"rocm"},"user":{"id":"cmoeye1gq0000le04ie2kqs58","username":"schuttdev","verified":false,"verifiedAt":null,"pro":true},"hardwareGroupLabel":"Radeon AI Pro R9700","hardwareGroupKey":"DISCRETE_GPU:radeon ai pro r9700","rank":50,"reactionCounts":{},"myEmoji":null}],"total":5805,"limit":50,"offset":0}