Qwen3.8-27B long-context test on dual RTX 3060 12GB
Qwen3.8-27BQwen3.8-27B on two RTX 3060 12GB cards, llama.cpp 0.4.1-dev. Measured 129,000-token synthetic prompt at 497.8 prompt tok/s and 42.7 output tok/s mean, with exact recall of 3/3 canary facts in two runs. Configured context 131,072; ngram-mod + MTP3; batch 2048, ubatch 512.
