▌ EXPERIMENT
Tesla V100 32GB Runs Qwen3.8-27B: 131k Context on a Single Card — Measured Benchmark
Benchmarking Qwen3.8-27B on a 2017 Tesla V100 32GB. 33.41 tok/s on Q4_K_XL, and a 131,072-token context fits in 25,264 MiB on one card. Plus the KV-quantization CPU fallback and why --reasoning-budget 1024 decides whether it is practical.











