Tesla V100 32GB Runs Qwen3.8-27B: 131k Context on a Single Card — Measured Benchmark
Benchmarking Qwen3.8-27B on a 2017 Tesla V100 32GB. 33.41 tok/s on Q4_K_XL, and a 131,072-token context fits in 25,264 MiB on one card. Plus the KV-quantization CPU fallback and why --reasoning-budget 1024 decides whether it is practical.







![[Solved] Deleting the ‘nul’ File from Claude Code on Windows](https://miyagadget.page/wp-content/uploads/2026/01/無題.jpg)

![[Flask + WebAuthn] Building a Mobile-Friendly Household Budget Web App with Passkey Authentication](https://miyagadget.page/wp-content/uploads/2025/11/ChatGPT-Image-2025年11月23日-10_41_02.png)


