727855c976fed67baed7cf0bceeb79ca83948e57
Measured in service: 128K decodes at 44.7 short / 42.9 after a 6.7K prompt, against 44.5/43.1 at 64K, at the same VRAM. Doubling the window is free here. 256K costs ~30% with q8 V-cache but only ~13% with -ctv q4_0. Corrects the earlier standalone llama-cli probe, which reported 18 tok/s at 128K and was wrong below 256K -- context cost must be measured in service.
Description
Utumno NixOS flake configuration
34 MiB
Languages
Nix
95%
Shell
3.5%
Python
1.5%