nyra 727855c976 llama-server-35b: raise context 64K -> 128K
Measured in service: 128K decodes at 44.7 short / 42.9 after a 6.7K prompt,
against 44.5/43.1 at 64K, at the same VRAM. Doubling the window is free here.
256K costs ~30% with q8 V-cache but only ~13% with -ctv q4_0.

Corrects the earlier standalone llama-cli probe, which reported 18 tok/s at
128K and was wrong below 256K -- context cost must be measured in service.
2026-09-24 09:11:56 -04:00
2026-08-04 09:12:30 -04:00
S
Description
Utumno NixOS flake configuration
34 MiB
Languages
Nix 95%
Shell 3.5%
Python 1.5%