Yan Su
Writing
A 200 OK does not mean your LLM used the GPU A low tok/s number cannot tell whether the model is slow or never made it onto the GPU. With 0/33 layers on the GPU, prefill fell 22x while the server kept returning 200.
[ SYS_TIME: 06.05.2026 ]
Who gets paid first in the AI boom? Chip vendors book the boom now. Software companies can book payroll savings now. The productivity gain can remain a forecast.
[ SYS_TIME: 05.11.2026 ]
© 2026 Yan Su
Writing RSS Feed