<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Yan Su</title><description>AI agents, tool use and RAG. Local LLM inference and retrieval. Computational law at Tsinghua.</description><link>https://yansu.me/</link><item><title>A 200 OK does not mean your LLM used the GPU</title><link>https://yansu.me/writing/a-200-ok-does-not-mean-gpu/</link><guid isPermaLink="true">https://yansu.me/writing/a-200-ok-does-not-mean-gpu/</guid><description>A low tok/s number cannot tell whether the model is slow or never made it onto the GPU. With 0/33 layers on the GPU, prefill fell 22x while the server kept returning 200.</description><pubDate>Sat, 06 Jun 2026 00:00:00 GMT</pubDate></item><item><title>Who gets paid first in the AI boom?</title><link>https://yansu.me/writing/ai-capital-flow-chips-eat-saas-starves/</link><guid isPermaLink="true">https://yansu.me/writing/ai-capital-flow-chips-eat-saas-starves/</guid><description>Chip vendors book the boom now. Software companies can book payroll savings now. The productivity gain can remain a forecast.</description><pubDate>Tue, 12 May 2026 00:00:00 GMT</pubDate></item></channel></rss>