Yan Su
Tsinghua University ยท MSc candidate, computational law
Open source
Did the GPU actually run it? A low tok/s number cannot tell whether the model is genuinely slow or just did not fit on the GPU. picchio checks engine placement against the OS GPU meter and records the memory added by the run before you change the model or buy hardware. One tool, three agent frameworks dialects runs the same page_stats contract on LangChain, Dify and OpenClaw. Its 53 offline tests show where each framework validates inputs, carries configuration, reports failures and handles cancellation. Which firms actually win cases like this? dockets fits law-firm strength on 1,284 federal civil cases, then exposes the evidence counts that break the headline ranking. Every score drills to case IDs, with a randomized-outcome control and a reproducible win-rate evaluation. Every figure traces to a source row crosstalk turns MEBOCOST output into linked network, matrix, dot-plot and table views. Share the exact filter state by URL, then drill each plotted signal back to its input row.Engineering work
Agent systems Built an MCP server over stateless Streamable HTTP with typed admission errors, approval metadata, stage-level tracing, Swift and Python clients, and regression coverage for concurrency, head-of-line blocking and contract drift. Built tool-calling loops over real datastores with alias resolution, query segmentation and golden-question evaluation.
Local AI measurement Benchmarked llama.cpp, Ollama and MLX across quantization and CPU/GPU execution paths, including silent CPU fallback. Built on-device speech with Apple frameworks and a streaming profiler that separates latency by stage.
Retrieval and model evaluation Built BM25, TF-IDF and hybrid keyword-vector retrieval with Transformer encoders, bilingual tokenisation and freshness-stamped crawler pipelines. Ran SFT and RFT tuning loops with data cleaning and prompt strategy. Built golden-question evaluation sets for retrieval and agent workflows.
Desktop systems Implemented features across a 346-file Rust backend, Tauri and Swift bridges to native macOS frameworks. Shipped cross-platform builds and CI for macOS, Windows and Linux, plus Chrome and Safari extensions from one shared source.
Services and security Built TypeScript and Python services with Astro, React, FastAPI and Flask. Moved long-lived cloud keys off clients behind a token broker that issues short-lived credentials and presigned object URLs, with HMAC verification implemented using only the standard library.
Writing
A 200 OK does not mean your LLM used the GPU A low tok/s number cannot tell whether the model is slow or never made it onto the GPU. With 0/33 layers on the GPU, prefill fell 22x while the server kept returning 200. Who gets paid first in the AI boom? Chip vendors book the boom now. Software companies can book payroll savings now. The productivity gain can remain a forecast.Experience
Institute for Studies on AI and Law, Tsinghua Core member Built and maintain the knowledge graph behind the institute's legal knowledge engineering work. Research on legal NLP: Transformer models for reading statutory and case text, and how large models behave under compliance rules.
Baidu AI Product Manager, internship Led tuning iterations with SFT and RFT: prompt optimisation strategy and a data cleaning pipeline that raised job-to-candidate matching accuracy. Designed and shipped multi-scenario agent workflows for tasks the single-turn setup could not hold.
Mentougou District Procuratorate, Beijing Software Engineer, remote internship Shipped a hybrid BM25 and TF-IDF retrieval model into a district prosecutor's office, with a legal-specific ranking algorithm, the backend API and its tests. Optimised text vectorisation on a Transformer encoder.
Qiji Technology AI Product Manager Built the domain terminology base for sports and turned failure analysis into iteration priorities for the algorithm team. Built and launched the first version of the text-generation module, including the PRD and interaction prototypes.
Competitions
LexGraph Sole developer. Compiles criminal statute text into an executable rule graph. Each element of the offence becomes a True/False node; the graph returns conviction, sentencing band and suspended-sentence eligibility, and routes between three overlapping road-traffic offences when the facts fit a different one. Third prize at the first national Legal Rules Computer Expression Competition, 943 entrants from 201 universities.
Bar Exam Agent Sole developer. Built on GLM-4.5 over a knowledge base of 19,399 authoritative legal documents, and every answer cites the statute it rests on. Ranked 8th of 109 agents in public voting, and took third prize at Tsinghua's first AI Agent Innovation Competition, out of the 100 teams that cleared the midpoint review. Admissions Radar Built for Tsinghua's second AI Agent Innovation Competition. It creates a dated admissions timeline from published departmental notices, projects the next milestones from the previous cycle, links every claim to its source, and marks both source freshness and milestones with no published announcement. Education
Tsinghua University MSc candidate in Computational Law. Thesis on who owns the data inside an LLM knowledge base and who is liable when it infringes, read against the EU AI Act, GDPR and the DSM text-and-data-mining exception. Expected graduation: 2027.
Communication University of China BA in Translation with a minor in Computer Science.