Author
Jungjun Hur 허정준
I write about how LLM inference systems actually work. The method is to rebuild them one mechanism at a time, starting from what breaks without that mechanism, and to check every piece against how the real engines do it: vLLM, SGLang, and TensorRT-LLM. Nothing here is a design I invented because it made a point easier to explain.
If something here is wrong, I would rather hear it than not. Corrections, questions and requests are welcome by email.