Author

Jungjun Hur 허정준

I write about how LLM inference systems actually work. The method is to rebuild them one mechanism at a time, starting from what breaks without that mechanism, and to check every piece against how the real engines do it: vLLM, SGLang, and TensorRT-LLM. Nothing here is a design I invented because it made a point easier to explain.

If something here is wrong, I would rather hear it than not. Corrections, questions and requests are welcome by email.

Books