• Open Daily: 10am - 10pm
    Alley-side Pickup: 10am - 7pm

    3038 Hennepin Ave Minneapolis, MN
    612-822-4611

Open Daily: 10am - 10pm | Alley-side Pickup: 10am - 7pm
3038 Hennepin Ave Minneapolis, MN
612-822-4611
Inside LLM Inference: From Silicon to Served Token - A Measured Guide to How Language Models Run on GPUs

Inside LLM Inference: From Silicon to Served Token - A Measured Guide to How Language Models Run on GPUs

Paperback

Series: The AI Singularity

General Computers

ISBN13: 9798187180769
Publisher: Independently Published
Published: Jul 13 2026
Pages: 180
Weight: 0.71
Height: 0.38 Width: 7.00 Depth: 10.00
Language: English
How does a language model actually become served tokens on a GPU - and why is it so much slower than the hardware's headline TFLOPS promise?

Inside LLM Inference follows a single request through the entire stack: from a PyTorch call down through the CUDA runtime, the streaming multiprocessor, the warp, and the memory system, to the arithmetic cores that finally do the work - and back up to the throughput and cost you pay for.

Every claim is measured, not asserted. A production model runs on 2026's newest silicon - NVIDIA's Blackwell generation - across four inference engines (vLLM, SGLang, TensorRT-LLM, and Hugging Face Transformers), profiled to individual GPU kernels with hardware counters and cross-checked against the physics that predicts them.

You start from zero - Chapter 0 defines every term - and finish able to read a live server:

  • classify a kernel as memory- or compute-bound, and prove it with a clock-sensitivity test
  • size a KV-cache pool and the batch it admits from a model's own dimensions
  • read sm_active and sm_occupancy off a running GPU and know what they mean
  • choose an inference engine on evidence, not vibes

Also in

General Computers