• Open Daily: 10am - 10pm
    Alley-side Pickup: 10am - 7pm

    3038 Hennepin Ave Minneapolis, MN
    612-822-4611

Open Daily: 10am - 10pm | Alley-side Pickup: 10am - 7pm
3038 Hennepin Ave Minneapolis, MN
612-822-4611
Deep Dive into SGLang, Volume II: Quantization, Distributed Serving, Kernels, and Measurement

Deep Dive into SGLang, Volume II: Quantization, Distributed Serving, Kernels, and Measurement

Paperback

Series: Foundation Books

General Computers

ISBN13: 9798183580419
Publisher: Independently Published
Published: Jun 21 2026
Pages: 530
Weight: 2.68
Height: 1.07 Width: 8.50 Depth: 11.00
Language: English

Deep Dive into SGLang, Volume II continues the source-guided explanation of SGLang's inference runtime where the core request path meets deployment pressure.

This volume follows the serving contracts introduced in Volume I into quantized weights, LoRA adapters, Mixture-of-Experts routing, distributed placement, collectives, prefill/decode disaggregation, load balancing, custom kernels, CUDA graphs, hardware backends, multimodal transformer serving, diffusion inference, benchmarking, correctness testing, and technical extension work.

Also in

General Computers