• Open Daily: 10am - 10pm
    Alley-side Pickup: 10am - 7pm

    3038 Hennepin Ave Minneapolis, MN
    612-822-4611

Open Daily: 10am - 10pm | Alley-side Pickup: 10am - 7pm
3038 Hennepin Ave Minneapolis, MN
612-822-4611
vLLM in Production: Running LLMs at Scale with GPUs, High-Performance Inference & Modern AI Infrastructure

vLLM in Production: Running LLMs at Scale with GPUs, High-Performance Inference & Modern AI Infrastructure

Paperback

General ComputersProgramming

Currently unavailable to order

ISBN13: 9798245694542
Publisher: Independently Published
Published: Jan 26 2026
Pages: 250
Weight: 1.30
Height: 0.53 Width: 8.50 Depth: 11.00
Language: English

LLM inference is no longer experimental-it is production infrastructure.
As models grow larger, applications become agent-driven, and real users arrive, the true bottleneck shifts from training to serving models reliably, securely, and at scale.

vLLM in Production is a hands-on, operator-first guide to running large language models in real environments-where GPUs are finite, latency matters, failures happen, and cost must be controlled.

Also in

Programming