• Open Daily: 10am - 10pm
    Alley-side Pickup: 10am - 7pm

    3038 Hennepin Ave Minneapolis, MN
    612-822-4611

Open Daily: 10am - 10pm | Alley-side Pickup: 10am - 7pm
3038 Hennepin Ave Minneapolis, MN
612-822-4611
Hands-On LLM Serving and Optimization: Hosting Llms at Scale

Hands-On LLM Serving and Optimization: Hosting Llms at Scale

Paperback

General ComputersProgramming

Publisher Price: $79.99

ISBN13: 9798341621497
Publisher: O'Reilly Media
Published: Jun 2 2026
Pages: 371
Weight: 1.31
Height: 0.77 Width: 7.00 Depth: 9.19
Language: English

Large language models (LLMs) are the reasoning engines of modern AI. Today, a major inflection point has arrived: as the world races to deploy AI at scale, model inference has moved to the center of the stack. Welcome to the inference era.

Without proper optimization, however, LLMs can be expensive and slow to serve. Hands-On LLM Serving and Optimization is a comprehensive guide to the complexities of deploying and optimizing LLMs at scale.

Also from

Wang, Chi

Also in

General Computers