• Open Daily: 10am - 10pm
    Alley-side Pickup: 10am - 7pm

    3038 Hennepin Ave Minneapolis, MN
    612-822-4611

Open Daily: 10am - 10pm | Alley-side Pickup: 10am - 7pm
3038 Hennepin Ave Minneapolis, MN
612-822-4611
AI Performance Engineering: From GPU Kernels to LLM Inference

AI Performance Engineering: From GPU Kernels to LLM Inference

Paperback

General Computers

Currently unavailable to order

ISBN13: 9798198692480
Publisher: Independently Published
Pages: 340
Weight: 1.30
Height: 0.71 Width: 7.00 Depth: 10.00
Language: English

A hands-on guide to making AI systems fast - from GPU kernels to production LLM inference.

Most AI systems run well below the speed their hardware allows - GPUs idle waiting on data, LLMs serve a fraction of their throughput, and adding hardware sometimes makes things slower. AI Performance Engineering: From GPU Kernels to LLM Inference is a practitioner's guide to diagnosing, profiling, and fixing those bottlenecks - systematically, with real tools and runnable code, from hardware first principles to production LLM serving.

Also from

Bittla, Srinivasa Rao

Also in

General Computers