• Open Daily: 10am - 10pm
    Alley-side Pickup: 10am - 7pm

    3038 Hennepin Ave Minneapolis, MN
    612-822-4611

Open Daily: 10am - 10pm | Alley-side Pickup: 10am - 7pm
3038 Hennepin Ave Minneapolis, MN
612-822-4611
Vision Language Models: Building Vlms with Hugging Face

Vision Language Models: Building Vlms with Hugging Face

Paperback

General Computers

Publisher Price: $79.99

ISBN13: 9798341624047
Publisher: O'Reilly Media
Published: Jul 14 2026
Pages: 406
Weight: 1.43
Height: 0.84 Width: 7.00 Depth: 9.19
Language: English

Vision language models (VLMs) combine computer vision and natural language processing to create powerful systems that can interpret, generate, and respond in multimodal contexts. Vision Language Models is a hands-on guide to building real-world VLMs using the most up-to-date stack of machine learning tools from Hugging Face, Meta (PyTorch), NVIDIA (Cuda), and others, written by leading researchers and practitioners Merve Noyan, Miquel Farré, Andrés Marafioti, and Orr Zohar. From image captioning and document understanding to advanced zero-shot inference and retrieval-augmented generation (RAG), this book covers the full VLM application and development lifecycle.

Also in

General Computers