- Home
- Computers
- Intelligence (AI) & Semantics
- Quantization and Fast Inference (A practitioner's guide to efficient AI)
Quantization and Fast Inference (A practitioner's guide to efficient AI)
List Price:
$59.99
| Expected release date is Dec 29th 2026 |
- Availability: Confirm prior to ordering
- Branding: minimum 50 pieces (add’l costs below)
- Check Freight Rates (branded products only)
Branding Options (v), Availability & Lead Times
- 1-Color Imprint: $2.00 ea.
- Promo-Page Insert: $2.50 ea. (full-color printed, single-sided page)
- Belly-Band Wrap: $2.50 ea. (full-color printed)
- Set-Up Charge: $45 per decoration
- Availability: Product availability changes daily, so please confirm your quantity is available prior to placing an order.
- Branded Products: allow 10 business days from proof approval for production. Branding options may be limited or unavailable based on product design or cover artwork.
- Unbranded Products: allow 3-5 business days for shipping. All Unbranded items receive FREE ground shipping in the US. Inquire for international shipping.
- RETURNS/CANCELLATIONS: All orders, branded or unbranded, are NON-CANCELLABLE and NON-RETURNABLE once a purchase order has been received.
Product Details
Author:
Vivek Kalyanarangan
Format:
Paperback
Pages:
350
Publisher:
Manning (December 29, 2026)
Imprint:
Manning
Release Date:
December 29, 2026
Language:
English
ISBN-13:
9781633433915
ISBN-10:
1633433919
Weight:
14.78oz
Dimensions:
7.375" x 9.25"
File:
Eloquence-SimonSchuster_09032026_P10573212_onix30_Complete-20260903.xml
Folder:
Eloquence
List Price:
$59.99
Pub Discount:
37
As low as:
$56.99
Publisher Identifier:
P-SS
Discount Code:
H
Overview
Today's AI models demand a lot of memory, compute, and server horsepower—which quickly translates into cost. This book show you how you can optimize AI models without architectural redesigns or task-specific compression. It reveals practical techniques for quantization, systematically reducing numerical precision to achieve faster inference, lower memory usage, and cheaper deployment—all with minimal accuracy loss.
From quantization fundamentals to runtime packaging, the book gives you a complete and comprehensive overview of the full quantization pipeline. It starts by deriving quantization mapping from first principles, and then builds your knowledge and skill through techniques for production-tested PTQ and QAT workflows and a fully compressed deployment. You'll learn to apply post-training quantization to production models, run quantization-aware training using fake quantization and straight-through estimators, and handle subtle tradeoffs like activation outliers in LLMs, KV cache pressure, and sub-8-bit formats like NF4 and FP4.
What's inside
• Applying post-training quantization to production models
• Deploying efficiently on CPUs, edge devices, and mobile
• Framework-agnostic techniques and real cross-framework parity testing
• Flowcharts and checklists for efficient decision making
About the reader
For ML engineers and researchers experienced in Python.
About the author
Vivek Kalyanarangan is an AI/ML architect, researcher, and educator with over twelve years of experience designing and deploying large-scale machine learning systems.









