Small Language Models for Efficient AI A Practical Guide to Training, Serving, and Scaling SLMs
Reliable shipping
Flexible returns
Small Language Models for Efficient AI
A Practical Guide to Training, Serving, and Scaling SLMs
Santanu Bhattacharjee
This book explains implementing Small Language Models (SLMs) with a practical, production-minded approach that goes beyond theory. This book gives AI engineers, forward deployed engineers and applied scientists a clear path to building, tuning, evaluating, and deploying compact language models that are efficient, adaptable, and ready for real-world use.
The book begins with the foundations needed to work confidently with SLMs, including model architecture, tokenization, training data design, compute constraints, and the key implementation choices that shape model quality and cost. It then moves into the full development lifecycle: preparing datasets, selecting model sizes, training from scratch, fine-tuning with supervised and preference-based methods, optimizing inference with quantization and caching, and designing evaluation workflows that measure both model quality and operational efficiency. The book also covers domain adaptation, retrieval-augmented patterns, tool use, agentic workflows, and deployment strategies for enterprise and edge settings.
Throughout, the content is organized around practical decisions, reproducible workflows, and code-driven examples that readers can apply directly in their own work.
By the end of the book, readers will be able to build and deploy SLM systems with a strong grasp of both the technical foundations and the production trade-offs. They will know how to choose the right approach for a given use case and how to design compact AI systems that are cost-effective, reliable, and future-ready.
What you will learn:
Designing and building small language models from the ground up, covering architecture, tokenization, data prep, and training workflows.
Fine-tuning SLMs for targeted tasks and domains using supervised tuning, preference optimization, and practical adaptation methods.
Optimizing SLM inference through quantization, caching strategies, and deployment-ready configurations for real-world performance.
Evaluating SLMs with both model quality metrics and system-level measures like latency, cost, memory use, and reliability.
Applying SLMs in production environments, including retrieval-augmented pipelines, tool use, agentic systems, and edge or enterprise deployments.
Who this book is for:
This book is for AI engineers and applied scientists developing compact language models for production use. It also serves practitioners and technical teams deploying efficient, cost‑effective SLM solutions across enterprise and edge environments.
Santanu Bhattacharjee is Director of AI and Product Engineering at Pretium, with more than 15 years of experience in artificial intelligence, machine learning, and product engineering. Previously, he served as Senior Chief Engineer at Samsung Research, where he contributed to the research and development of advanced AI systems. His work spans building small language models for agentic systems, deploying SLMs on edge devices, and integrating language models with knowledge graphs. He has deep expertise in model distillation, SLM development, agentic AI systems, and scaling AI solutions in production. Santanu holds a BE from Jadavpur University and an MS from Dublin City University. His research includes patents and publications on episodic knowledge graph construction for edge devices, natural language querying for big data platforms. He has received multiple recognitions as an AI engineering leader.
| Publication Date: | 12 March 2027 |
| Publisher: | Apress |
| Imprint: | Apress |
| ISBN-13: | 9798868833618 |
| Format: | Paperback softback |