Book detail
Library

Cuantum trackFull access
Under the Hood of Large Language Models
8 chapters and 47 canonical sections synced from the Cuantum content database.
Author
Cuantum Tech.
Chapters
8
Reading time
~ 22h
Level
Professional
Language
English
Edition
2025
Your progress0%
Chapters & sections
8 chapters - 47 sectionsChapter 1: What Are LLMs? From Transformers to Titans
0/5Chapter 2: Tokenization and Embeddings
0/5Chapter 3: Anatomy of an LLM
0/5Chapter 4: Training LLMs from Scratch
0/64.1 Data Collection, Cleaning, Deduplication, and Filtering136m4.2 Curriculum Learning, Mixture Datasets, and Synthetic Data44m4.3 Infrastructure: Distributed Training, GPUs vs TPUs vs Accelerators85m4.4 Cost Optimization & Sustainability in Large-Scale Training76mChapter 4 Summary – Training LLMs from Scratch3mPractical Exercises – Chapter 46m
Chapter 5: Beyond Text: Multimodal LLMs
0/5Quiz
0/2Project 1: Build a Toy Transformer from Scratch in PyTorch
0/8Project 2: Train a Custom Domain-Specific Tokenizer (e.g., for legal or medical texts)
0/110. Setup3m1. Gather a Representative Mini-Corpus3m2. Train a BPE Tokenizer (🤗 tokenizers)6m3. Train a SentencePiece Tokenizer (Unigram or BPE)5m4. Wrap Your Tokenizer for Transformers3m5. Evaluate Tokenizer Quality5m6. Add a User Vocabulary (optional but powerful)4m7. Save, Load, and Version3m8. Plug Into a Small Model (sanity run)4mPitfalls & Tips3mLearning outcomes1m