← All courses

Transformers & LLMs

Build a real understanding of the transformer architecture, then train and fine-tune LLMs — theory and practice.

Chapter 01

The Transformer Architecture

  1. The Big Picture: What a Transformer DoesComing soon
  2. Tokenization: Text → IntegersComing soon
  3. Embeddings: Tokens → VectorsComing soon
  4. Positional Information: Order MattersComing soon
  5. Attention I: Query, Key, ValueComing soon
  6. Attention II: Scale, Mask, Softmax → WeightsComing soon
  7. Multi-Head AttentionComing soon
  8. The Feed-Forward Network (MLP)Coming soon
  9. Residual Connections & NormalizationComing soon
  10. The Full Transformer BlockComing soon
  11. The Output Head: Logits → Next TokenComing soon
  12. Putting It Together: A Forward Pass, End to EndComing soon
  13. Chapter examComing soon
Chapter 02

Building & Training a Small LLM

  1. Next-Token Prediction as ClassificationComing soon
  2. The Data Pipeline: Text → BatchesComing soon
  3. Backprop Through the TransformerComing soon
  4. The Optimizer: AdamW in PracticeComing soon
  5. Initialization & Numerical StabilityComing soon
  6. The Training LoopComing soon
  7. Watching a Model LearnComing soon
  8. Sampling & Generation RevisitedComing soon
  9. Scaling Laws: Predicting Loss Before You TrainComing soon
  10. Capstone: Your Tiny LLMComing soon
  11. Chapter examComing soon
Chapter 03

Fine-Tuning

  1. Why Fine-Tune?Coming soon
  2. Full Fine-TuningComing soon
  3. PEFT: The IdeaComing soon
  4. LoRA: Low-Rank AdaptationComing soon
  5. QLoRA & Quantization (intro)Coming soon
  6. Instruction Tuning & Chat TemplatesComing soon
  7. Preference Data & Reward ModelingComing soon
  8. RLHF, DPO & GRPO in PracticeComing soon
  9. Fine-Tuning Gemma / Qwen (capstone)Coming soon
  10. Evaluation: Did It Work?Coming soon
  11. Chapter examComing soon
Chapter 04

Modern Architectures

  1. RoPE: Rotary Position EmbeddingsComing soon
  2. Grouped-Query & Multi-Query AttentionComing soon
  3. The KV Cache & InferenceComing soon
  4. FlashAttentionComing soon
  5. Mixture of ExpertsComing soon
  6. DeepSeek MLA: Low-Rank KV CompressionComing soon
  7. Multi-Token PredictionComing soon
  8. Quantization: int8, int4, and BeyondComing soon
  9. State Space Models & MambaComing soon
  10. Speculative DecodingComing soon
  11. The Modern Stack AssembledComing soon
  12. Chapter examComing soon