Skip to content

Study Plans

Structured schedules for self-paced learning. Pick the plan that matches your goals. For a degree-shaped sequence with assessment and specialization, use the CS curriculum. These shorter schedules are introductory passes or refreshers.

Build Your Own LLM System (about 67 weeks)

Section titled “Build Your Own LLM System (about 67 weeks)”

Follow the course path at roughly 10 to 12 hours per week. Each pass ends at a runnable milestone. P0 to P11 are staged; P12 is a design plan whose modules and implementation are not built yet.

PassFocusPathApprox. time
0Repository, Python, shell, Git, Make, and CISetup1.5 weeks
1Byte bigram, Rust engine, Go gateway, and deploymentTracer5 weeks
2Autograd and training foundationsFoundations7.5 weeks
3Tokenizers and dataTokens and data7.5 weeks
4Recurrent and sequence-to-sequence modelsSequence models4 weeks
5Transformer models and fine-tuningTransformer6 weeks
6Python inference and optional standalone C exercisesInference and kernels7 weeks
7Serving platform and observabilityServing platform8 weeks
8Durable data and training workflowsDurable4.5 weeks
9Training and evaluating the capstone modelCapstone training5 weeks
10Agents, retrieval, and usage policyAgents5 weeks
11Operability, security, and release readinessOperate5.5 weeks
12Planned image and audio extensionsMultimodal planNot estimated

Python and Rust exchange model and tokenizer data through files; the Rust engine uses candle. C modules are optional standalone exercises and are not required by core pass gates.

~10-12 hours per active subject per week, or ~20-24 hours while two subjects are paired. At 10-12 hours total, take subjects sequentially and extend the schedule.

WeekCalculus 1Linear Algebra
1Limits and continuityLinear systems, Gauss’s method
2Derivatives (rules and computation)Vector spaces, subspaces
3Applications of derivatives (optimization, MVT)Linear transformations, matrices
4Integration and FTCDeterminants, eigenvalues intro
WeekCalculus 2Discrete Math 1
5Integration techniques (by parts, partial fractions)Sets, logic, truth tables
6Applications of integrationDirect proof, contrapositive
7Sequences and convergence testsProof by contradiction, induction
8Power series, Taylor seriesRelations, functions, cardinality
WeekCalculus 3Discrete Math 2
9Vectors in space, dot/cross productGraph theory fundamentals
10Partial derivatives, gradientGraph coloring, Euler/Hamilton
11Multiple integralsNumber theory, modular arithmetic
12Vector calculus (Green’s, Stokes’)Recurrence relations, generating functions

Revisit the hardest topics from each subject. Work through challenge problems.

WeekFocus
14Probability axioms, conditional probability, Bayes’ theorem, discrete distributions
15Continuous distributions, joint distributions, expectation & variance
16Law of large numbers, CLT, hypothesis testing, confidence intervals, regression
  1. Read (30 min) — textbook sections for the day’s concept
  2. Work problems (60-90 min) — essential problems first, then challenge problems
  3. Review (30 min) — revisit previous day’s concepts, rework one problem from memory
  4. Connect (15 min) — read the “Connections to CS” section, think about how the math applies

~2-3 hours per day. Each day targets specific patterns and problems.

DayFocusProblemsPatterns
1Arrays & HashingTwo Sum, Contains Duplicate, Missing Number, Product Except SelfHash map complement lookup, frequency counting, index-as-hash
2Two Pointers & Sliding Window3Sum, Container With Most WaterOpposite-end pointers, fast/slow pointers, window expand/shrink
3Binary SearchFind Min in Rotated Array, Search in Rotated Array, LISSInvariant-based search, search on answer space
4Linked ListsAdd Two NumbersRunner technique, dummy head, reversal
5TreesAVL + Range QueryDFS/BFS traversal, path problems, BST properties
6Graphs (traversal)Number of Islands, Clone GraphDFS/BFS on grids, graph copying
7Graphs (dependencies)Course Schedule, Pacific AtlanticTopological sort, multi-source BFS
DayFocusProblemsPatterns
8DP: 1D basicsClimbing Stairs, House Robber, House Robber II, Coin ChangeFibonacci-style, take/skip, unbounded knapsack
9DP: sequencesDecode Ways, Word Break, LCS, Max SubarrayPartition DP, two-string DP, Kadane’s
10DP: advancedMax Product Subarray, Unique Paths, Knapsack, Edit DistanceGrid DP, min/max tracking, 0/1 knapsack
11DP: hard variantsKadane, DP on Graph, Optimal BST, AlignmentInterval DP, Knuth optimization, reconstruction
12GreedyBuy/Sell Stock, Jump Game, Huffman Coding, Examples 1-3Greedy choice property, exchange argument, priority queues
13Backtracking & SearchCombination Sum, Map ColoringChoice/explore/unchoose, pruning, constraint propagation
14Math, Bits & RecursionCount Bits, 1 Bits, Reverse Bits, Sum Without +, Tree Path, Max Subarray D&CBit tricks, divide & conquer, Recursion patterns

Week 3: Specialized Tracks & Interview Prep

Section titled “Week 3: Specialized Tracks & Interview Prep”
DayFocusProblemsPatterns
15Probabilistic structuresBloom FilterBloom filters, HyperLogLog, skip lists
16Company-targeted reviewPick 2-3 companies from interviews/ and study their guidesShared concepts across companies
17Mock interview dayRevisit 1 problem from each of topics 01-11 under timed conditions (45 min each)Review all pattern READMEs as quick reference
  1. Read patterns first (15 min) — read the topic README before touching code
  2. Solve problems (60-90 min) — attempt each problem for 20 min before looking at the solution
  3. Review and annotate (30 min) — understand the solution, note edge cases, trace through examples
  4. Spaced review (15 min) — re-solve one problem from a previous day without looking at code

~10-12 hours per week. Requires math foundations (linear algebra, calculus, probability).

WeekFocusReference
1Statistical learning framework, bias-varianceISLR Ch 2
2Linear regression, model selectionISLR Ch 3
3Classification (logistic regression, LDA, Naive Bayes)ISLR Ch 4
4Resampling, regularization (ridge, lasso)ISLR Ch 5-6
5Trees, random forests, boosting, SVMs, unsupervisedISLR Ch 8-9, 12
WeekFocusReference
6ML basics, feedforward networks, backpropagationGoodfellow Ch 5-6
7Regularization, optimization (SGD, Adam)Goodfellow Ch 7-8
8CNNs — architectures (LeNet to ResNet), applicationsGoodfellow Ch 9
9RNNs, LSTMs, sequence-to-sequenceGoodfellow Ch 10
10Practical methodology, debugging, hyperparameter tuningGoodfellow Ch 11
11Autoencoders, generative models (VAE, GAN)Goodfellow Ch 14, 20
12Transformers, attention, BERT, GPT, scaling lawsBeyond book — papers
13Representation learning, self-supervised, transfer learningGoodfellow Ch 15 + papers
WeekFocusReference
14Bandits, MDPs, Bellman equationsSutton & Barto Ch 2-3
15DP, Monte Carlo, TD learning, Q-learningSutton & Barto Ch 4-6
16Function approximation, DQNSutton & Barto Ch 9-10
17Policy gradient, actor-critic, A2C/A3CSutton & Barto Ch 13
18PPO, DDPG, TD3, SACSpinning Up
19RLHF, alignment, model-based RLResearch papers
WeekFocusReference
20Transformer serving view, prefill vs decode, TTFT/TPOTML Systems + vLLM docs
21KV cache, PagedAttention, prefix cachingvLLM (PagedAttention) paper
22Continuous batching, FlashAttention, kernel optimizationTGI + FlashAttention papers
23Quantization (GPTQ/AWQ/FP8), speculative decodingLilian Weng inference blog
24Tensor/pipeline/expert parallelism, multi-GPU servingTensorRT-LLM docs
25GPU clusters, K8s deployment, autoscaling, SLOsK8s + KServe/Ray Serve docs

Weeks 26-30: Foundation Models & Architectures

Section titled “Weeks 26-30: Foundation Models & Architectures”
WeekFocusReference
26Foundation models, the AI-engineering shift, adaptation ladderHuyen AI Engineering Ch 1-2
27Transformer architecture variants (RMSNorm, RoPE, SwiGLU, MoE)Llama/Gemma tech reports
28Attention mechanisms (MHA/MQA/GQA/MLA, sliding-window, sparse)DeepSeek + Mistral papers
29Modalities: text vs audio vs vision, multimodal fusion (CLIP, VLMs)HF Transformers + Whisper/CLIP
30Diffusion models (DDPM, latent diffusion, DiT, flow matching)HF Diffusers + Lilian Weng

Weeks 31-33: Neural Architectures & DL History

Section titled “Weeks 31-33: Neural Architectures & DL History”

Foundations deep dive — best done alongside Deep Learning (weeks 6-9). Derive each architecture from its math and code it from scratch.

WeekFocusReference
31History (perceptron → XOR winter → backprop); foundational math; code perceptron + MLP backpropCS231n + Rumelhart 1986
32CNNs: convolution math, pooling, LeNet→AlexNet→VGG→ResNet; AlexNet deep diveCS231n + AlexNet/ResNet papers
33RNNs, BPTT, vanishing gradients; LSTM/GRU gate equations; CNNs vs RNNs vs TransformersLSTM paper + Goodfellow Ch 10

How a model is made before it is served. Pairs with RL (weeks 14-19) for the RLHF and GRPO sections.

WeekFocusReference
34Pretraining: cross entropy, 6ND FLOPs, Chinchilla vs inference-optimal overtraining, MFUChinchilla + code/memory_calc.py
35Training memory and parallelism: Adam state, ZeRO/FSDP2, TP/PP/CP/EP, JAX shardingZeRO + Megatron papers, PyTorch/JAX docs
36Post-training: SFT, PPO-RLHF, DPO derivation, GRPO and verifiable rewardsInstructGPT, DPO, DeepSeekMath papers
37PEFT and scoping: LoRA/QLoRA/DoRA math, run a LoRA fine-tune, eval base vs tunedLoRA, QLoRA papers + code/lora_from_scratch.py

~8-10 hours per week.

WeekFocusReference
1Scalability, load balancing, caching, CDNByteByteGo
2Database design, SQL vs NoSQL, CAP theorem, replicationByteByteGo + DDIA
3Distributed systems primitives (consensus, distributed transactions)ByteByteGo
4Messaging, event systems, API designByteByteGo
5Classic problems (URL shortener, chat, news feed, etc.)ByteByteGo
WeekFocusReference
6Layered, event-driven, microkernel patternsSoftware Architecture Patterns
7Microservices, space-based architectureSoftware Architecture Patterns
8ADRs, quality attributes, trade-off analysisSoftware Architecture Patterns
WeekFocusReference
9Containers, Docker, multi-stage buildsCloud Native DevOps with K8s
10Kubernetes fundamentals (pods, deployments, services)Cloud Native DevOps with K8s
11Advanced K8s (StatefulSets, CRDs, operators, RBAC)Cloud Native DevOps with K8s
12Helm, GitOps, CI/CD, service meshCloud Native DevOps with K8s
WeekFocusReference
13Observability vs monitoring, instrumentation, OpenTelemetryObservability Engineering
14SLOs, error budgets, debugging with observabilityObservability Engineering + SRE Book
15Production excellence, incident response, chaos engineeringObservability Engineering

~8-10 hours per week. Requires SQL and system design basics.

WeekFocusReference
1Data engineering lifecycle, OLTP vs OLAP, ETL vs ELTData Engineering Cookbook + DDIA Ch 1-2
2Storage formats (Parquet/Avro/Arrow), schemas, compressionFundamentals of Data Engineering
3Build an end-to-end mini pipeline (ingest → store → transform → query)Hands-on
WeekFocusReference
4Storage engines (LSM vs B-tree), columnar executionDDIA Ch 3
5Warehouse / lake / lakehouse, separation of storage & computeIceberg/Delta docs
6Open table formats, partitioning, clustering, file sizingIceberg docs
7Dimensional modeling (star schema, grain, SCDs)The Data Warehouse Toolkit
WeekFocusReference
8Distributed processing, partitioning, shuffle, SparkSpark docs
9Kafka — log, partitions, consumer groups, replayKafka docs
10Event time, watermarks, windowing, the Dataflow modelStreaming Systems
11Delivery semantics, exactly-once, checkpointing, FlinkFlink docs + DDIA Ch 11
WeekFocusReference
12Orchestration (Airflow/Dagster), DAGs, idempotency, backfillsAirflow docs
13dbt — models, ref(), tests, materializations, lineagedbt docs
14Data quality, contracts, governance, data observabilityGreat Expectations docs

~8-10 hours per week. Patterns-first: the goal is to internalize the fundamental patterns (like math, everything downstream is derivative) and treat each trendy tool as an instance. Requires Deep Learning and LLM Systems, plus System Design and Data Engineering foundations (OLTP vs OLAP, replication). Tip: Coding & Design Patterns can be read first as vocabulary.

WeekFocusReference
1PyTorch deep-dive: eager, autograd, nn.Module, torch.compilePyTorch docs
2JAX: jit/grad/vmap/pmap, Flax/Optax, PyTorch vs JAXJAX docs
3Distributed training I: data parallel (DDP), all-reduce, mixed precisionPyTorch DDP + Ultra-Scale Playbook
4Distributed training II: FSDP/ZeRO, tensor/pipeline/expert parallelism, Megatron/DeepSpeedFSDP tutorial + DeepSpeed docs
5Embedding models: contrastive fine-tuning, hosting (TEI), vector storesSentence-Transformers + pgvector
6Small language models: distillation, LoRA/QLoRA (PEFT), quantize, serve with vLLMHF PEFT + vLLM docs
WeekFocusReference
7RPC model, partial failure, serialization formats (JSON/Protobuf/Avro/Thrift)DDIA Ch 4 + protobuf.dev
8gRPC: HTTP/2, the four call types, streaming, interceptors, gatewaysgRPC docs
9Schema evolution, choosing the protocol per boundary, Arrow FlightAvro spec + gRPC docs
WeekFocusReference
10SSE protocol, EventSource, SSE vs WebSockets vs long-polling vs gRPC streamingMDN SSE + HTML spec
11End-to-end token streaming, cancellation/backpressure, proxies and scalingOpenAI streaming + vLLM docs
WeekFocusReference
12OLTP vs OLAP, sharding (hash/range/consistent hashing), hot shards, rebalancingDDIA Ch 5-6 + Citus/Vitess
13In-memory stores & caching patterns (cache-aside/through/back), invalidationRedis docs + Scaling Memcache
14Caching failure modes (stampede, herd, penetration); Cassandra (LSM, ring, quorum)Redis + Cassandra docs
15Knowledge graphs (Neo4j, Apache AGE on Postgres), bitemporal/temporal dataAGE + Neo4j docs

Weeks 16-18: Durable Orchestration, Workers & Patterns

Section titled “Weeks 16-18: Durable Orchestration, Workers & Patterns”
WeekFocusReference
16Durable execution & checkpointing, event-sourced replay, the worker patternTemporal + DBOS docs
17Reliability patterns (idempotency, saga/compensation, outbox); distributed observability + profilingTemporal docs + OpenTelemetry + Gregg
18Coding & design patterns: decorator, facade, closures, currying, compositionRefactoring Guru + Mostly Adequate Guide

Weeks 19-25: Retrieval, Authorization & Evaluation

Section titled “Weeks 19-25: Retrieval, Authorization & Evaluation”
WeekFocusReference
19Encoders (bi- vs cross-encoder), embedding retrieval, ANN/HNSWRAG paper + HNSW + sbert
20Chunking patterns (fixed/recursive/semantic, parent-child), metadatasaige RAG + pgvector
21Lexical retrieval (BM25, bag-of-words, TF-IDF); hybrid fusion (RRF) + rerankingBM25 review + RRF paper
22Multimodal handling (Document→Section→Variant), GraphRAG, production retrievalGraphRAG + saige
23Authorization: RBAC/ABAC/ReBAC/NGAC, the pushdown complexity ladder, ZanzibarNIST + Zanzibar + OpenFGA
24Secure RAG (permission-filtered retrieval, multi-tenancy); PDP/PEP, policy-as-codeNIST ABAC + OPA
25LLM evaluation: cross-entropy, perplexity, bits-per-byte; benchmarks & contaminationMacKay + HELM + The Pile

Weeks 26-28: Edge, Realtime & On-Device Inference

Section titled “Weeks 26-28: Edge, Realtime & On-Device Inference”
WeekFocusReference
26The deployment spectrum; llama.cpp + GGUF end-to-end (convert → quantize → llama-server); Ollama; MLX/ExecuTorch on-devicellama.cpp + GGUF + Ollama
27Realtime/streaming encoders: offline vs causal/chunked; speech (Whisper offline, Conformer/RNN-T streaming); real-time factorWhisper + Conformer papers
28Efficiency architectures (the Mistral clinic): sliding-window attention, GQA, Mixtral MoE, Codestral Mamba/SSM, VoxtralMistral 7B + Mixtral + Mamba papers
WeekFocusReference
29The three routing decisions; router families (rules, predictive, cascade, verifier-gated, session-level); oracle vs deployed router; model recallRouteLLM + FrugalGPT + LLMRouterBench
30Cache-aware switching, escalation policy (category, verifier, confidence), token-basis economics, the SFT/DPO/RFT ladder, evaluating a router, System One decision models, specialised models in the poolFireRouter docs + RFT docs + TypeSafe docs

LLM-as-judge & Elo arenas, RAG metrics (faithfulness, context precision/recall, NDCG/MRR), agent metrics (TTFT/TTLT, tool success), and building a composable Scorer eval harness gated in CI. References: Ragas, Chatbot Arena, lm-evaluation-harness, saige eval/.

  1. Name the pattern, then the tool — for every tool, ask “what pattern is this an instance of?” That’s the transferable knowledge.
  2. Build, don’t just read — train it, shard it, cache it, stream it, orchestrate it, retrieve it, authorize it, score it.
  3. Measure — GPU memory, tokens/sec, bytes-on-wire, TTFT, cache hit rate, retrieval recall, bits-per-byte; numbers beat intuition.
  4. Break it on purpose — kill a worker mid-workflow, expire a hot cache key, reuse a Protobuf tag, leak a doc past an ACL; learn the failure modes.

~6-8 hours per week. Pairs with Software Craftsmanship (the principles behind treating artifacts like source) and Infrastructure (the systems you draw). Uses the Streamflow event-driven platform as the running example.

DayFocusReference
1C4’s four levels (Context/Container/Component/Code); the map analogyc4model.com
2Install d2; render the Streamflow C4 set; edit and git diff a changeD2 docs
3Mermaid inline (sequence, state, ER); pick the diagram by the questionMermaid docs
4-5Redraw a system you know at L1+L2; set up render-in-CI for living diagramsHands-on
DayFocusReference
1Diátaxis: tutorial / how-to / reference / explanation — don’t mix themdiataxis.fr
2Google Technical Writing One; active voice, BLUF, tables over proseGoogle Tech Writing
3Docs-as-code: version, review, test samples in CI, deprecateSWE at Google Ch 10
4-5Write one ADR + one runbook for a real systemADR / Write the Docs

~8-10 hours per week. The math first (lambda-calculus ladder → Hindley-Milner inference), then the same concept across five type systems. Run the code/ files as you read.

DayFocusReference
1What a type system is: judgments Γ ⊢ e : τ, soundness = progress + preservation, Curry-HowardTAPL Ch 1-9 / Software Foundations
2The polymorphism taxonomy (Strachey, Cardelli-Wegner): parametric, ad-hoc, subtype, boundedCardelli-Wegner paper
3The lambda-calculus ladder: STLC → System F (parametric polymorphism, parametricity)TAPL Ch 22-23
4-5Hindley-Milner + Algorithm W: unification, the occurs check, let-generalization, principal types. Read & run code/hindley_milner.pyDamas-Milner paper
DayFocusReference
1Structural typing + variance + type-level programming — code/polymorphism.tsTS Handbook
2Go generics: type sets/unions, GC-shape stenciling — code/generics.goGo generics tutorial
3Rust traits: static (monomorphized) vs dynamic (dyn) dispatch, associated types — code/traits.rsRust Book Ch 10, 17
4Haskell typeclasses (dictionary passing) + higher-kinded Functor — code/typeclasses.hsWadler-Blott paper
5OCaml HM inference + modules/functors — code/inference.ml; compare its principal types to Week 1Real World OCaml
DayFocusReference
1-2Subtyping & variance in depth: declaration- vs use-site, the array-covariance soundness holeTAPL Ch 15-16
3How generics compile: monomorphization vs erasure vs dictionaries — the runtime-cost trade-off—
4-5The frontier: HKTs, associated types, GADTs, dependent types (the lambda cube top)PLFA

~6-8 hours per week. Individual craft → organizational craft → the testing mentality that ties them together → the failure modes that show what happens without them.

Weeks 1-2: Pragmatic Programmer & SWE at Google

Section titled “Weeks 1-2: Pragmatic Programmer & SWE at Google”
WeekFocusReference
1Pragmatic philosophy; DRY, orthogonality, reversibility, tracer bullets, design by contractThe Pragmatic Programmer
2HRT & culture, code review, Hyrum’s Law, deprecation budgets, build/CI as leverageSWE at Google

Week 3: The Testing Mentality & Lessons from Practice

Section titled “Week 3: The Testing Mentality & Lessons from Practice”
DayFocusReference
1Test pyramid; test behavior not implementation; fakes > mocksSWE at Google Ch 11-14
2Property-based testing + fuzzing; add one property testHypothesis / proptest
3Testing infra, data, docs, diagrams (the mentality everywhere)Conftest / dbt tests
4-5Canary, synthetic monitoring, chaos experiment design; SLOsPrinciples of Chaos
6Naming, data modeling, reliability in small code (Themes 1-2)Lessons from Practice
7Judgment signals, scope discipline, evidence and closure (Themes 3-5); audit one of your own repos against the catalogLessons from Practice

~8-10 hours per week. How distributed systems actually run. Topics 1-3 stay anchored to the Streamflow event-driven platform.

Weeks 1-2: Containers, Kubernetes & Workloads

Section titled “Weeks 1-2: Containers, Kubernetes & Workloads”
WeekFocusReference
1Containers (namespaces/cgroups/layers), multi-stage + distroless, 12-factorDocker / 12factor.net
2K8s as a reconcile loop; objects; stateful vs stateless; caveats (limits/OOM, probes, SIGTERM, PDB); scaling (HPA/VPA/Cluster Autoscaler/KEDA)Kubernetes docs

Weeks 3-4: Messaging & Distributed Workers

Section titled “Weeks 3-4: Messaging & Distributed Workers”
WeekFocusReference
3Queue vs log, push/pull/long-poll, pub-sub vs competing consumers, delivery semantics, Little’s Law; ZooKeeper vs KRaft vs RedpandaKafka docs + DDIA Ch 11
4Consumer groups, partition cap, rebalancing, idempotency/outbox, DLQ, lag → KEDA scalingKafka consumer docs
WeekFocusReference
5Inverted index, TF-IDF → BM25, Lucene segments/merges, Elasticsearch vs OpenSearch vs SolrLucene / ES Guide
6The Apache data ecosystem map: messaging, storage/tables, OLAP, search, orchestrationApache docs + DDIA

~4-6 hours per study. Each is a worked build: read the README, run the implementation, then break it deliberately and watch the tests catch you. Best taken after the track that carries the theory, as the applied pass.

DayStudyDo this
1-2Order Book MatchingWrite the flat-list version yourself first, then do the BUD analysis before reading the heap layer
3-4Grounded SQL AgentTry to get a write past the safety gate; design an eval harness for a domain of your own
5-6Exactly-Once Event APIReplace the constraint dedup with check-then-act and watch the concurrency test fail
7-8K-Means OptimizationPredict the timing sweep before running it; explain why dropping sqrt is safe

Pairs with Lessons from Practice, which catalogues the failure modes these builds were avoiding.


The full journey from foundations through PhD-level depth and Staff+ engineering expertise.

WeeksFocusDetails
1-4Calculus 1 + Linear AlgebraParallel math study
5-8Calculus 2 + Discrete Math 1Parallel math study
9-12Calculus 3 + Discrete Math 2Parallel math study
13Math review & catch-upRevisit hardest topics
14-16Probability & StatisticsCompletes math track
17-19Algorithm Mastery17-day plan (one week per study-plan week)
20-21Competitive ProgrammingAdvanced techniques, contest practice
22-26Statistical Learning + Deep LearningISLR then Goodfellow
27-30Deep Learning (continued) + Info TheoryResearch-level depth
31-34Reinforcement LearningSutton & Barto + Spinning Up
35-39Systems & ArchitectureSystem design through observability
40-46AI Platform EngineeringPatterns-first: training, RPC, streaming, data & caching, orchestration, design patterns, RAG, authorization, evaluation
47-48Case StudiesFour worked builds, applied pass over the theory
49-51Interview preparationCompany-targeted study, mock interviews

For those pursuing research depth in ML/AI. Assumes math foundations and statistical learning are complete.

WeeksFocusReading
1-8Deep Learning (full book)Goodfellow Parts I-III, all 20 chapters
9-10Information TheoryMacKay — entropy, KL divergence, channel coding
11-14Reinforcement LearningSutton & Barto full + Spinning Up key papers
15-17Transformer architecturesAttention Is All You Need, BERT, GPT, scaling laws papers
18-19Alignment & SafetyConstitutional AI, RLHF, DPO papers
20+Research explorationPick a subfield, read 10+ papers, implement one
  1. Pick your starting point — if you have a math background, skip to algorithms or ML
  2. Parallelize where possible — math topics are designed for parallel study
  3. Don’t skip the connections — the cross-references between tracks are where real understanding lives
  4. Adjust the pace — these timelines assume ~10 hrs/week; scale up or down as needed
  5. Use interview guides last — they’re most effective after building deep understanding