AI Full Stack
What sits between the endpoint and the user: token streaming, retrieval, evaluation, and routing across models.
For: AI application engineers, and anyone proving that a model swap did not break a product.
Included in: Superstar FDE. Progress made here counts there too.
The stages, their chapters, and what “done” means for each are in path.tsv.
practice/bin/ol learn ai-full-stack # stages and progresspractice/bin/ol learn ai-full-stack next # read the next unfinished stageStages
Section titled “Stages”Generated from path.tsv. Track progress locally with practice/bin/ol learn.
| Stage | Read | Done when |
|---|---|---|
| 1 | Streaming and SSE Streaming & SSE | You stream tokens from an OpenAI-compatible endpoint to a browser over SSE and handle disconnects and backpressure. |
| 2 | Retrieval and RAG Retrieval & RAG | You build hybrid retrieval (BM25 plus vectors with reranking) over a real corpus and measure recall before touching the generator. |
| 3 | LLM evaluation LLM Evaluation | You run an eval suite against two models on your own endpoint and report the delta with confidence intervals. |
| 4 | Model routing and cascades Model Routing & Cascades | You can design a cascade that sends easy traffic to a small model and show the cost and quality trade-off on a held-out set. |