About this course
The general-education course for the age of commanding agents: why LLMs reason, where
agents came from, how multi-agent frameworks and compound AI systems work, how agents land
in software, enterprise workflows and robotics, and how to measure capability and safety.
It follows the 12-lecture skeleton of Berkeley's renowned LLM Agents course (Fall 2024),
with every session rewritten as original hands-on teaching for the online sandbox.
What you'll learn
- Explain what each LLM reasoning technique (CoT, self-consistency, least-to-most) solves and when it fails
- Trace the lineage of ReAct and the agent paradigm; run the reason-act-observe loop yourself
- Compare multi-agent frameworks and compound AI systems (conversational collaboration, declarative optimization)
- Explain how agents land in software development, enterprise workflows and robotics — and the bottlenecks
- Assess agents through capability evals and responsible scaling; explain injection risks and defenses
Syllabus
1Foundations: Reasoning & a Brief History of Agents2 sessions
2Frameworks: Multi-Agent & Compound AI Systems3 sessions
3Applications: Code, Workflows & Embodiment4 sessions
4Ecosystem & Safety: Openness, Evals & Trust3 sessions
Same series · 世界名校知名实验室系列
Flagship University Lab Series
Modeled on Stanford / MIT / Berkeley syllabi — learn from scratch with a mentor Agent guiding you in real time.
◎ Modeled on Stanford CS336
LLM from ScratchHand-build every piece of an LLM, along the 17-lecture skeleton of Stanford CS336View course →◎ Modeled on Stanford CS329A (Self-Improving AI Agents)
Self-Improving AI AgentsHand-build AI agents that make themselves better, along the skeleton of Stanford CS329AView course →◎ Modeled on MIT’s How to AI (Almost) Anything (multimodal)
Multimodal AI: Teaching Models to See, Hear, and ConnectTurn images, sound and text into tensors by hand — then build alignment, fusion, cross-modal retrieval and multimodal models piece by pieceView course →◎ Modeled on Princeton COS 511 (Theoretical ML)
Learning Theory: PAC, Boosting & Online LearningHand-derive PAC bounds, hand-build AdaBoost, hand-run regret curves — turn learning theory from formulas into experiments you runView course →◎ Modeled on Stanford CS234 (Reinforcement Learning)
Deep Reinforcement Learning: MDPs to Policy GradientsProve convergence by hand, backprop a DQN from scratch, give policy gradient a critic — pure numpy, zero gym, zero GPUView course →◎ Modeled on CMU’s ML in Production (MLIP)
Machine Learning in ProductionTurn a model that merely runs into a system that survives productionView course →