
Series Overview — Mastering Language Models: From Architecture to Optimization
The map of the journey: shared expert mental models, the field's real fights, and the forks every LLM builder faces
节目笔记
Maya and Leo open the series with the map: seven stops from the Transformer blueprint to the machinery under massive models, anchored by a three-person startup building an insurance-claims assistant on eight GPUs. They lay out the mental models every LLM expert shares — trust curves, find the bottleneck, separate capability from behavior — then stage the field's cleanest fight on air: bigger models versus more data, from OpenAI's 2020 scaling curves to Chinchilla's flip to the serving-cost era that ran past both camps. Plus trailers for the live attention debate and the alignment fight to come.
来源材料
- Attention Is All You Need
- Kimi Linear: An Expressive, Efficient Attention Architecture
- Scaling Laws for Neural Language Models
- Training Compute-Optimal Large Language Models
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
- LoRA: Low-Rank Adaptation of Large Language Models
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Constitutional AI: Harmlessness from AI Feedback
- Llama 2: Open Foundation and Fine-Tuned Chat Models












































