
- SE0 — Series Overview — Mastering Language Models: From Architecture to Optimization
- T1E0 — T1E0 · Foundations of Sequence Modeling: The Transformer Revolution
- T1E1 — T1E1 · Attention Is All You Need
- T1E2 — T1E2 · Kimi Linear: An Expressive, Efficient Attention Architecture
- T1E3 — T1E3 · Attention Residuals
- T2E0 — T2E0 · Scaling and Training Large Models Efficiently
- T2E1 — T2E1 · Scaling Laws for Neural Language Models
- T2E2 — T2E2 · Training Compute-Optimal Large Language Models
- T2E3 — T2E3 · Scaling Data-Constrained Language Models
- T3E0 — T3E0 · Advanced Distributed Training: Overcoming Bottlenecks
- T3E1 — T3E1 · GPipe: Efficient Training of Giant Neural Networks Using Pipeline Parallelism
- T3E2 — T3E2 · Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- T3E3 — T3E3 · ZeRO: Memory Optimization Towards Training Trillion Parameter Models
- T3E4 — T3E4 · Fully Sharded Data Parallel: Faster AI Training with Fewer GPUs
- T3E5 — T3E5 · Research on Distributed Training Architecture for Large Scale Models for Natural Language Processing
- T3E6 — T3E6 · FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
- T3E7 — T3E7 · FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- T3E8 — T3E8 · Embarrassingly Simple Self-Distillation Improves Code Generation (SSD)
- T3E9 — T3E9 · Test-Time Scaling Makes Overtraining Compute-Optimal
- T4E0 — T4E0 · Fine-Tuning and Specialization: LoRA and Beyond
- T4E1 — T4E1 · LoRA: Low-Rank Adaptation of Large Language Models
- T4E2 — T4E2 · QLoRA: Efficient Finetuning of Quantized LLMs
- T4E3 — T4E3 · LowRA: Accurate and Efficient LoRA Fine-Tuning of LLMs under 2 Bits
- T4E4 — T4E4 · Continual Learning of Large Language Models: A Comprehensive Survey
- T5E0 — T5E0 · Reinforcement Learning from Human Feedback (RLHF)
- T5E1 — T5E1 · Proximal Policy Optimization Algorithms
- T5E2 — T5E2 · Learning to Summarize with Human Feedback
- T5E3 — T5E3 · Training Language Models to Follow Instructions with Human Feedback
- T5E4 — T5E4 · Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- T5E5 — T5E5 · Constitutional AI: Harmlessness from AI Feedback
- T5E6 — T5E6 · Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- T5E7 — T5E7 · RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs
- T5E8 — T5E8 · Robust Reinforcement Learning from Human Feedback for Large Language Models Fine-Tuning
- T5E9 — T5E9 · Reinforcement Learning from Human Feedback: Progress and Challenges
- T6E0 — T6E0 · The Latest in Fine-Tuned and Open Models: From LLaMA to DeepSeek
- T6E1 — T6E1 · Llama 2: Open Foundation and Fine-Tuned Chat Models
- T6E2 — T6E2 · The Llama 3 Herd of Models
- T6E3 — T6E3 · DeepSeek-V3 Technical Report
- T6E4 — T6E4 · DeepSeek-V4-Flash
- T7E0 — T7E0 · Mixture-of-Experts Models and Handling Massive Models: Sparsity, Data, and Optimization
- T7E1 — T7E1 · Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- T7E2 — T7E2 · Mixture-of-Experts Operator Transformer for Large-Scale PDE Pre-Training
- T7E3 — T7E3 · The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- T7E4 — T7E4 · AutoML from Basics to State-of-the-Art Techniques
- T7E5 — T7E5 · Recent Advances in Optimization Methods for Machine Learning: A Systematic Review