6 pages tagged
#post-training
constitutional-ai
concept
Anthropic alignment: written constitution + AI critique + RLAIF instead of per-example human labels.
deliberative-alignment
concept
OpenAI alignment: model reasons about safety policy at inference time; reasoning-model-enabled.
eval-grader-reward
concept
Training-time eval/grader/reward loop; ORM vs PRM; the grader is the critical failure point.
grpo
concept
Group Relative Policy Optimization; drops PPO's value network; default for verifiable-reward RL.
post-training
concept
SFT/RLHF/DPO/RFT routes; DeepSeek-R1 four-stage recipe; SFT teaches style as much as knowledge.
2026-04-27-llm-training-principles-paths-practices
source
Tw93's third essay: post-training, eval/reward, harness, distillation — the pipeline's back half.
esc