4 pages tagged

#alignment

constitutional-ai concept Anthropic alignment: written constitution + AI critique + RLAIF instead of per-example human labels. deliberative-alignment concept OpenAI alignment: model reasons about safety policy at inference time; reasoning-model-enabled. post-training concept SFT/RLHF/DPO/RFT routes; DeepSeek-R1 four-stage recipe; SFT teaches style as much as knowledge. reward-hacking concept Reward overfitting → hacking → tampering → alignment faking; acute in self-improvement loops.
esc