4 pages tagged
#alignment
constitutional-ai
concept
Anthropic alignment: written constitution + AI critique + RLAIF instead of per-example human labels.
deliberative-alignment
concept
OpenAI alignment: model reasons about safety policy at inference time; reasoning-model-enabled.
post-training
concept
SFT/RLHF/DPO/RFT routes; DeepSeek-R1 four-stage recipe; SFT teaches style as much as knowledge.
reward-hacking
concept
Reward overfitting → hacking → tampering → alignment faking; acute in self-improvement loops.
esc