The practice of shaping a model's behaviour to match human intent and values — via RLHF, DPO, constitutional AI, or filtered training data.
safety
The practice of shaping a model's behaviour to match human intent and values — via RLHF, DPO, constitutional AI, or filtered training data.