Umbrella term for research and engineering practices that reduce the risk of AI systems producing harmful, deceptive, or unaligned outputs.
Umbrella term for research and engineering practices that reduce the risk of AI systems producing harmful, deceptive, or unaligned outputs.