GEOZ

Entity

RLHF(Reinforcement Learning from Human Feedback)

基于人类反馈的强化学习,用于对齐大模型行为与人类价值观,提升模型安全性和可控性。

相关文章