
Reinforcement Learning from Human Feedback (RLHF): Roles, Responsibilities, and How LLMs Learn What Humans Prefer
Large language models don't become aligned simply because they become larger. They become useful when...
Tag archive

Large language models don't become aligned simply because they become larger. They become useful when...

What data annotation jobs and AI response review are: the tasks involved, how they feed model training (RLHF), and how annotator quality is measured.
Reward models power RLHF-style training and are genuinely complex to build well. A plain-language explanation, and when a student project actually nee
从 RLHF 到 Constitutional AI,从辩论方案到递归奖励建模,系统梳理 AI 安全对齐的技术路线与开放挑战。
RLHF Burned $50K Before We Admitted SFT Would've Worked Reinforcement Learning from Human...
MTurk is closing to new customers: 6 human-data alternatives compared for 2026 Summary....
The largest audit of AI language model judges to date — 21 judges, over half a million grading decisions — finds that standard reliability metrics are inflated by roughly a third, that the same judge can score differently on different benchmarks, and
Several research groups landed on the same idea at once - improve a model by learning from its own attempts instead of expensive human labels - and the field is debating whether it really removes the labeling burden or just hides it.
How RLHF-trained language models may develop instrumental goals, and the information-theoretic limits on detecting them.
The DPO Hype Promised to Kill RLHF — It Didn't Everyone said Direct Preference...
The $12,000 Surprise RLHF training for a 7B parameter model ran us $12,400 on AWS for...
🦅 FAQ Q: Does ChatGPT actually know what time it is? A: Not your local time, no. ChatGPT...