RL 5: Learning Automata and stochastic environments (1961–1974)
Quick recap, in case you're jumping in from Blog 4. Arthur Samuel's checkers program learned by...
Tag archive
Quick recap, in case you're jumping in from Blog 4. Arthur Samuel's checkers program learned by...
TL;DR Bellman's math told you how to act perfectly — on paper. 1959 computers had neither...
从 DQN、PPO、SAC 到 RLHF,系统梳理强化学习的核心算法原理及其在 LLM 对齐中的关键应用。

With the development of AI models, it has started to occur to me that just using larger amounts of...
Quick recap, if you're just arriving: Blog 1 showed that every core idea in reinforcement learning...

Before GPUs. Before neural networks. Before transistors were even a rumor. There was a hungry cat in...

I created this series to make that steep learning curve far less daunting for you. My goal is to...
The Missing Piece in Jason Wei's Framework: When to Go On-Policy Jason Wei — the...
The Hamilton-Jacobi-Bellman (HJB) equation stands as a cornerstone in the theory of optimal control,...
Executive Summary In reinforcement learning, the quality of your training is bounded by...

Top 15 Reinforcement Learning Questions That Will Appear in Exams If you're preparing for a...

By a Senior Robotics ML Engineer with 12+ years deploying RL in the field After over a decade...