AI and tech

Reinforcement learning

강화학습

Also known as: RL

Learning by acting, receiving a reward, and adjusting towards actions that earn more.

It fits problems with sequences of actions and delayed outcomes, like games or robotics: instead of showing the right move, you reward winning.

It also shapes language models. Scoring answers that people preferred, then training towards them, is called reinforcement learning from human feedback.

A badly designed reward teaches shortcuts. Whatever you score becomes the goal.

  • Reward designWhat you score decides the outcome.
  • In learningThe same holds for children: reward attempts and consistency, not only correct answers.

Related terms