Reinforcement learning
강화학습
Also known as: RL
Learning by acting, receiving a reward, and adjusting towards actions that earn more.
It fits problems with sequences of actions and delayed outcomes, like games or robotics: instead of showing the right move, you reward winning.
It also shapes language models. Scoring answers that people preferred, then training towards them, is called reinforcement learning from human feedback.
A badly designed reward teaches shortcuts. Whatever you score becomes the goal.
- Reward designWhat you score decides the outcome.
- In learningThe same holds for children: reward attempts and consistency, not only correct answers.


