Transformer
트랜스포머
An architecture built on attention, which lets each word decide which other words to look at. It underlies most language and image models today.
Older models read a sentence in order, which made distant relationships hard to hold. A transformer lets every word look at every other word at once and weigh what matters (Vaswani et al., 2017).
Because nothing waits its turn, the computation parallelizes, which suits training on huge data. Most language and image models now sit on this architecture.
Attention cost grows quickly with input length, which is why long documents get expensive.
- AttentionThe mechanism that decides what to look at, and how much.
- Length costLonger inputs grow the computation sharply.


