Inference
추론
Also known as: serving
Running a trained model to get a result, as opposed to training it.
Training happens once at scale; inference happens on every request. That is where the running cost lives.
It also decides perceived speed: how many seconds one photo takes, how many users can be served at once.
You lower it with smaller models, lighter computation, or caching, though caching is off the table for material like personal photos.
- CostIt grows with usage, not with training.
- DesignSeparate screens that need a fast answer from jobs that can wait.


