AI and tech

Inference

추론

Also known as: serving

Running a trained model to get a result, as opposed to training it.

Training happens once at scale; inference happens on every request. That is where the running cost lives.

It also decides perceived speed: how many seconds one photo takes, how many users can be served at once.

You lower it with smaller models, lighter computation, or caching, though caching is off the table for material like personal photos.

  • CostIt grows with usage, not with training.
  • DesignSeparate screens that need a fast answer from jobs that can wait.

Related terms