AI and tech

Speech recognition

음성 인식

Also known as: STT · speech to text

Turning spoken language into text.

It powers dictation, captions and voice commands. Accuracy is high in quiet conditions with clear speech and drops with noise or overlapping speakers.

Children's voices are harder than adults': the pitch differs and pronunciation is still forming, so features involving child speech need more validation.

Scoring features such as pronunciation assessment deserve extra care, so that a number does not discourage the child.

  • DifficultyNoise, overlapping speech and children's pronunciation, in that order.
  • DesignIf you show a score, show the next step too.

Related terms