Mean Opinion Score (MOS)

Mean Opinion Score (MOS) is a summary measure of how listeners perceive voice quality, produced either from human ratings or from an algorithm that estimates how people would rate the audio.

MOS compresses several aspects of a listening experience into one comparable result. Human evaluation asks a group of listeners to judge recorded samples under controlled conditions, traditionally on the five-point scale standardized by the ITU, then averages their opinions. Operational voice systems often use an estimated MOS derived from network and audio measurements because collecting human ratings for every call is impractical.

The score can reflect effects such as packet loss, jitter, delay, codec behavior, background noise, clipping, and distortion. However, MOS is not a diagnosis. Two calls can receive similar ratings for different reasons, and a single average can hide short periods of severe degradation. The calculation method, listening conditions, and point in the call path also influence what the result represents.

For AI phone agents, perceived quality affects whether callers can understand the generated speech and whether the system receives intelligible audio in return. Poor audio may contribute to incorrect transcripts, repeated questions, awkward turn-taking, or transfers that would not otherwise be needed. MOS can help teams compare call paths, identify broad trends, and prioritize investigation, but it should be paired with call recordings, transcripts, transport metrics, and outcome data.

It is important to distinguish conversational quality from audio quality. A clear call can still contain an unhelpful response, a long pause, or an incorrect action. Conversely, a useful conversation can complete despite imperfect audio. MOS addresses the listening experience; it does not directly measure task completion, caller satisfaction, reasoning accuracy, or agent behavior.

When using MOS in monitoring, keep the method consistent. Comparing values from different estimators or different capture points can create a false sense of precision. Segment-level data is also more actionable than a call-wide average when the problem is intermittent. The most useful question is not simply whether a score is high or low, but what changed in the network or audio path and whether callers experienced a meaningful effect.

Related terms