Some things you measure. You do not ask a model to guess them.
How long the customer waited on hold, how much of the call was silence, how fast the agent spoke, how strained their voice was by the end of a shift — these are properties of the waveform, and we measure them. The work runs on a service we operate ourselves, built on Praat and librosa, on our own infrastructure. No second AI vendor, and nobody else holding your audio.
What you get
Measured, not estimated
Talk ratio, speaking pace, filler words, dead air and time on hold come out of the signal and the transcript, deterministically. A supervisor can re-count them by hand and get the same answer — which matters when the number is about to be said to a person.
What the voice itself is doing
Pitch, jitter, shimmer, harmonics-to-noise ratio, spectral flatness and energy. The measures a speech scientist would use, computed with Praat and librosa rather than inferred from a text summary of the call.
On infrastructure we run
The acoustic service is ours, in our own network. Your recordings are not sent to a second AI provider for this, because nothing about measuring a waveform needs one.
It knows what is not speech
Ringback at the start of a recording and hold music in the middle are detected and excluded, so 'dead air' means the customer waiting in silence and not the two seconds of ringing before anyone picked up.
How it works
The original file, once
The measurement runs on the audio you uploaded, before it is compressed for storage — a codec's job is to throw away exactly the noise-like content these measures are made of.
The waveform is measured
Our service computes the acoustic figures and the timeline of holds and silences, and hands back numbers.
It sits beside the score
On the call, on the dashboard, and in the coaching plan — with the company's own averages next to it, so one call reads against how this team normally sounds.
Questions people ask about this
Because a language model reads a transcript, and a transcript has already thrown the audio away. Ask one how long a hold lasted and you get a plausible number, not a measurement. These figures end up in coaching conversations and occasionally in disputes, so they have to be reproducible — and the only way to do that is to measure the waveform.
The acoustic analysis happens on our own infrastructure and the recording is not sent to a third party for it. Transcription is the one step that uses Google's enterprise AI platform, under terms that exclude your data from model training. We would rather draw that line clearly than let you assume either more or less than is true.
The call still opens and the metrics taken from the transcript are still there; the acoustic panel is simply absent until it comes back. A quality tool that refuses to show a call because a side service is restarting is worse than one that shows a little less.
Related capabilities
See it on your own calls
Send us a handful of recordings and your current scorecard. We will score them, show you the result next to what your team scored, and tell you honestly what you would and would not gain.
No automated demo. A real conversation, usually within one business day.