When a caller talks over an AI voice agent, stopping the audio is only the first of five things that have to stop. The frames already queued, the sentence still being rendered, the open model stream and the conversation record all have to be cut too. Get the last one wrong and the rest of the call is broken.
We know because we got it wrong. The AI Voice Agent in ICTContact has been through six releases since we opened the underlying sidecar as asterisk-ai-voice-agent, and barge-in was the single richest source of bugs in the lot. Most of them were invisible in the logs. This is what we found, and what we would test before trusting any voice agent on live traffic.
Barge-in looks like it works long before it does
From the first release, playback stopped when the caller spoke. By the usual definition that is barge-in, and a demo call sounds right. The problem is everything still running behind the silence.
An engineer who works on the streaming side of a voice API put it to us as a question rather than a bug report: once your text to speech is streaming into a pacer, what happens to the part already rendered? The honest answer at the time was that we had never looked. Once we did, it turned out that four separate pieces of work carried on after the caller had heard silence.

One: playback, and the queue behind it
Stopping the writer is obvious. Dropping what you already handed to the phone system is not, and it sets a hard floor on how fast a barge-in can feel.
If your media path takes audio ahead of real time, every second you have queued is a second the caller keeps hearing after they started speaking. That is not a tuning problem you can fix later. Queue depth is your barge-in latency, so pick the depth deliberately rather than discovering it on a customer call.
Two: the sentence still being rendered
Our stop flag was only checked between frames, and synthesis sat one step earlier than that. So a caller who interrupted while the sentence was still being built had no effect at all. The whole sentence got rendered, then playback stopped at the very first frame. Perfectly quiet, completely wasteful.
With a cloud voice, that waste is on your bill. With a local voice it is worse, because rendering holds a lock that the next caller is waiting on. Passing the stop check down into synthesis itself was the fix, and on a metered voice it means an interrupted sentence stops being charged.
There is a trap in how you stop it, though. Cancelling the render outright releases the lock while the engine underneath is still working, and the next call walks straight into a busy engine. We measured both: cancelling left a second caller waiting 2.14 seconds, while letting the doomed render finish and throwing the result away left them waiting 1.32 seconds. Finishing work you intend to discard is the faster option, which is not where intuition points.
Three: the model stream, and the record of what was said
This is the one that ruins calls. Breaking out of the model stream skipped the step that writes the assistant turn into the conversation, so as far as the model was concerned it had never spoken.
Worse, if it had already asked for a tool before the caller cut in, the tool still ran and its result was filed against a request that was never recorded. The model provider rejects that turn, and every turn after it. One badly timed interruption, and the caller hears the fallback apology line for the remainder of the call. No error in the application log, no alert, nothing to see on a dashboard.
What it takes to be correct: record only the sentences that actually reached the caller, drop tool calls belonging to an interrupted turn, and close the abandoned stream immediately instead of leaving it to be tidied up later.
The dead air problem nobody reports as a bug
Separate from interruptions, there is a gap that callers feel but rarely mention. If you render one sentence at a time, sentence two does not start rendering until sentence one has finished playing, so every sentence boundary carries a pause the length of the next render.
We timed a three sentence answer at the far end of a real call rather than at the server, because departure times flatter you and arrival times do not. Sequential rendering came in at 2.65 seconds with two half second holes in it. Rendering the next sentence underneath the current one came in at 1.88 seconds with a worst gap of 120 milliseconds.

Overlap moves work earlier, it does not create capacity. Where a sentence costs more to render than it does to play, the shortfall is still audible and no queue hides it. At that point you are buying cores, not writing code.
The obvious optimisation that made things slower
Someone raised the local voice lock as a cap on concurrent calls, and we agreed with them. A pool of worker processes, one voice each, would obviously let several callers render at once. We built it and measured it with four concurrent calls, five trials each, because single runs disagreed with each other badly enough to be meaningless.
The pool was worse on every axis. Total time 2.30 seconds against 1.99 for the single lock. The first caller waited 1.89 seconds instead of 0.48. The inference runtime was already spreading one render across every core, so the lock had been serialising work that was already using the whole machine, and the pool only added contention and copying. Restricting each worker to one thread made it worse again at 5.41 seconds.
We reverted it and closed the issue with the numbers attached. The lesson is not that pools are bad. It is that on a voice path, the intuitive fix and the measured fix disagree often enough that you should not ship either one without a stopwatch.
The small things callers actually notice
Two of these did more damage to how the agent sounded than any of the above.
The first was sentence splitting. The splitter that feeds text to the voice while the model is still generating fired on every comma, semicolon and colon. So a name with a title became two utterances, a numbered list item turned into a spoken number followed by a fragment, and every clause arrived as its own choppy piece with the voice losing the shape of the sentence at each cut. It now splits at sentence ends, knows about abbreviations, initials and list numbers, and only falls back to a clause break once a sentence runs past 120 characters.
The second was the clipped first word after an interruption. Voice activity detection only marks a frame as speech once the talk spurt has enough energy behind it, so a quiet onset is already gone by the time the utterance opens. A commenter suggested holding a rolling window of the frames just before onset. We said we had gone another way. He was right and we were not. The agent now keeps 300 milliseconds of pre-onset audio per call and prepends it when an utterance opens, which costs under five kilobytes however long the call runs.
One detail there is easy to miss: that prepended audio has to stay out of the voiced duration count, because that number feeds a check for transcription hallucinations. Padding it made short answers look like noise and got them thrown away.
What to test before you trust a voice agent with real calls
If you are evaluating an AI voice agent for a contact center, the demo will not show you any of this. These four tests will.
- Interrupt mid-sentence, then ask a follow-up question. If the agent answers as though it never said the interrupted sentence, the conversation record is broken.
- Interrupt during a turn where the agent is looking something up. If the rest of the call turns into apologies, an orphaned tool call has poisoned the conversation.
- Listen to a long answer end to end. Count the pauses between sentences. Half a second each means nothing is overlapping.
- Interrupt with a quiet word rather than a loud one. If the first syllable goes missing, there is no lookback buffer.
All four are things a caller will do in the first week. None of them need a lab.
The agent side of this sits inside ICTContact and is documented in the AI Voice Agent guide, where you configure personas, the tools they can call, and where a call hands off to a person. If you are weighing it against a hosted platform, the CCaaS overview covers where each model makes sense.
Frequently asked questions
What is barge-in on an AI voice agent?
It is the caller interrupting the agent mid-sentence and the agent yielding. In practice it means five things stopping together: writing audio, the frames already queued, the sentence being synthesised, the open model stream, and the record of what the agent said. Only the first is audible, and the last is the one that breaks later turns.
Why does my voice agent apologise for the rest of the call after an interruption?
Almost certainly an orphaned tool call. The agent asked for a lookup, the caller interrupted, the request never made it into the conversation history, and the result arrived attached to nothing. The provider rejects that turn and every one after it, so the agent falls back to its error line and stays there.
How much dead air between sentences is normal?
Under about 150 milliseconds is fine and most callers will not register it. Half a second at every sentence boundary means the agent is rendering one sentence at a time. That is a design choice in the agent, not a network problem, and it is worth asking a vendor about directly.
Does a local voice limit how many calls an agent can handle?
Less than you would expect. We measured a worker pool against a single lock with four concurrent calls and the pool was slower in total and made the first caller wait roughly four times longer. Modern inference runtimes already use every core for a single render, so the ceiling is CPU. Add cores or add servers rather than adding processes.
Is the AI Voice Agent in ICTContact the same as the open source project?
They share the engineering, not the packaging. The version in ICTContact is wired into campaigns, personas, routing and the rest of the platform. The open source sidecar is the media and conversation layer on its own, built to run against a plain Asterisk with your own model and voice keys. Fixes flow both ways, which is why the public issue tracker has improved the product version.
Can I test any of this without a lab?
Yes, and you should. Interrupt mid-sentence and ask a follow-up, interrupt while the agent is looking something up, listen to a long answer for the gaps, and interrupt with a quiet word. Four calls tell you more about how an agent will behave in production than any scripted demo will.
ICTContact is built by ICT Innovations. The voice agent work described here is open, and the corrections that shaped it came from engineers who had never seen our code.
