As voice AI moves from experimental demos to real customer conversations, latency is emerging as a critical factor in determining whether an AI agent sounds genuinely conversational or only like a machine. In this conversation with Tech Achieve Media, Subhash Kalluri, Founder, FreJun, explains why sub-200ms latency matters, where delays accumulate across the voice AI stack, and why model-agnostic infrastructure, natural interruptions, live call controls and auditable event pipelines will become increasingly important for enterprise adoption. He also looks ahead to the next phase of voice AI, where speed alone may no longer be enough and infrastructure will need to support more human-like conversations, seamless human escalation and compliance at scale.
TAM: Voice AI is moving from demos to real customer conversations. Why does sub-200ms latency matter so much for making AI agents feel genuinely conversational?
Subhash Kalluri: Teler runs at independently verified sub-200ms latency, and that number isn’t arbitrary. This is important because the human voice takes about 200 milliseconds to pause. After that time has passed, the brain discovers something is incorrect, even if the person cannot articulate what the problem may be. If latency goes beyond that threshold to over 500 ms to 800 ms, which is normal in a traditional voice AI technology setup, the voice agent stops being perceived as a person and rather as a machine executing an order. For applications in the field of debt collection, lead qualification, or customer service, the distinction can be the reason for the voice agent to remain useful or let the subject continue the conversation with a real agent.
TAM: Many organisations are stitching together carriers, telephony, STT, LLM and TTS providers. Where does latency actually accumulate in this architecture, and why is the infrastructure layer becoming the bottleneck?
Subhash Kalluri: Many teams working on voice agents concentrate on optimising the product, such as utilising more efficient models that produce faster results and require less space. However, the latency budget is consumed long before the model gets into action. An incoming call must be processed in multiple steps: upon receiving a call, the telephony device must pick it up, decode the carrier audio, transmit it to the STT engine, pass it to the LLM, then convert it back through the TTS, and send it back to the telephony device. Each of these processes requires some milliseconds, and usually, the telephony and media-streaming equipment remains the most missed source of overall latency. This is why Teler transmits the audio directly via WebSocket to a customer’s AI application instead of using several different systems for signal transmission. The key to success here is to reduce the number of intermediate links between the carrier and the AI machine as much as possible.
TAM: FreJun is positioning Teler as a model-agnostic voice infrastructure layer. Why is vendor independence becoming important as enterprises experiment with multiple LLM, STT and TTS providers?
Subhash Kalluri: The LLM, STT, and TTS landscape is moving too fast for anyone to lock into one stack today and expect it to be the right choice in six months. Teler is built model-agnostic by design: any custom in-house model can plug into the same voice infrastructure. This is very important for companies, as they can replace components depending on price, accuracy, and speed changes without renewing the whole telephony system. Also, a company can use several models depending on the case, such as high-accuracy BFSI working processes on one model and using a cheaper and faster model for low-stakes enquiries. Independence from specific vendor technologies in the infrastructure allows companies to experiment with the economics of AI.
Also Read: FreJun Launches Teler to Build the Voice Layer Under India’s AI Agents
TAM: Barge-in and natural interruption are often overlooked in voice AI. What needs to happen across the infrastructure stack for an AI agent to respond naturally when a caller interrupts it?
Subhash Kalluri: Creating natural interruption may seem easy, but it can get difficult while building it. Natural interruption requires a corresponding infrastructure that will provide the opposite speech flow because the system has to listen while speaking. Real-time voice activity detection must be in place at the media level so that the system detects when the caller starts speaking immediately. The moment that’s detected, TTS should stop in moments, keeping the context information of the interrupted speech so that the LLM could react to what happened. Echo cancellation should be efficient enough to guarantee that the system won’t confuse its own speech with the speech of the caller. Making any mistake, the system may either interrupt its speaker or keep quiet; any of these outcomes will ruin the illusion.
TAM: As voice agents move into customer service, sales and other mission-critical workflows, how important are live call controls, transfers and an auditable event pipeline for enterprise adoption?
Subhash Kalluri: When voice agents enter the realm of customer services, sales, or collection calls, the companies cannot simply settle for “it always works in demonstration”. The enterprises demand to control every step of the conversation, including cold and warm transfers, as well as what the caller and the agent are or aren’t able to hear during the call transfer. The procedure must also allow sending the call to a human representative, as it exceeds the agent’s ability to deal with it. The call itself is not everything. All the events during the entire call lifecycle must be recorded through signable and replayable webhooks, and a complete call history and recordings must be available through the API. This is not a good-to-have option for any regulated field such as banking or insurance but rather a must-have requirement.
TAM: Looking ahead, will latency become a competitive differentiator for AI agents, and what do you believe the next generation of voice infrastructure needs to solve beyond simply making AI respond faster?
Subhash Kalluri: Latency has become a key differentiator because the majority of the market has not yet figured it out, but in the next 18 to 24 months, sub-200 ms latency will be the minimum requirement rather than something that stands out. The next frontier is what happens after speed is no longer an issue, which includes being able to deal with interruptions and edge cases in a human-like manner, matching the emotional tone during difficult conversations, and having a platform that can scale properly in different regions while ensuring compliance is built in from the beginning and not just added later on. The next generation of voice solutions will also need to deal with orchestration and not just solving conversations, which means being able to integrate AI-enhanced communications with a possibility of a human escalation without the caller noticing it. Speed will help AI get through the door, but what matters in the long run will be the infrastructure being able to manage everything that needs to be done beyond the first few seconds of a call.















