EP 58: Every Millisecond Matters: Diffusion LLMs and the Future of Voice AI | Aditya Grover, Inception
Inside AsembleAI: DeepTech, AI & Science
This episode of 'Show: Inside AsembleAI' features Aditya Grover, CTO and co-founder of Inception, discussing the future of voice AI and the
Key takeaways
- Traditional autoregressive LLMs generate tokens sequentially, creating compounding delays that hinder real-time performance.
- Diffusion-based models enable parallel generation of text, allowing for faster response times while maintaining high-quality output.
Main topics
- Diffusion-based LLMs
- Latency optimization in AI
Notable quotes
Every millisecond matters — latency compounds across reasoning chains, making real-time applications sluggish.
Conclusion
Aditya Grover envisions a future where voice agents are seamlessly intelligent,
Transcript preview
Speaker 1 (0:00) Welcome to Assemble.ai. Subscribe and follow us on YouTube, Spotify, Apple Podcasts, iHeartRadio, and Substack. Speaker 2 (0:08) Hello, hello, hello. Hello, Assemble.ai listeners out there. So this is the last podcast of the day one from the AI4 Podcast Pavilion. I'm your co-host, Sam, and here I joined with Aditya Grover from Inception. Adi, welcome to the show. Thanks Speaker 1 (0:27) for having me. Speaker 2 (0:28) So Aditya, before we start, could you please tell a little bit about yourself for our listeners? Instead of me reading your bio, I don't want to do that. So, good to see you. Speaker 1 (0:38) Yeah, so my name is Aditya Grover. I'm one of the co-founders and CTO at Inception. And I have a background in doing research in genitive AI and reinforcement learning. I've been doing that for over 10 years. So I grew up in India, did my undergraduate at IIT Delhi, and then I moved here. to Stanford University where I did my PhD. And during my PhD, I got into this field of genitive modeling. And this field is today what has resulted in all the advances that we see in genitive AI broadly. And since my PhD, I also got a chance to then scale up some of these ideas in industry. So I spent some time in the early days at OpenAI, also at Google DeepMind and Meta before I decided to start my own research lab at UCLA. as a faculty where I worked at the intersection of genitourine models and reinforcement learning. And sometime two years ago, my co-founders and I, we were convinced that the next generation of LLMs would look very much more parallelizable and efficient than they have been before. And to make that happen, we decided to start our own company called Inception. Speaker 2 (1:47) Great. You had a fascinating journey. Like you're someone who literally lived up to the academic tenure and then moved to industry, but you're still juggling between those two things. So there are a lot of listeners out there who might be thinking about how can I balance between my academia job and being an investor or like a founder of an industry company, right? So please do reach out to Aditya. He will give you some tips. Moving on to my first real question about inception. So I went to your website. And I kind of like did some research. Your voice starts with a tagline, every millisecond matters. So what does it mean? Could you please walk us through that? Yes. So Speaker 1 (2:28) I think today's AI is in a place where intelligence is everywhere, but it's still not being delivered to us in a way that actually makes our applications run at the efficiency we'd want them to be. So think about a voice agent. So in the case of a voice agent, Every quarter of a second lost from an automated voice chat bot means you're going to be asking for customer service. That's true. And that's where AI has failed to deliver on its promise. And this is not just a phenomenon that's more unique to voice now. If you start thinking about agents, they're not making one tool call, they're making dozens of tool calls before you get an answer. So even if every individual tool call might be well optimized, the fact that they're going to compound together in a very long reasoning chain means that your end application is going to be extremely slow. And that's where we feel that milliseconds of latency with LLM turns is a compounding problem. And that's what we're trying to solve at Inception. That's Speaker 2 (3:31) great, Chuna, because... There are a lot of companies out there. Yes, we are trying to make the model bigger and better. But what about the efficiency? Like you mentioned, like every millisecond matters. Like there are the compounding factor to that if LM is going to provide you answer in a slowness rate, right? So not a lot of companies are working on an LM model to address that.