Neil Movva - Making AI 10x Cheaper
Invest Like the Best with Patrick O'Shaughnessy
In this episode of Invest Like the Best, Patrick O'Shaughnessy interviews Neil Movva, founder of SAIL Research, a company building a 'token
Key takeaways
- The future of AI lies in long-horizon tasks where latency matters less and cost matters more.
- Open-source models are gaining traction due to user demand for sovereignty over intelligence.
Main topics
- AI inference cost optimization
- Long-running AI agents and background processing
Notable quotes
"The best latency is no latency at all. When you wake up in the morning, the work's already been done overnight."
Conclusion
Neil Movva envisions a future where AI agents operate continuously and autonomously
Transcript preview
Speaker 3 (0:00) Ramp is the only platform built to make your finance team leaner, faster, and better, saving businesses 5 % annually on average so you can stay focused on growth. Ramp customers grow revenue 3.2 times faster than the average American business. Visa, Vercel, Cursor, Stripe, Notion, 11Lab, Shopify, and 70,000 other businesses all now run on Ramp. Mine does too, and so should yours. Learn more at ramp.com slash invest. OpenAI, Cursor, Anthropic, Perplexity, and Vercel all have something in common. They all use WorkOS. To achieve enterprise adoption at scale, you have to deliver on core capabilities like SSO, SCIM, RBAC, and audit logs. Instead of spending months building these mission-critical capabilities yourself, you can just use WorkOS APIs to gain all of them on day zero. That's why so many of the top AI teams you hear about already run on WorkOS. WorkOS is the fastest way to become enterprise ready and stay focused on what matters most, your product. Visit WorkOS.com to get started. Felix by Rogo is a personal finance agent that turns a single prompt into finished client ready work using your firm's own templates, context, and standards. Send Felix an email like, take these comments and turn them for me, or update my tracker with the context of these emails. And Felix sends back finished PowerPoint decks, Excel models, and sourced research. Felix works the way your team already does, delivering work quickly and accurately around the clock. Learn more at rogo.ai Speaker 2 (1:22) slash Felix. Speaker 3 (1:28) Hello and welcome everyone. I'm Patrick O'Shaughnessy and this is Invest Like the Best. This show is an open-ended exploration of markets, ideas, stories, and strategies that will help you better invest both your time and your money. If you enjoy these conversations and want to go deeper, check out Colossus, our quarterly publication with in-depth profiles of the people shaping business and investing. You can find Colossus along with all of our podcasts at colossus.com. Speaker 1 (1:51) Patrick O'Shaughnessy is the CEO of Positive Sum. All opinions expressed by Patrick and podcast guests are solely their own opinions and do not reflect the opinion of Positive Sum. This podcast is for informational purposes only and should not be relied upon as a basis for investment decisions. Clients of Positive Sum may maintain positions in the securities discussed in this podcast. To learn more, visit psum.vc. Speaker 3 (2:19) My guest today is Neil Nova, the founder of SAIL Research. SAIL is building what Neil calls a token factory, an inference company designed for a specific kind of future, one where AI agents run in the background for hours or days at a time rather than answering a human in real time. In that world, latency matters less and cost matters much more. And Neil has built the entire company around driving the cost of a token as low as it can possibly go. What makes this conversation special is that it's one of the most detailed tours I've ever done through the full stack of intelligence, the software, the chips, the power, and how the three connect. Along the way, we cover the trade-off between speed and cost that lives inside of every GPU, his scavenger strategy for buying the chips and power no one else wants, his contrarian view on NVIDIA, and why the premium the Frontier Labs charge for being three to six months ahead may not last. Please enjoy my conversation with Neil Mova. Speaker 2 (3:08) I think it's important early in these Speaker 3 (3:09) conversations to just say the thing, literally what you're building and what it does today. So maybe just orient us there with a brief description, like literally what the system is that you're building Speaker 2 (3:20) and why it should exist. CL Research is a token factory. We have an API where anyone can send us requests, where they can use large language models, open source large language models for any task they want. We will serve those tokens to them at a price that is unbeatable in the market. We also support their ability to build agents on top of this. We host what we call sale boxes, which are long-running agent virtual machines hosted in the cloud that are designed for agents that run for hours, days, or weeks. So I should think about Speaker 3 (3:46) you as a peer company to others that serve different kinds of inference. You're serving one specific kind of inference, and your goal is to be the absolute cheapest provider and enabler of a certain kind of use of intelligence. Speaker 2 (3:58) Exactly. The theme of our company is abundance. We want to deliver this new commodity of intelligence to as many people as possible at a cost that is sustainable for almost every industry. We think that whenever you make something 10 times cheaper, it's a new product category. We aspire to do that for tokens. We think it's so profound that the machine can think, and now our job is to make as many machines as possible in the world work towards thinking. Speaker 3 (4:20) So if you think about the theme of the day being token costs, is token cost the right way to think about this? Is there some other way you'd put it? Speaker 2 (4:27) To start with, absolutely, token cost. Today, my North Stars, I want to have the lowest cost per token in the industry and do that by a mile. I don't think tokens are the final unit of work or intelligence, but they are what we use today. After tokens, you start to move more towards more outcomes, which is like a vague direction. You can imagine, for example, today when you consume tokens through an agent, you don't actually control how many tokens the agent reasons for. It can reason for a certain amount of time or it can call a certain number of tools. And increasingly, I think we will have... agents do some unit of work, take as many shots on goal as they can. And however many tokens they use to get there is going to be a dependent variable depending on the task. So you think about like agents that self-administer a token budget as opposed to a company setting a budget for how many tokens engineers can spend per month. Speaker 3 (5:12) Why is there an opportunity that you can tackle? It seems like the entire world is oriented around more, better, faster, cheaper tokens right now. It seems like the world is trying to solve this problem very aggressively. What was the unique opening that you saw that's maybe the market's not being efficient in its attempt to tackle this? So I think there's two things that Speaker 2 (5:32) are tailwinds for our company. One has got to be the rise of open source. I had to talk about that first. I think we're starting to see an increasing number of our customers and the broader market care about owning intelligence. They want to have control, sovereignty over the thing that they depend on. That created a much more robust market for our customized models. Or even just like these vanilla open source models that no one can ever take away from you. You always have the weights. You always have the right to deploy them however you like. In that world, there's been a reasonably robust market for the past couple of years serving these models at large scale. The challenge is all those companies, you could take your pick, Base 10, Fireworks together, they all focus on low latency inference. And they were pulled in that direction by one very important customer, Cursor. I think that that was the right choice about a year ago. And as of six months ago, it started to look like maybe low latency wasn't the only thing you wanted from an agent. You wanted more persistence, more long horizon tasks. And now it's, to me, very obvious that the future of agentic inference is long horizon tasks. You're going to run the machine for hours or days at a time. It doesn't matter if it's been set tokens at 100 tokens per second. Maybe 10 is just fine if that comes with corresponding advantages and efficiency. Why are you so confident in that? To me, it seems like... I want everything as fast as possible. When you're waiting on it, you absolutely deserve the fastest answer possible. My trick is I don't want you to be waiting on it. I want it to be proactive. I want it to be in Speaker 3 (6:52) the Speaker 2 (6:52) background. One way to say it is like the best latency is no latency at all. When you wake up in the morning, the work's already been done overnight. You didn't even have to ask for it. That's the dream. We're not quite there yet. More importantly, I think the more you're in the loop as you prompt agents and wait for a response. In fact, you're the bottleneck in having the agent do more or less work. What we'd like is the agent to operate on more human time skills. You don't manage your colleagues every five minutes. You ask them to do a high-level task, and you come back and check in maybe every day, but more likely once a week. And that, to me, is the future of human-agent collaboration, more like human time skills. Speaker 3 (7:25) Say more about the early indications that this is happening, and therefore you should be building Speaker 2 (7:30) this company. So the first and most important thing is the idea of test-time compute scaling, the idea that you can give an agent more time, and it will give you a better answer. So that was theorized about two years ago now, but it wasn't really something that we could do. actually bet on until I would say late last year with Opus 4.5. Opus 4.5 was the first agent that was at all suitable for longer horizon tasks. It was pretty mediocre when it first came out, but you look at the more recent models and what we've done on open source as well, and you see that agents are capable of running for an hour at a time. I wouldn't say it's days, but definitely an hour is quite suitable today. Seeing that like average task length get longer and longer, it doesn't take many points to have you draw out the exponential and see that agents are worth running for longer periods. What do you think will be the market share of long-running agents in three years or something like this? I love this market because it's unbounded. There's no human in the loop, so you can consume as many tokens as you like in the background versus human attention span. If you tell me to consume 10x as many tokens at Codex or at Cloud Code, I'm actually not sure if I can anymore. I'm already in the loop and locked in coding for most of the day that I'm at the laptop. What is unbounded is how many tokens can be consumed in the background or proactively. Long term, I think we're going to end this year at maybe 50-50, background and real-time workloads. But I see this going to 90-10 in Speaker 3 (8:45) favor of background. What are your favorite examples of something that gets accomplished much better as a background task than as a human-in-the-loop task? Speaker 2 (8:52) Most deep research. Most questions where you want to have a definitive answer over not 100 sources, not 1,000 sources, but 10,000 sources. or more. If you want to build an authoritative index of information, like for example one of our customers Parallel Web Systems seeks to do, they want to build an index over the whole internet and they want to monitor the internet in real time for changes. That is the kind of crazy exabyte scale task that you need a very different kind of intelligence or scale of intelligence to achieve. Deep research is a top category for us and then increasingly we see cyber security following this direction. If you think about, yes there's so much code you can generate but there's exponentially more ways to break that same code. than it is to generate that code. There are some great customers out there who are working very hard to find agents that can break any piece of software and proactively patch them. When Fable first came out, for example, or Mythos first came out, basically there was this push in the cybersecurity community to run Fable against every line of code we've ever written and look for bugs in 20 different ways. Meaning you're looking for both memory errors, you're looking for business logic errors, and looking for network vulnerabilities. All these things. And these are all actually... things that you would write specialized agents for. You wouldn't just have Fable look at the source code once, you'd have it actually set up environments where you can pen test these applications. At some point, people started to make this joke that security has become proof of work. When you want secure software, it's really a question of how many dollars did you spend on Anthropix APIs trying to break into your software. That is the best indication for how secure it is, because that's the best tool in the world. And increasingly, we found that open source models, well, the frontier of intelligence here is quite jagged. It's not the case that Fable finds a superset of all bugs in software. You would find some bugs with a very small model that you don't find with a large model. You'd find some bugs with haiku that you would find with table and vice versa. So it encouraged this very diverse approach to sampling and trying to build cybersecurity agents that break software autonomously such that you can patch them. If you were Speaker 3 (10:45) to get speculative and imaginative about the sorts of things that very cheap, very long running agents can enable, we talked about some very practical examples, deep research. cybersecurity, et cetera. If you get a little bit dreamier about the use cases, new product category that this sort of inference will unlock, I guess the question is just like, so what? If you're maximally successful, dream a little bit about what that might enable. I Speaker 2 (11:10) think for individual users, what I'm excited about most is this idea of proactive, intelligent agents. You can imagine a Siri that is running in the background all the time to understand all the emails you received in a day, all the text messages you receive in a day. And It has a much more encyclopedic view of your life and how to be helpful in that life. Right now, there's still point solutions. You end up doing a lot of prompting. Siri is not very proactive. It's something we can fix with abundant inference. If you trust the machine enough that it's reliable and also trustworthy as in private, you might even imagine the machine can understand how you interact with it and proactively surface your next action. Whenever you open your phone, can we build a good model of what you're going to do next? My estimation is yes, we totally can. And the key to that is incredibly cheap intelligence. You have to be willing to spend tokens without any promise of return. That is the unlock. The long lens view to take on this is that we have a form of intelligence that can tackle any verifiable problem. Any verifiable problem means most software. It means a lot of formal math proofs and similar. And it could also mean scientific discovery. These are all... relatively verifiable problems. And all those things currently have a dollar cost attached to them, essentially, that's a hidden one. It's like, how many tokens could you possibly harness to make this work? We have started to bring it within view, a dollar cost for these long horizon tasks that is reasonable. It's not millions, it's thousands. And maybe it could be hundreds or even tens of dollars in the near future to have a definitive answer to any scientific question, to any research problem. Speaker 3 (12:39) So if we dream about that future, we then become limited just by the questions that Speaker 2 (12:44) people can ask, basically? Pretty much. The questions we can ask, the models are on the cusp of basically taking even a high-level question and chasing it down. Every possible follow-up, you can have the model essentially take that on its own. And the question is, what is your token budget? And we will solve the token budget problem. What about non-verifiable tasks? Those are basically the entire category of human taste into that category. We have not solved human taste yet, and I don't know that it fundamentally can be. I'm excited to be surprised here, but we are focused on very quantitative problems. We leave the quality of writing, we leave the beauty of art to people. Speaker 3 (13:20) Vanta automates security and compliance for over 16,000 fast-moving companies like Ramp, Cursor, and Harvey, keeping them audit-ready around the clock. It's the number one agentic trust platform, and it now helps companies like yours watch for the risks that show up between audits, across your vendors, your AI tools, and your whole environment. Every new tool your team signs up for, every vendor that turns on AI features, is an opportunity for something to go wrong, and most security programs weren't built for AI's pace of growth. The Vanta agent works like a 24-7 GRC engineer in the background, finding issues, drafting fixes for you, and cutting vendor assessment time by up to 50%. Whether you're a fast-growing startup or a global enterprise, Vanta helps you earn and prove trust. Invest like the best listeners get a special offer for $1,000 off at vanta.com slash invest. Speaker 3 (14:10) Ridgeline is the first end-to-end system of record with embedded AI for investment management firms running portfolio accounting, reconciliation, reporting, trading, and compliance on one unified platform. Firms are moving off legacy technology and onto Ridgeline because of how far ahead Ridgeline's AI features are compared to anything else in investment management software, which is why I believe that firms that come out ahead in the AI era will be the ones running on Ridgeline's unified platform. If you're serious about your firm's AI strategy, Ridgeline should be part of that conversation. You can request a demo at ridgeline Speaker 2 (14:42) .ai. Speaker 3 (14:48) All right, now let's talk about the very clever stack of solutions that you hope to build. Ultimately, they have this giant token factory, supplier of extremely low-cost intelligence. I think you think about this in terms of software, hardware, and power. Talk through what your master plan is to approach this challenge that's so different from what others are thinking about doing. We Speaker 2 (15:07) always have to start with software. Where is the opportunity on today's data centers to improve efficiency? And the first thing we did was we tried to build the entire LLM software stack. around peak GPU efficiency, meaning we're using NVIDIA GPUs. We wanted to squeeze out more tokens from the same chip than anyone else in the world. And that starts with the lowest level of programming, kernels. It's actually my background. I spent my whole professional life working on GPUs and kernels. NVIDIA was my first job while I was in college. I got to see how the Tensor Cores got to earn their right to be on the chip. What does that mean? Like, what is a Tensor Core? Tensor Core is a specialized unit on the GPU that accelerates matrix multiplication. Simple as that. There's been a long history of how we evolved that tensor core over time that we'll get into. And why is matrix multiplication so important? I cannot say that there is a divine truth of the universe that explains why matrix multipliers seem to be the atomic unit of computation. But one way I've heard it described to me is, well, it's a really succinct way to mix two blocks of numbers together and have them interact in some interesting way. That's as much as I can say about it. It is really convenient that linear algebra turns out to be a very compact representation of arbitrary relationships in data. So NVIDIA... Great graphics company. Had market share dominance in GPUs and gaming graphics for quite some time. And then starting in like the mid-2010s, they started to actually start these like Skunkworks projects to make the graphics processor more suitable for machine learning tasks that they were tracking. I remember actually reading some of the lab notebooks of some of my managers when I was at NVIDIA. They would visit these small ML conferences like ICML or NeurIPS at the time. They would just take note of these papers like, oh, this deep learning thing seems to be catching on. And what's really interesting is that these grad students are using gaming and video GPUs in order to train their large models. We should double-click on this and figure out what's going on here. By 2015, 2016, at least Jensen had the conviction to kind of double down on, hey, this usage of our chips is only going to grow. Let's start allocating more and more precious silicon dye area to this capability that seems to be emerging. Let's put the first version of Tensor Cores on the chip. So we're talking about taking this gaming chip. which is designed for painting pixels on a screen and adapting it to do matrix multiplies. It was early. You would be competing against the graphics teams, essentially. When you ask for more silicon area at any chip company, there's always competition for that. It is something that the designers guard so carefully. You don't ever want to invest in the wrong technology because that's opportunity cost that you could have allocated to some other functionality. We fought tooth and nail and got just a tiny bit of diarrhea, maybe like 5%, 10%, something like that, for the first generation of these chips. to get some amount of acceleration for basic convolutions, which were the fundamental operation for computer vision models of the day. And then we had a software team that was trying to squeeze all the performance we could out of the chip. And I think on that software team, which is where I worked, that's what actually taught me the most about the ethos that NVIDIA has around this term called speed of light. They always chase the speed of light for any piece of hardware that they make. It is so ingrained in every engineer's mind that... If the machine can do it, we're going to push the machine to the frontier until it does what we think. And the speed of light is the edge of what's possible. The speed of light is the edge of what's possible. Exactly. If we think the chip can run at this frequency and produce this many multiplies per cycle, we're going to get there. We're going to break every bottleneck and get to that peak level of performance. To this day, I tell all my engineers, we're chasing 100 % speed of light. I don't care about relative numbers versus the competition. I only care about absolute numbers. Speaker 3 (18:25) What are we able to do on the chip? How do we achieve that? Before we leave that chapter of your time at NVIDIA, anything else beyond that? cultural touchpoint that changed the way you think about things or that stood out the most about how the business ran back then or its Speaker 2 (18:37) culture? I have a ton of stories about NVIDIA. I could tell you a few of them. One of my favorites is that on the tenure side, a lot of people I worked with in NVIDIA in 2015, 2016 are still there today. That company has incredible retention and these are the best engineers on the silicon side, at least I've worked with in my whole career. They're extremely, extremely motivated and passionate. They believed in parallel computing as a concept through its various incarnations and have loved seeing the chip evolve. This is their life's work, and they're extremely competent in that direction. They're also a very frugal company. NVIDIA and all the Silicon Valley companies, after 2008, they had some cutbacks and perks. So no free lunch, for example. NVIDIA took it one step further. There was no free milk in the fridge. So if you wanted to drink coffee at NVIDIA and you wanted some milk, you'd actually have to chip in a dollar every month to the milk club, and the milk club would stock Costco milk in the fridge. And I remember that distinctly. We don't do that at sale. It's Speaker 3 (19:27) a frugality that permeates. So coming out of this time there, you get this experience of what it's like to develop more efficient usage of the underlying hardware through software. So link that to today's environment. The GPU is fundamentally Speaker 2 (19:40) a throughput machine. The GPU is happiest when you give it a lot of work to do and let it chew through that work at peak utilization of its compute units. But that's actually not the way that we've taken AI in the last couple of years. We've pushed AI to be an interactive chatbot tool is the most common form of AI usage today. In that world, you care a lot about actually spitting answers out to the person at the keyboard as quickly as possible. To your point about don't make the user wait, I want things as fast as possible. That's actually quite interesting for the GPU. It's