95% of AI Agent Projects Fail to Reach Production. Here's Why | Manoj Saxena, TrustWise
Eye On A.I.
Manoj Saxena, CEO of TrustWise, discusses the critical challenges facing AI agent deployment, emphasizing that 95% of AI projects fail to re
Key takeaways
- AI has evolved from generating outputs to taking autonomous actions, creating new risks in enterprise environments.
- Multi-agent systems are becoming standard, but lack of runtime governance leads to potential chaos and misalignment.
Main topics
- Runtime governance in agentic AI systems
- The trust gap between intent and action in AI models
Notable quotes
We're building bigger and bigger nuclear cores, but no one's thinking about putting a dome on top of these things.
You've got digital workforce that's being introduced that can act now. These actions could run for minutes and hours and days.
Conclusion
Manoj Saxena argues that without a robust runtime governance layer—like TrustWise's AI control
Transcript preview
Speaker 2 (0:00) Trust has become an issue. It's not simply trusting that the model is accurate, but trusting that the model is going to do what you want it to do and not what it wants to do. Last Speaker 1 (0:11) month, the first time ever, The traffic on the internet, agent traffic exceeded human traffic. Just like when on the AT &T network, data traffic exceeded voice traffic. This is a very big deal. What I saw was this focus on building more and more intelligent models and bigger models, almost like you're building bigger and bigger nuclear cores, but no one's thinking about putting a dome on top of these things. That's what got me going with TrustWise about three, a little over three years ago. That's an interesting Speaker 2 (0:37) analogy, a dome over a nuclear core. Speaker 1 (0:44) Why Speaker 2 (0:44) don't I have you introduce yourself to listeners and give as much of your background as relevant, of course, the IBM Watson thing and how you got to TrustWise. And then we'll talk about the runtime governance and the control layer over agentic AI systems and all of that. Speaker 1 (1:12) I'm Manoj Saxena. I'm the CEO and founder of TrustWise. This is my fifth startup. So some people call me a certified masochist. But the reality is that I love building things and love building companies and assembling people together to go after important problems. And I've been working on the intersection of enterprise AI, responsible AI, and... operationalizing trust in digital systems for a long time, going back to the IBM Watson era about 10 years ago. I had the privilege of building a startup that was acquired by IBM. And then after that, the IBM board asked me to commercialize the Jeopardy playing game called Watson, which I, yeah, so I had, I spent about three years taking that game that was a size of a... master bedroom. And I live in Texas. It's a big master bedroom. And reducing it down to one pizza box level system that we deployed. And so, yeah, that's kind of been my journey. I was teaching responsible AI at the University of Texas, Austin, and seminars at Cambridge University in London when ChatGPT got launched. And what I realized, I was surprised how fast this technology had moved. And what I saw was this focus on building more and more intelligent models and bigger models, almost like you're building bigger and bigger nuclear cores, but no one's thinking about putting a dome on top of these things. I decided to put on a jersey and come back and said, this is a problem that needs to be solved. Otherwise, my grandkids are going to ask me, saying, you were there. You made all this money with your companies. Why didn't you solve it? That's what got me going with TrustWise a little over three years ago. Speaker 2 (3:03) Yeah, and that's an interesting analogy, a dome over a nuclear core. You're focused on building a platform that sits above everything, that is a control layer both for generative AI and agentic AI. Is that right? And before we get into that, can you talk about... why trust has become such a critical issue, trust and governance. And I'll just say that personally, you know, this is moving so fast, you know, now with mythos, you know, from Anthropic, which is so powerful, they can't even release it. the issue of trust goes in many directions. It's not simply knowing, trusting that the model is accurate, but trusting that the model is going to do what you want it to do and not what it wants to do. So can, anyway, talk about how trust has become an issue, how it's changed in the last year or so. Speaker 1 (4:26) Absolutely. I think you nailed it. At the end of the day, trust cuts down to, does the AI and the model do what you intend it to do? That's kind of the heart of, is it aligned to your business and personal intent? So the alignment problem is not spoken about enough. And the reason trust has become more and more important is three things. Number one, AI has now moved from generating output to now taking actions. Watson and deep learning and chat GPT is not just giving you a better email or a better picture. Now it's taking actions on your behalf. And these actions could run for minutes and hours and days. So essentially, you've got digital workforce that's being introduced that can act now. Second, enterprises and enterprise AI is moving from a single model workflow to a multi-agent system. So the rise of things like OpenClaw, I think OpenClaw is going to be a thousand times more impactful than ChatGPT was in terms of its impact on business and society. So businesses are now moving towards applying not just single model, but multi-agent system, multi-model workflows that if things break, you could have real chaos on your hand. And third is policy in terms of driving the behavior of these systems. is no longer something you can enforce only at deployment. It has to be continuously enforced at runtime. So it's not enough to say, well, you go do this. You've got to make sure that you're checking every tool call, every action, every output. It is staying on policy. So as a result, the architecture has now become the risk surface. The whole architecture of how you're building an AI stack has become a risk surface. And that's why... Trust becomes like the silver thread that runs through your entire stack. And there needs to be a new class of infrastructure to manage and deploy these agents at scale. I sort of talk about it almost like an HR department of agents. You know, you would not deploy a company today with humans in it without an HR and finance department. And now you're about to introduce a whole bunch of digital labor. with no HR and finance, with no drug testing, with no employee manuals on how to behave, with no performance appraisals. So we look at what we're doing at TrustWise. We call it the AI control tower. It's almost like the HR and finance department for hundreds of thousands of agents from multiple vendors. That Speaker 2 (6:57) layer of the control tower, is that in itself an agentic system? That's a Speaker 1 (7:04) great question. And yes, at the end of the day, when you have to scale, the system, you'll need AI to control AI. But the difference is that the agentic system that we deploy, we call these guardian agents, but these guardian agents are built with human in the loop, and these are built with deterministic outcomes. So it's not probabilistic systems. So what we are looking at is a whole new class of agents and humans working together to be able to define, assess. control and then optimize these systems as they're running because they're not enough humans. And one of the large global companies told me by the end of this year, they'll have 100,000 agents in the company. And you're not going to have enough employees being able to onboard each of these manually and test it manually. So in the control tower, these guardian agents, there are two types of agents, guardian agents and genesis agents. Guardian agents make sure that these agents are onboarded well and they're working well. And genesis agents... convert some of these agents into super workers, what is called as AGI. So within our control tower, there are two types of agents, guardian agents and genesis agents. But most of the focus right now is guardian agents to make sure that the agents are aligned and are behaving properly so that you build the confidence to be able to then launch them. Speaker 2 (8:26) This is to provide runtime governance. I mean, it's monitoring what agents do in real time. Is that right? Speaker 1 (8:35) Yeah. So one of the issues is this is a wide open space today because runtime control means the control decision happens at the moment of action, not weeks before in a policy document or not hours later in an audit log. So today, if you look at the problem, there is a wide space here. There are three types of software that enterprises have. None of them are able to manage runtime control. One is security. And security mostly is outside-in defense focus, right? And agents are the new threat factor. Agents are the new insider threat. And there is no software for that. I like to say that I can build you the world's most secure prison and defend it from outside. But if you have a bunch of Chucky's and Hannibal Lecter's on the inside, you're still going to have chaos on your side. So security doesn't protect it. Security is mostly outside-in and not inside-out. Second, governance only defines policies, but it doesn't implement policies at runtime. What I call as declarative governance versus runtime governance. And third, observability tells you, that's the third class of software that companies have, it tells you what's going on, but it doesn't let you influence and shape it. So one way to think about it is you're launching these supercars into your company with a giant, you know, a thousand horsepower engine. without any steering wheel, without brakes, without seatbelts, and without emission control. So this is the area, this new layer. To give you an example, each of these agents that you deploy have to perform differently. So in the UK, there is a law called the UK FCA, Consumer Duty Law, for distressed customer. So if you have a financially distressed customer, say an 85-year-old who lost his wife is applying for a loan. versus a 24-year-old who just got a new job and she's applied for a loan. By law, the tone and clarity and helpfulness of your agent has to be different. Today, there are no systems that are able to differentiate and drive, like I call it the steering wheel, to align the behavior of the agent at runtime to meet the requirements of those different individuals. So that's an example of runtime control, which means the system checks the action before it happens. Is this customer financially distressed? Is this action allowed? Is this tone allowed? Is there a human approval required? Is this action is not stopped? If it's something I don't have an answer to, do I say no to or do I escalate it to a human being? All of these things is what the runtime control implements at the time of action. Speaker 2 (11:18) Yeah. And you have something called modular AI shields to watch agent actions and constrain tool use if necessary and ensure compliance. with rules and regulations, as you mentioned. How do you differentiate this modular AI shield from a guardian agent? Are they the Speaker 1 (11:45) same thing? Yeah, great question, Craig. So think of guardian agent as the AI that is overseeing and supervising different worker agents or functional agents. And that guardian agent can be equipped with one or more shields. And so it's like adding on a superpower pack. If you're Mario Brothers, you're adding a superpower pack. And broadly, there are three types of shields. There are shields for safety and security. There are shields for compliance and regulations. And then there are shields for cost and carbon management. So typically, when you are guarding these worker agents, you are looking at combining all these three together. You want it to be safe and secure. You want it to be compliant with internal policies and external regulations. And you want it to be cost and carbon efficient. All three together is what we define as trust posture management. And that's the new space we believe is required is you need to be able to manage trustworthiness of these agents as they are acting at scale and be able to make those tradeoffs across safety, compliance, and efficiency. And that's what the shields allow you to do. Speaker 2 (12:56) I see. And the shields are, they're not agents themselves. They're just watching who are they talking to when it is something that's questionable. Does it send it to the guardian agent? Yeah, the shields are always Speaker 1 (13:13) attached to a guardian agent. I see. So shields by themselves don't do anything. The shields get plugged in like a breastplate, like an armor. So a guardian agent will have multiple shields, which will give it the mission as to what it is allowed to do in terms of controlling your AI workforce. And Speaker 2 (13:34) this is all within a platform. How does the human interact? Because you said, depending on the policy or action, you may want a human in the loop. Is this a dashboard that sends alarms when a human needs to, or is it something more hands-on that the human is watching the agents perform? Speaker 1 (14:09) That's a great question. So the product is a platform, but the platform is a collection of APIs and what we call a CLI's command line interfaces. So it is not yet another dashboard in a UX. It's a collection of APIs that you can call and orchestrate because people are not looking for one more pane of glass. That's right. I already have my security infrastructure. I want to plug this into security, into governance, into observability. So at the core level, the product is a collection of APIs and CLIs, but... We do have a reference implementation of a dashboard. We call that the UX of the product. And that's where the humans will go in and they primarily do four things. Number one, they onboard their agents and the workforce. So this is where you connect, you point to it, you discover it, you bring them on board. Number two, you then assess them and to say that you evaluate them. You say, okay, I'm going to put you to this job at a customer support board versus... The other one, I'm going to put you to work for complaints handling. And the third one, you're going to do HR questions and answers. So based on that, we will then apply the right policies and controls, and we would run evaluations on it, like a simulator, like a Formula One simulator in a car. You would run the agent through hundreds of thousands of flight paths. And then you will see where is it breaking, what policies and controls need to be changed. So one, you're on board. Second, you assess and evaluate. Third thing you do is you then put that into production and you