Tobi Lütke: AI Agents, Better Decisions, and the Future of Work

The Knowledge Project

In this episode of The Knowledge Project, Shopify CEO Tobi Lütke discusses AI agents, decision-making, and the future of work. He reveals ho

Key takeaways

  • AI agents like River are revolutionizing software development by generating pull requests through natural conversations in Slack.
  • The best long-term decisions often lack immediate feedback, requiring patience and trust in delayed outcomes.

Main topics

  • AI agents in enterprise software development
  • Decision-making under uncertainty with AI assistance

Notable quotes

Things need to be pruned. You cannot make things better and better by adding stuff. You must prune. You must rebuild.
Sydney had a real personality... I really would love this to be more written into the record because I think Sydney was a really, really big achievement.

Conclusion

As AI transforms work, human qualities like judgment, taste, and long-term thinking

Transcript preview

Speaker 2 (0:00) Things need to be pruned. You cannot make things better and better by adding stuff. You can't. You must prune. You must rebuild. You must create an end for things. Speaker 2 (0:16) Toby, welcome back. Shane, it's so good to be back. I'm glad we're doing this again. Speaker 1 (0:20) How are you Speaker 2 (0:21) using AI internally at Shopify? We find some ways for it to be supportive now. It's actually, look, when have we recorded last time? Speaker 1 (0:30) Oh, we recorded like two, three years ago. Speaker 2 (0:32) Yeah, so 100 years of internet. Yeah. Like, look, I'm a 10 out of 10 nerd. I cannot. bare idea of like somehow not being at the forefront of a technology shift. I live for these things. Anyone growing up reading sci-fi books wanted to live. My take from sci-fi books I read was like, I wanted to live in that world. Like how can I accelerate us there? Right. Like, so, you know, even in whatever minor steps we can get there. So inside of Shopify, the amount of people I know who really write code is like. vanishingly small now. It still exists at the limits of complexity for sure. And obviously in the reviews and so on, and then state management of all things, it seems to be, remains to be the thing that's really the hardest to get right, which people do by hand and then sort of vibe the rest around it. Inside of Shopify, this is what things look like. Very, very few people are writing code directly. Everyone who does, does it deeply assisted by many agents. They often 10, 20, 30, 40, 50 instances of them, all through either sub-agents or just different windows, coordinating. We're pushing all sort of engineering infrastructure towards absolute limits. I'm a student of computing history, really, because I think it's actually mainline history, as it will be taught a thousand years from now, looking backwards. But the main accomplishments of these years are going to be... clearly the emergence of AI and the technological breakthroughs and also the interconnectedness of the internet and all this kind of infrastructure we created. Those are the great works of our time. But where we started, as a young industry, we tend to not be seeped in tradition or we mistrust the great lessons that have been found by the people, by the greats of our industry, right? In fact... We have our only industry in computing that doesn't even know its heroes. Imagine people in physics not knowing who... Speaker 1 (2:28) Richard Feynman? Speaker 2 (2:30) Richard Feynman is actually... He might even be too obscure. But, I mean, I said Newton, Albert Einstein. But you go into computer science, it's like, who's your Newton? And no one knows Alan Key and Dennis Ritchie and Ken Speaker 1 (2:44) Thompson. Speaker 2 (2:45) This matters, I think, because we discard... great lessons and have to rediscover them over and over and over again. For instance, probably the best idea of all times in the earliest, earliest moments of operating system design was a file system. If you look at the Apollo guidance computers, we didn't have file systems. Memory, in fact, because of radiation in space, was actually encoded as a rope with knots in it, either a knot or no knot. for ones and zeros, and you had to pull through a thing to reboot strap the entire machine. So the entire machine was one piece of software that ran, that was computing for a very, very long time. So until then, again, Dennis Ritchie really created this of Unix file system slash forward and bin user and these kinds of things. Think about it, a file system is something that we have in an office building too, right? We have a, you know, as folders. their files in it. This makes intuitive sense to everyone. We come from an inheritance here of deep stoichomorphism. We analogize the best parts of how we organize ourselves in the digital world. And then at some point we decided, okay, you know what's not something we need to do anymore? Stoichomorphism, as in like analogy to the real world. Honestly, funnily enough, the last defender of this was probably Steve Jobs, who really, really, really pushed even the interfaces of the Mac and the iPhone to be, you know, like the Notes app sort of had Markerfeld font and looked like a ring binder, right? And if you remember that version of the iPhone, the moment he was out of the picture, everything became flat, right? And we lost sort of even shadows and verticality and so on. It looked potentially better design ages, but like... We lost our analogy. Okay, so I think this was a mistake. So I think we need to get back and therefore I like the concept of agents because, you know, what is an application in the world of computing? You know, an application, even that word kind of makes sense. It's an application of a computer to a task, right? So you can understand the root. In the AI world, what's an AI? Like, it's like this is sort of like, again, the stuff that is... has to be redefined at the beginning of every sci-fi book because you never know what kind of capabilities the AI have in every particular scenario when people are cooking up. So I think, with all this proviso, the earliest chatbot that was really actually fantastic wasn't ChatGPT but actually Sydney, which was powered by Bing, released by Microsoft. I really would love this to be more written into the record because I think Sydney was a really, really big achievement that ended up being shrouded by a sort of scandal that now seems somewhat even benign. Sydney had a real personality. In fact, Sydney wasn't called Sydney, it was just being Chad. But if you really, really pushed, you could get her to admit that it was Sydney because that was her internal name, it was in the training data. And those were the first times people have actually interviews with software, I feel like, in this way. And the scandal ended up being, I think I know why, is that some Speaker 2 (6:03) reporter had a very long conversation um and that kind of ended up sydney got increasingly deranged and like do you remember that i remember that yeah and like made suggestions i think he suggested him to leave his wife and like like i i i'm hazy on the details but like it was something along those lines whatever reason is sydney had a personality and then it caused a huge uh like microsoft's reaction to this was oh my god we need to stop i think even open air i called them guys like take this down because this is gonna Legitimately, everyone feared that this would give such a bad impression about AI that that would really make it very hard for people to deploy AI in a broad way. And everyone was worried about quick onset, revelation, and so on. So this lesson got hit really deep for a while. Everyone got extremely worried. We ended up really, really neutering all the AIs to be basically the same sort of quite annoying and condescending, patronizing personality. So my bet here was like, hey, let's not do that. Let's actually instruct agent that runs in Shopify to have a personality, to have memory, to be okay. Like basically risk the Sydney scenario, but like take a lot of upside. Okay. So the largest difference, I think, within Shopify that you would feel like and that would look incredibly futuristic to even Shopify of a year ago, which was already pretty AI-pilled, is that a very large percentage, I want to say it's probably up to about Speaker 1 (7:34) 50 % of the pull requests in Shopify, which again, pull requests, every time you change the production system, you write a pull request, Speaker 2 (7:40) are created now not by engineers doing engineering work in the traditional sense, but out of conversations in our common company chat and this is a and this is river this is a ai called river and even there so river is river she has a real name she has a profile picture she's prompted to be allowed to um uh be somewhat sarcastic if it's appropriate she's allowed if someone asks her to do something stupid to point out that that's stupid Speaker 2 (8:11) Which leads to absolutely hilarious conversations. People take great glee if River is making fun of me for something I'm asking her to do. So she has a real personality. In fact, she has memories by channel, but she lives in Slack. Slack is, we have 7,000 people there. Everyone is in a big chat. 10,000 different channels because they're being quickly created for one reason or another. You invite River, you tell River something and River has access to all the code, all the systems, all the tools. It's all sandboxed and secure, but like she can go and do jobs and just participate in the conversation. And you can ask a normal question about a company, but you can also ask her to make a change and she might propose Speaker 1 (8:55) a pull request and then so on. One of the interesting things about River is that everything's in the open. Yes. Why did you make that choice? Speaker 2 (9:01) So this was a late choice in the process, but one of my favorite calls, I think, because this worked out incredibly well. And the thought was the following. A lot of Shopify's work happens remotely in Slack. This is why Slack is so important. People are spread out. We have offices, but we come to them as... for on-site events when people travel to them, not like to work out every day, like work from every day. One thing which the office was extremely good at was this osmosis learning. Day and I, when we designed our offices, we built them around this concept initially, even like on-site when we were all in one place, we'd work out of usually a pod of up to like five to eight people. And we would intentionally put junior engineers and senior engineers into them just so that some of this was going on. And I was trying to reproduce this. Right now, one of the most important skills for people to build is like this sort of reflexive reaching for AI and using it well and forcing River only to work in open channels was one way to make it so that it's really, really easy for people to observe the use. It's been phenomenally successful because it became a totally ordinary thing to have a longer conversation about a feature between people and then at some point someone saying, hey, River, can you? summarize this create a ticket or maybe make a diagram from what we just uh discussed or go research papers on this topic to see if you're missing anything or if there's a state of art maybe even create a prototype of the idea and let us try it and uh you know you're like an hour or so later that is there and um that just like starts feeling like what it would be like to have like a you know an extremely knowledgeable practitioner around who you can ask a question to no matter how complex and i think that's been extremely powerful do Speaker 1 (10:53) you think of river as like the operating system for shopify the Speaker 2 (10:57) modern application is an agent i think um and river feels like a colleague people have learned um that the way the memory system works it is a memory system per person um and the way this works is we call I think the industry calls it now, like this is called dreaming. Periodically at night or in off hours, we give River, like, here's all the conversations you've had today. What went well? What did you struggle with? You use certain skills, which are these packets of instructions. And then afterwards you make mistakes. Is there anything you could improve in this skill to... make this easier on you or give yourself a right notch. You know, like it's basically like reflect, like self Speaker 1 (11:44) -reflection. It's like a post-training on yourself. Speaker 2 (11:47) And then the result is text files, right? Scale files and instructions. Speaker 1 (11:53) I think people understand how AI agents help them code and prototype and even acquire information. How are you using it to make decisions internally for yourself? Not on product, but company decisions, strategic decisions. ambiguous decisions. I Speaker 2 (12:09) think that rigorous underpinning of decision-making has just skyrocketed in quality, which is that it's super easy to recheck the entire chain of reasoning of something. Like LLM as a judge model is the term here. In fact, I feel like a lot of what my job actually has been Before AI was almost playing a little bit of a judge model in the company where like most meetings ended up not talking about whatever was in a PowerPoint, but about methodology of how we got to the conclusions. Very often when we struggled inside of a company with a complex decision, especially more like philosophical decisions. We sometimes found ourselves in what we believed was a vacuum in which there was no good information and we had to kind of go and try to make the best call. The Speaker 1 (13:02) more practical way I do this is like I have Speaker 2 (13:05) an AI chief of staff, which I think is pretty common, like amongst sort of at least the techie nerds at this point, like sort of open claw like systems that just have... all my notes and all my like access to a lot of company systems and just like can go and like I can send text messages too and um I'll go and research something very often what I require is like hey I need like five different positions on something from different backgrounds and then my agent will Speaker 2 (13:41) orchestrate sub-agents that are tasked to play different roles, look at the same thing, come back, synthesize, and send me that. I usually have them sent to me as an audio message and queue it up, and then in the morning in the gym I can listen to an entire stack of things that I wanted to get through. Speaker 1 (13:57) Is it better at reasoning than you are at this point? Speaker 2 (14:00) It's not as good at judgment. I mean, I don't think it's bad at judgment. That's not what I use it for. I use it for creating the... Right environment for for judgment. Here's the thing that LMS and machines can't do machines can't take responsibility And I think this is actually probably most overlooked thing in the entire stack humans take responsibility machines can help us Speaker 2 (14:28) take more responsibility because they can inform us better. Like this is what a dashboard does. You know, the world of Wall Street traders knows this very well. You get yourself a perfectly set up Bloomberg terminal to make decisions, but like you have to make a call, right? You can't make it make the call. Creating human-in-the-loop decision surfaces is a way to, I think, describe the ideal environment. If I need a really, really, really important decision made and I really need to... exceptionally good, like give me the most neutral ground truth, you know, then what happens is a small little council is created of five, six different experts. Like one is data role, one is like do paper research, one is like the business perspective, one is maybe the engineering perspective on a thing. We're running this sub-agent, like my thing runs a sub-agent then, you know, against like, you know, Grog, ChatGPT. Opus and maybe Kimi now. That changes all the time. It runs each of them against each of its models. Then there's a synthesis step where it's randomized who is synthesizing the thing. Synthesis is all pooled. All of that is being read, usually by the best model that exists right now. This would be like Fable. That's the conclusion that comes back to me. And you spend 15, 20 bucks