EP 65: The AI Bottleneck Nobody's Watching | Troy Liljedahl, Backblaze
Inside AsembleAI: DeepTech, AI & Science
This episode of 'Inside AsembleAI' explores the overlooked yet critical role of storage in AI infrastructure, featuring Troy Liljedahl from Backblaze. The
Key takeaways
- Storage bottlenecks are a primary reason GenAI projects fail to scale beyond pilot stage.
- Idle GPUs due to insufficient data throughput result in lost time and opportunity cost—such as missed chances to improve or monetize training data.
Main topics
- AI storage bottlenecks
- GPU idle time and opportunity cost
Notable quotes
Storage is not the sexiest thing to talk about, but it's a really important part of that workload.
Conclusion
Storage infrastructure is a critical but often ignored pillar in AI scaling. Solutions like Backblaze's B2
Transcript preview
Speaker 4 (0:00) What actually happens? The GPU just sits idle or does it like get forced that back? Speaker 3 (0:05) So yeah, I mean, you're right. GPUs can sit idle if they can't be fed fast enough, right? But there's an opportunity cost associated with that as well. Speaker 4 (0:12) What do you think is the main call trait because of which... The storage scenario can slow down any GNA adoption in any enterprise framework. Speaker 3 (0:21) There's only a handful of foundational model folks out there that are doing things. Everyone else is creating niche models. We have folks creating small language models, right? Very specific use cases. So it's very competitive. Anytime that... your bottleneck to your workflow, whether that's from the performance standpoint, it could also be from a financial standpoint. It could be do or die for your company. Speaker 4 (0:42) What are the main skills that they should hold on to survive and even like prosper in this current AI pivot world? Speaker 3 (0:49) Really understanding how to manage agents is where the engineering industry is going. What one advice would you like to give? Give it to them. Probably all of them were sold out of flash for the next two years, right? It's all been purchased up. And that's a very limited capacity. How Speaker 4 (1:04) advice do you have for any like fresh graduate out of school that if they want to kind of prosper in the IT industry and storage industry in the infrastructure world, we've seen a lot of new roles started coming up. We'll definitely see more roles are coming through in the next few years. Speaker 1 (1:21) Joining me today is Troy, Senior Director of Solution Engineering at Backblaze. With nearly a decade of experience in cloud storage and AI infrastructure, Troy brings a unique perspective on where the industry is headed. Troy, helping AI customers integrate cloud storage, solving data and infrastructure challenges. He is the company's co-founder and co-host, Assemble AI. Sam Day. Speaker 2 (1:51) Welcome to Assemble AI. Subscribe and follow us on YouTube, Spotify, Apple Podcasts, iHeartRadio, and Substack. Speaker 4 (2:00) Welcome back to another Assemble AI episode coming directly from the Podcast Pavilion at AI4 Conference. I'm your co-host, Sam. Today, I'm joined by Troy, Senior Director of Solution Engineering at BlackBlaze AI. Troy, welcome to the show. I'm really happy to have you here. Please tell us a little bit about yourself to our audience. Speaker 3 (2:24) Absolutely, Sam. I'm so thrilled to be here. So I run... the solution engineering team at Backblaze. What we do is we're responsible for helping customers in the AI space and really across the board integrate Backblaze into their solution. If you're not familiar with Backblaze, we provide cloud objects to them. We're really part of the AI data workload. And kind of where we fit is being a capacity tier, data lake tier. for customers that are building out their own models. We provide high performance, always hot object storage in the cloud, always available, no tiering, very simple pricing, no egress, no API transactions, none of that kind of stuff. And fully S3 compatible Speaker 4 (3:12) API fitted to all the workloads and customers we're doing today. The Speaker 3 (3:16) fun job of helping people set that up, try us out, running, pardon me, running people concepts. integrating Backblaze into their workflows. It's been a blast. I've been at Backblaze for about 10 years. It's at various positions and now by myself kind of leading this field engineering team, which has been super fun to work with. Getting to work with a bunch of different types of customers, helping solve problems that are really unique, I think, to where we are at this stage, kind of where AI has become so prevalent. It's been really, really cool. Speaker 4 (3:48) Thanks for enlightening our audience with the your background, what you guys are doing at that place. So now, moving on to the first question. As we all know, storage is the least glamorous, most overlooked item in the AI stack, right? This is true. Yet, it's often the actual reason that the Gen-AI pipelines stall between the pilot and the productions. So now, when it comes to those bottlenecks that you're talking about, what do you think is the main culprit? because of which that storage scenario can slow down any GNAI adoption in any enterprise further. Speaker 3 (4:25) Yeah. I mean, I think one of the things we all recognize is that the GNAI market is so competitive. There's only a handful of foundational model folks out there that are doing things. Everyone else is creating niche models. We have folks creating small language models, right, for very specific use cases. And so it's very competitive. Anytime that... You are bottlenecked in your workflow, whether that's from the performance standpoint, you can also be from a financial standpoint. It could be do or die for your company. And so what we do, where we come in, is to help people solve those bottlenecks and bridge those gaps. You're right that storage is not maybe the sexiest thing to talk about. Speaker 4 (5:05) But Speaker 3 (5:05) it's a really important part of that workload to getting from starting training a model. and then having a model that you can deliver and then sell to your customers. And planning out how that store works can be vitally important. Not just from a, hey, I have data I need to back up and store and protect standpoint. But from doing it in an efficient way, that means that you're not letting GPUs run idle. That you're able to go to market as fast as you possibly can. Speaker 4 (5:35) So that kind of brings me to my next question. Like, you know, with storage... and keep up with the GPU cluster. And what actually happens is the GPU just sits idle or does it get forced at that? Think about that there's an AI engineer listening to that trying to bring some technical lengths for the conversation. Absolutely. So Speaker 3 (5:54) yeah, I mean, you're right. GPUs can sit idle if they can't be fed fast enough, right? But there's an opportunity cost associated with that as well. It's not just the time that it's taking to complete a training graph. But there are the other things that could be happening during that lost time. There are perhaps improvement of the data, whether we're talking about labeling, cleaning, right? Making that data more valuable to the model training. There are opportunities to monetize that data, right? There's some folks here at Comforts today that their whole business is monetizing training data. That's something that other companies that have started out training their own models are thinking about as well that we talked with. And so... Really having something that's going to be performant in a way