Effortless Audio Cleanup for Broadcasts

Reshaping Workflows with Dell Pro Precision and NVIDIA RTX PRO GPUs

In this episode of 'Reshaping Workflows with Dell Pro Precision and NVIDIA RTX PRO GPUs,' host Logan Mahler interviews Jessica Powell, CEO o

Key takeaways

  • AudioShake's technology enables real-time separation of mixed audio into individual components like vocals, instruments, or speakers—even when they overlap or bleed into each other.
  • The company has won the Sony Demixing Challenge twice, demonstrating leadership in both technical benchmarks and perceptual quality for human listeners.

Main topics

  • AI-powered audio separation
  • Source separation in post-production

Notable quotes

"We make audio more usable... separating an audio recording in real time or in post into its different components." – Jessica Powell

Conclusion

AudioShake is redefining how audio content is processed and utilized across industries by turning complex, messy recordings into

Transcript preview

Speaker 2 (0:04) Welcome to Reshaping Workflows with Dell Pro Precision and NVIDIA, where innovation meets real-world impact in high-performance computing. Speaker 5 (0:19) Welcome back, everyone, to an exciting and exhilarating episode of Reshaping Workflows with Dell Pro Precision and NVIDIA RTX Pro GPUs. I'm your host, Logan Mahler. Well, it's ironic that we're having this episode today because I logged in to start the episode. My audio isn't working. There were things blocked on my computer. So ironically enough, we have a guest today that's talking all about audio. It's just how the world works. Life is funny like that. So with that, please welcome Jessica Powell, the CEO of Audioshake. Jessica, how are you doing? Speaker 3 (0:55) Good, thanks. Thanks for having me. Speaker 5 (0:58) Wonderful. Well, give everyone just to start a little bit, you know, your elevator pitch about you, right? Who you are, where you've been, what you've done. So Speaker 3 (1:07) my name is Jessica Powell. I run Audioshake. Prior to Audioshake, I was at Google really for most of my career. Did like a startup. and some other stuff here and there, but essentially was really at Google for a very, very long time and worked across a broad range of areas. But in my final years at Google, I was running communications across the company. Speaker 5 (1:29) Okay. Communications. Okay. Awesome. I mean, that's not what I was expecting you to say. Speaker 3 (1:35) Well, I can go a different angle. Wait, where would you? No, I Speaker 5 (1:38) mean, don't lie. Don't lie about it. I was just not what I was expecting. Sorry. Like a little bit of a frustrating word. So now the next question, move from Google, now AudioShake. Tell us high level, because I don't want to spoil the details. I got a bunch of questions. But what is AudioShake? Speaker 3 (1:54) Right. So what we do fundamentally is we make audio more usable. So we make it really easy to edit audio, interact with audio, access information that's inside audio. And we do that all through something called source separation, which is where we are separating. an audio recording in real time or in post into its different components. So you could think about it as sound separation or stem separation, but the field is known as source separation. So you take a track, for example, a music track, and you're splitting the vocals, the drums, the bass, you're taking a podcast and you're splitting my voice from your voice, those kinds of things. Speaker 5 (2:31) Okay. So I make some notes before the episode and it's saying that you kind of separate the audio like you just described. You know, I'm not an audio expert. Clearly, I can't get my computer running in the morning. But it basically says that that's an impossible task. I don't want you to share like technical details. You have to get too nerdy. But like how? Speaker 5 (2:53) They always record on separate tracks. Even I know that. No, most of the recordings in the world are not recorded on separate tracks. So, yeah. Speaker 3 (3:01) So Riverside that we're using will have us on separate tracks. But most other situations in the world, you are not. And even when you are recording on separate tracks, you might have other issues. So, for example, right now, your audio, because of some of the weird issues you were having this morning, isn't really perfect. And that's going to have to be cleaned up in post. Someone's going to have to remove the hiss and that kind of thing. If you were on a movie set, you might be on a high budget production where they have a lot of audio budget and everyone's mic'd up, but then your voice might bleed into my voice. If you're doing unscripted TV, you're doing reality TV, there's a ton of, say, commercial music bleeding into the background, a lot of voices bleeding into other voice tracks. You need to be able to separate all of that. And then if you think about actually all the world's data just like around you, the audio around you, no one's sitting there micing everything up. So we do a ton. And then there's plenty of workflows like in music where historically the different components didn't exist. And then even when they did exist, they weren't passed along as part of the like contracted assets. And even now, a lot of times they are not passed on or go missing or sitting on someone's computer. And so being able to demix them is incredibly useful. Speaker 5 (4:13) Yeah. So. I will get into kind of the compute piece here in a second. But, you know, if you are someone watching this, because we, our audience spans, you know, IT developers to ITDNs, like, tell us what this looks like. You know, for example, we bring a file to Audioshake and you just work your magic until we pass it back. Is there anything that's done on the end of the consumer or the end customer? Meaning, do I have to do anything to repair the audio file? Or can I just take this hunk of junk Riverside file that's probably terrible because I'm on a different computer? running Linux and stuff and just pass it to you. And then I get something back to further. Speaker 3 (4:48) Yeah. So it really depends on the use case. Essentially, we're B2B infrastructure tech. And a file that you might want to do, like, let's say you want to remove music from the background of your track because somehow your guest was playing music while they were doing the podcast with you and you don't have the rights to that music. That same file for someone else for some other purpose, they may want to retain that music. So you can do some upfront guesswork at what the user is trying to get, but you do need a little bit of... context or user intent to know what someone wants to do. So just to give you two wildly different use cases that use exactly the same model. If you, again, like if you were doing, if we were sitting in the studio together and maybe we're both mic'd up or we're not, but we're together in the same room, our voices are going to bleed into each other's tracks or we're going to be on a single track. AudioShake would be able to separate our voices even in the moments where you and I are talking over each other and we have like cross talk or overlap. That same technology which right there, what you and I just described would be what we would call post-production, right? We're just going to process the entire file. You could do that on our on-demand site in the cloud where you just pass it through. That same technology is going to be also used in ASR and transcription and captioning workflows at scale where you need to, again, you've got people that are talking over each other that are talking to an AI. It could be an AI agent that they're talking to. It could be another human they're talking to, but you have a lot of issues of. People talking over each other and you need to be able to disentangle that overlap. So totally different. So if we're talking about inference, totally different in terms of the user profile and the end result. In one case, you want that output because people are going to listen to the output, that podcast case. In the other case, you're talking about, you know, millions of minutes or hours and no one's going to listen to that output. Instead, that output is going into downstream workflows like ASR. So very, very different. Speaker 5 (6:39) Very different. Okay. And it's funny because you say talking already. That's Speaker 3 (6:42) what Speaker 5 (6:42) my wife says. I do all the time. I was talking too much. So you won the Sony demixing challenge, which is like an industry award. Can you describe that? Because I feel like that's kind of a big deal. And tell us a little bit about that. Speaker 3 (6:55) Oh, sure. So Sony, I don't know if they're still hosting it, but they have held a contest a few times called the demixing challenge that we've topped the leaderboard on both times, which is trying to find the best demixing technology. One year they did it for music, another year they did it for music and film. There you have, you know, there are benchmarks in the field. It's called the SDR score. And that's really what you are competing on against other teams, including big tech. We also spend, I mean, SDR and those kinds of benchmarks are important. But I think particularly you have a developer crowd. Anyone say working with voice AI models, for example, know that there are the benchmarks and then there's reality. So we went on the benchmarks, but just as important to us is actually how these things sound to real human ears, not just how they sound to a machine. And so we put equal, if not more weight into our own perceptual benchmarks and so forth. So the Sony thing is very cool and we're very proud of it. But what's more important to us is that our customers, you know, that we meet our customers quality. Speaker 5 (7:54) Yeah, that's who's buying it. But still, I mean, you're humble as well. But Speaker 4 (7:57) I mean, I mean, sounds like a pretty cool thing, at least when