ENacross all this stuff the virtual scavenger how do you think about that question of like which which type of these businesses to be You're doing real life factorial basically. we have to be more pragmatic. I think that the capital we're we're looking at for owning everything is like you said, it's insane. Yeah, software has high leverage, so we have to start with software, but ultimately, you know, do we own power generation or can we get great power purchase agreements with uh utilities? I'm more inclined to pursue like letting other people specialize in the things that they're historically good at and then see if we can get to the scale. I think of it as like I want to get to the scale where I earn the right to take this under our wing. I absolutely think that there's efficiencies to be gained everywhere in the stack. If you can break the assumption that people I would be buying from, they made assumptions about who their customers would be. And I maybe break those assumptions. It's a pretty optimistic view. Uh I think it's only possible because we're actually trying to underwrite the largest market for compute in the history of computing. we're actually going to build so many billions, trillions of dollars of investment into inference. Uh, and because of that focus, it makes sense to build a lot of things that are custom for inference. And it's my job to seek all the places where that's possible. And then as as they become obvious to me and my and my partners, I will get my partners to build custom things for me. And if they can't do it for me, I will do it myself. Yeah. I think compute scaling is actually like very efficient. Uh as in like you give me more flops and I will use more flops. And I would say we're actually fairly judicious already with our use of flops. Uh if you look at a modern model, there are very few models that are more than 10% dense, meaning 10% of the possible number of experts you can activate are activated. And I think the frontier models are closer to like 1%. So fairly sparse already. I don't think that we're wasting too much on the MOE side. People have been working with for quite some time. They're pretty good at squeezing. Where we are not good is attention and its use of memory. Specifically, the KV cache is quite uncompressed right now. I think if you look at the entropy in a KV cache, it's nowhere near it's not earning its keep. Like we're storing many kilobytes of data in the KV cache per token. Um, and that's probably off by an order of magnitude or two. And I I don't know what the Frontier Labs do, but Deep Seek certainly publishes really interesting work to compress that further and further. And they're making good progress. And I think the fact that they're able to make order magnitude progress here every year or so signals that there's a lot more room to go. I guess this all on the micro scale. If you zoom out further, I think that we actually don't marshall our compute effectively at all. Like we have all this compute in the world. Nvidia is pumping out 5 million Blackwell chips this year. Where are they all going? Are they all being used at all all the time? I certainly doubt it. I think that at some level we just need better orchestration of compute across the world. Uh this is very difficult to do because a lot of the comput disappears into private pools of compute that will never see the light of day and those GPUs sit very sadly idle. Uh it's actually it pains me physically to see that those GPUs are just you know silicon and power went into that and it's just sitting idle and I want to fix that. how we or organize and orchestrate the world's compute as a shared resource and and pack it more efficiently. I would estimate that, you know, we all make fun of XAI for having, you know, some challenges with total flop utilization on its clusters, but um the reality for the rest of the world is it's far far worse. A ton of GPUs just sit in warehouses or sit in private pools allocated to a specific customer um just don't get utilized. That's way more effective. Yeah. Um, what about fabs? Like what do you think is the future of fabs themselves? Like I think everyone is wondering be able to how will they expand capacity basically? Will we do it here in the US? Yeah. Well, it's interesting. Everything grows in balance with each other, right? If we snap our fingers and double all those things, you might fix a TSMC bottleneck that there you're just going to run into another bottleneck. You make 20% more chips, then you have another bottleneck immediately. I will say though, it is interesting what they consider to be a mustd deliver. uh like what they consider to be like an invariant that their customers me are always going to want versus what I think of as like a more fluid relationship. I think that if a fab exposes more of their trade-offs to me, I'm able to make more intelligent decisions about what I think I can I can do. Can we talk about how you designed the system of your own business? But like bring me into the culture and how you structure a team and a business where this is the northstar. Curiosity. It's 100% curiosity. You know the one thing I cannot teach is love for performance, love for uh digging into every microscond that the machine is working and understanding what's happening on the machine at that time. That to me is the most important trait for a performance engineer and it's what I look for. I don't look for lots of AI experience. I don't look for, you know, CUDA experience at all. That's actually a huge red herring. I mean, CUDA as a concept or GP as a concept have evolved so much in the last 5 years. There's no point asking for 10 years of experience. I want to teach that, but I cannot teach the love for performance engineering. That is what I seek. in a line. I would say the labs pay an immense premium to be 3 to 6 months ahead of of everything else. Uh and I think that's probably still worth it. I think it makes perfect sense for open and anthropic to do what they do. You know there's a sensitive topic around distillation which I think is part a very core piece of the relationship between closed and open frontier. And you know I'd like to offer an alternative view on that which is there is the sense that distillation is theft that you are taking something from the frontier models when you distill on their outputs. And in fact, even if that's not your intent, even if you don't ever try to, you know, scrape data from anthropic, one thing I'll offer is that an increasingly large percentage of the artifacts we put out on the internet are AI generated. Even if you just look at GitHub alone, you know, what percentage of repos created in the last year do we think were created by cloud code? Um, do we consider that to be distillation? Because that's probably all we need. I would not be surprised if you could train a fable glass model only on the outputs of code you consider good on GitHub that's open source. And certainly if we take the position that users own the outputs of their interaction with AI and they choose to put that up on GitHub, which a lot of them do, we're going to have latent distillation for a long time. It seems fundamentally impossible for me. Like I I don't think it's fundamentally possible to prevent the diffusion of of information or model capabilities. It will happen. The question is just how fast. And so then the question becomes, do scaling and improvement laws hold forever or for a really long period of time? And if they do, then there's value to being three and six months ahead and that will just last as long as it lasts and they can charge a huge premium for those tokens relative to a very cheap open source token. Is that the right way to think about it? And so your hope of what the future looks like is what like what balance between closed and open, you know, what balance between model companies doing everything because they have the advantage of owning the stack or whatever. You know, Enthropic can do that, you know, is like the new Google could Google just do that or something. What do you hope the future looks like? every company every user even make the agent your your own. Uh I think we're we're very not that far away from that level of customization and capability. I want people to own their intelligence and I want that intelligence to be customized probably not through weight fine-tuning but probably through more in context learning. That's a more technical detail. But the underlying input to this abundance future is about is basically cheap tokens. My job is to make the tokens as cheap as humanly possible. I will achieve that and I will do it through every layer in the stack available to me. I love the supply side levers. I will use every chip. I'll use every source of power and I will use every piece of land in the United States that's you know suitable for this. And in return, people will have the incentive to explore what it's like to have abundant intelligence. We still treat the agent as a person that is expensive to consult and you should ask them when you have a hard question. That's not the way to think about intelligence. It's incredible that the machine can think and we should try to get that into as many hands as as many people as possible. Most of the ideas on chips, I would say. You know, when I talk about building custom chips and they ask me, "Oh, so what's different?" Basically, it's it's about sidestepping the HPM shortage and focusing on more extreme offload to other forms of memory such as flash. Um, I'm quite passionate about that idea. Everyone on my team knows that I keep banging the drum around like what would we have to change about the model architecture to make offloading KB cache to flash work at a much greater level. And um I'm whiteboarding that all the time. That's like in the community of like inference people. you know we have some divergent views on what you can do if you design a system around serving at you know one to 10 tokens per second which is our whole north star more broadly I think there is this like larger sense around you know what do you do how do people consume a trillion tokens per day like that's the world we want to create the capability for them to do that a trillion tokens well okay at openi pricing that's at least $5 million at the very least for 5.5 or 5.6 six. Yeah, I think the dollars was probably the most good metric. Yeah. Yeah. So, what's the world in which we consume what currently costs $5 million per person per day? Are you at all worried that just like the average person just can't and won't do that like doesn't do that now with their own brain? Like there actually isn't that much demand for intelligence in the world. You want to enable those people. What about the inverse question? Not what you think is craziest, but like what consensus thing you think is wrong? so the consequence of this is people lose their minds over geopolitics like what what happen if we lost access to DMC for any reason. And um my contrarian take is that it wouldn't be that bad. Supply would take a shock for sure, but the best processes that we have in the west uh like Intel not that far behind at worst like maybe 2x uh worse performance per watt and the gap is just far smaller than than you would make it out to be if you talk if you follow like the chipboard dialogue. What else is happening in the AI world that is not in your path? Meaning it's not like a component of this whole system that you would end up doing something in that interests you most. if you had a 100 entrepreneurs in a room, all of whom wanted to create some new compute startup. um and let's say they were specifically wanted to make hardware chips or systems or racks or whatever. because it seems like we're going to try everything and that will be great for the world. You know, some stuff will work. But if you had to give them advice on how to orient their business to be successful in this coming world, what advice would you give them? It's all about the bottlenecks on supply chain. So, you need to first convince me or convince an investor that you understand the like three to five bottlenecks that dictate modern chip supply. There's TSMC wafer capacity, there's HPM capacity, and there's um like advanced packaging, and maybe a fourth one would be power. Like, where will you get the power? How will you build these racks? Uh and I I want to hear like you should have a great answer to each of those four bottlenecks and how you're going to work around them because it's all arbitrage at the end of the day. You're building a chip because you think that Nvidia has made some choices that are difficult for them to change, which is true. Nvidia makes a lot of choices that are difficult for them to change. They're not perfect. They're just really well balanced. And so, you want to be spiky. You want to pick something and say, I think they've underpriced the impact of how short we're going to be on HBM. We're going to push really hard in this other direction instead. Which, you know, as a as an aside, I do think is probably the thing to attack most. so it's going to be a while until we They they've been burned on that many times. I think they're gonna make everything else more expensive. Think that iPhones will cut their memory. iPhones are going to go up in price and um we're just going to deal with it. Why doesn't Nvidia go all the way to the end and sell tokens? Do you think It's great to have competition amongst his buyers. The kindest thing I mean I my immediate first thought is like all the mentors that I've had over the years. It's a rare person who takes a lot of time out of their their schedule and um and you know makes it like their personal interest essentially to to make sure that you understand something that uh or teach you something or or like ingrain some value in you that they think that you're on the cusp of understanding but just push you over the line for understanding. uh a lot of the people in Nvidia that I mentioned earlier who instilled that like love of performance engineering in me but also my professors in college who I remember like my adviser in like sophomore year I was very impatient student so I show up at his office hours and say like I I want to build AI chips I know what I want to do why am I wasting time taking all these like other basic classes and networking and you know operating systems and he just looked at me and said like you know he laid out basically like the whole stack and showed me the depth of or the beauty of like understanding every piece in the puzzle like he he took my entire path of like trying to focus on one piece of the the system and said that you know it's so rare that someone can actually understand the entire stack from the gate level silicon all the way to building a great internet scale service and you know you should aspire to be someone who over the course of your lifetime achieves that level of understanding. Thank you so much for having me.