Invest Like The Best · JA + tiny EN

元NVIDIAエンジニアが語る、AIが1000倍安くなる理由

Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper

英語はYouTube自動字幕ベース。日本語は学習用の整文対話。一次情報は動画。

パトリックニール
FAQ 一覧 YouTube
パ
パトリック#1

こういう対談の序盤では、いま何を作っていて、今日それが何をしているのかを、そのまま言っておくのが大事だと思う。システムそのものと、なぜ存在するのか、短く案内してくれ。

EN

My job is to make the tokens as cheap as humanly possible. I will achieve that and I will do it through every layer in the stack available to me. I love the supply side levers. I will use every chip. I'll use every source of power and I will use every piece of land in the United States that's, you know, suitable for this. We still treat the agent as a person that is expensive to consult and you should ask them when you have a hard question. That's not the way to think about intelligence. It's incredible that the machine can think and we should [music] try to get that into as many hands as as many people as possible.

ニ
ニール#2

僕の仕事は、トークンを人間に可能な限り安くすることだ。スタックの使える層は全部使って、それをやり遂げる。供給側のレバーが好きだ。あらゆるチップを使う。あらゆる電源を使う。アメリカ国内でこの用途に適した土地も、全部使う。

僕たちはまだエージェントを、相談するのに高い「人」として扱っている。難しい質問のときだけ聞け、という発想だ。知能の捉え方としては間違っている。機械が考えられること自体がすごい。だから、できるだけ多くの人の手に渡すべきだ。

セイル(Sail Research)はトークン工場だ。誰でもリクエストを送れるAPIがある。オープンソースの大規模言語モデルを、好きなタスクに使える。市場で負けない価格で、そのトークンを届ける。その上にエージェントを組むことも支える。サンドボックスと呼んでいるものをホストしている。何時間、何日、何週間も動き続けるエージェント向けに設計した、クラウド上の長時間稼働の仮想マシンだ。

EN

[music] [music]

パ
パトリック#3

つまり、別の種類の推論を出している会社のピアだ、と。特定の種類の推論を担い、その使い方では絶対最安の提供者であり、イネーブラーになるのが目標だ、ということか。

EN

I think it's important early in these conversations to just say the thing like literally what you're building and what it does today. So maybe just orient us there with a with a brief description like literally what the system is that you're building and why it should exist.

ニ
ニール#4

そのとおり。うちの会社のテーマは豊富さだ。知能という新しいコモディティを、ほぼあらゆる産業が持続できるコストで、できるだけ多くの人に届けたい。何かを10倍安くすると、それは新しい製品カテゴリになる。トークンでそれをやりたい。機械が考えられることは、それほど根源的だ。だから世界中で、できるだけ多くの機械を「考える」方向に動かすのが、いまの仕事だ。

EN

Sal research is a token factory. We have an API where anyone can send us requests where they can use large language models, open source large language models for any task they want. Um, we will serve those tokens to them at a price that is unbeatable in the market. We also support their ability to build agents on top of this. We host what we call sandboxes, which are longunning agent virtual machines hosted in the cloud that are designed for agents that run for hours, days, or weeks. And so you should think about you as a peer company to others that serve different kinds of inference. You're serving one specific kind of inference and your goal is to be the absolute cheapest provider and enabler of a certain kind of use of intelligence.

パ
パトリック#5

今日のテーマをトークンコストだとすると、まずはそれで考えるのが正しいのか。別の言い方もあるのか。

EN

Exactly. The theme of our company is abundance. We want to deliver this new commodity of intelligence to as many people as possible at a at a cost that is sustainable for almost every industry. We think that whenever you make something 10 times cheaper, it's a new product category and uh we aspire to do that for tokens. We think it's so profound that the machine can think and now our job is to make as many machines as possible in the world work towards thinking.

ニ
ニール#6

まずは、そのとおりだ。いまのノーススターは、業界でぶっちぎり最安のトークン単価を持つことだ。トークンが仕事や知能の最終単位だとは思っていない。ただ、今日使っている単位なので、話は単純だ。トークンの次は、もう少し「成果」のほうへ移っていく。まだぼんやりした方向だが。

たとえば今日、エージェント経由でトークンを消費するとき、エージェントが何トークン推論するかは、実は自分では制御していない。一定時間考えることもできるし、ツールを何回か呼ぶこともできる。これからは、エージェントがある単位の仕事をやり、ゴールに向けてできるだけ多くシュートを打ち、そこに至るまでに使ったトークン数はタスク次第の従属変数になる、という形が増えると思う。会社がエンジニアの月間トークン予算を決めるのではなく、エージェントがトークン予算を自分で管理する、というイメージだ。

EN

So if you think about uh the theme of the day being token costs is token cost the right way to think about this like is there some other way you'd put it

パ
パトリック#7

なぜ、そこに入り込める隙があるのか。世界中が、より良く、より速く、より安いトークンに向かっているように見える。この問題を、かなり本気で解こうとしている。市場が効率的に取り組めていない、君たちが見つけた独自の隙間は何か。

EN

to start with? Absolutely. Token cost today my north star is I want to have the lowest cost per token in the industry and do that by a mile. I don't think tokens are the final unit of uh of work or intelligence but they are what we use today and so it's very straightforward. I think after tokens you start to move more towards um more outcomes which is like a vague direction. Uh you can imagine for example today when you consume tokens through an agent you don't actually control how many tokens the agent reasons for. It can reason for a certain amount of time or it can call a certain number of tools and increasingly I think we will have agents do some unit of work take as many shots on goal as they can and however many tokens they use to get there is going to be kind of a dependent variable depending on the task. So you think about like agents that selfadminister a token budget as opposed to a company setting a budget for how many tokens engineers can spend per month.

ニ
ニール#8

追い風は二つある。まずオープンソースの台頭だ。これは最初に言わないといけない。顧客も市場全体も、「知能を自分のものにしたい」と考える人が増えている。依存するものに対する主権、コントロールが欲しい。だからカスタムモデルや、ごく普通のオープンソースモデル向けの市場が、かなり厚くなった。誰にも取り上げられない。重みは常に自分のものだ。好きなようにデプロイする権利もある。

二つ目は、その市場の向き方だ。ここ数年、大規模にモデルをサーブする市場はそれなりに育ってきた。課題は、BasetenでもFireworksでもいい、どれを取っても低レイテンシ推論に寄せていることだ。その方向に引っ張った、とても重要な顧客がCursorだ。1年くらい前なら、それは正しい選択だった。半年前くらいから、エージェントに欲しいのは低レイテンシだけではない、と見え始めた。持続性、より長いホライズンのタスクが欲しい。いまはもう明らかだ。エージェント推論の未来は、長時間タスクだ。機械を何時間も、何日も動かす。100トークン毎秒で吐き出す必要はない。10トークン毎秒で十分かもしれない。そのぶん、効率の利点がついてくる。

EN

Why is there an opportunity that you can tackle it? It seems like the entire world is oriented around more better faster cheaper tokens. Right now it seems like the world is trying to solve this problem very aggressively.

パ
パトリック#9

なぜそこまで確信できる。自分としては、何でもできるだけ速く欲しい、と思うが。

EN

What was the unique opening that that you saw that's maybe the market's not being efficient in its attempt to tackle this? So I think there's two things that are tailwinds for our company. One is got to be the rise of open source. I had to talk about that first. I think we are starting to see an increasing number of our customers and the broader market care about owning intelligence. They they want to have control sovereignty over the thing that they depend on. Uh and so that created a much more robust market for customized models or even just like these vanilla open source models that no one can ever take away from you. You always have the weights. You always have the right to deploy them however you like. uh in that world there's been a reasonably robust market for the past couple years uh serving these models at large scale. The challenge is all those companies, you could take your pick, base 10, fireworks together, they all focus on low latency inference, and they were pulled in that direction by one very important customer, uh, cursor. And [clears throat] I think that that was the right choice about a year ago, and as of six months ago, it started to look like maybe low latency wasn't the only thing you wanted from an agent. You wanted more persistence, more long horizon tasks. And now, it's to me very obvious that the future of agentic inference is long horizon tasks. You're going to run the machine for hours or days at a time. It doesn't matter if it spits out tokens at 100 tokens per second. Maybe 10 is just fine. That comes with corresponding advantages and efficiency.

ニ
ニール#10

待っているなら、いちばん速い答えを受け取る権利がある。僕のトリックは、待たせないことだ。先回りして動いてほしい。バックグラウンドで動いてほしい。言い方を変えれば、最良のレイテンシは、レイテンシがないことだ。朝起きたら、仕事はもう夜のうちに終わっている。頼む必要すらなかった。それが夢だ。まだそこまでは来ていない。

もっと重要なのは、プロンプトを投げて応答を待つループに人が入れば入るほど、実は人がボトルネックになることだ。エージェントがどれだけ働くかを、人が制限してしまう。エージェントには、もっと人間の時間スケールで動いてほしい。同僚を5分おきに管理したりはしない。大きなタスクを頼んで、戻って確認するのは毎日か、むしろ週に一度だ。人間とエージェントの協働の未来は、そういう人間の時間尺度だと思う。

EN

Why are you so confident in that? It's it like to me it seems like I want everything as fast as possible.

パ
パトリック#11

それが実際に起きている、だからこの会社を建てるべきだ、という初期の兆候をもう少し聞かせてくれ。

EN

When you're waiting on it, you absolutely deserve the fastest answer possible. My trick is I don't want you to be waiting on it. I want it to be proactive. I want it to be in the background. One way to say it is like the best latency is no latency at all. When you wake up in the morning, the work's already been done overnight. you didn't even have to ask for it. Uh that's the dream. We're not quite there yet. But more importantly, I think the more you're in the loop as you prompt agents and wait for a response, in fact, you're the bottleneck uh in helping the in having the agent do more or less work. What we'd like is the agent to operate on more human time scales. You don't manage your colleagues every 5 minutes. You ask them to do a high level task and you come back and check in maybe every day, but more likely once a week. And that to me is the future of human agent collaboration, more like human time skills. say more about the early indications that this is happening and therefore you should be building this company.

ニ
ニール#12

いちばん大事なのは、テスト時計算スケーリングだ。エージェントに時間を与えるほど、答えが良くなる、という考え。理論としては約2年前からあった。ただ、実際に賭けられるようになったのは、昨年末のOpus 4.5あたりだ。Opus 4.5は、長いホライズンのタスクにまともに向いた最初のエージェントだった。出た当初はかなり普通だった。でも最近のモデルと、オープンソース側で僕たちが進めたものを見ると、エージェントは1時間単位で走れる。何日とは言わない。でも1時間は、今日もう十分に適している。平均のターンやタスク長が伸びていくのを見れば、点はそんなにいらない。指数関数を引いて、もっと長く走らせる価値があると分かる。

EN

Well, so the first and most important thing is the idea of test time compute scaling. Uh the idea that you can give an agent more time and it will give you a better answer. So that was theorized about 2 years ago now and uh but it wasn't really something that we could actually bet on until I would say late last year with Opus 45. Opus45 was the first agent that was at all suitable for longer horizon tasks and you know it was pretty mediocre and when it first came out but you look at the more recent models and what we've done on open source as well and you see that agents are capable of running for an hour at a time. I wouldn't say it's days but definitely an hour is quite suitable today. And so just seeing that like average turn or task length get longer and longer uh it doesn't take many points to have you kind of draw out the exponential and see that agents are worth running for longer periods. What what do you think will be the market share of longrunning agents in 3 years or something like this?

パ
パトリック#13

3年後くらい、長時間エージェントの市場シェアはどうなると思う。

EN

You know, I love this market because it's unbounded. There's no human in the loop. So, you can consume as many tokens as you like in the background. Uh versus human attention span. If you tell me to consume 10x as many tokens at codeex or at cloud code, I'm actually not sure if I can anymore. I'm already in a loop and locked in coding for most of the the day that I'm at the laptop. What is undowned is how many tokens can be consumed in the background or proactively. So longterm, I think, you know, we're going to end this year at maybe 50/50 background and uh and real-time workloads, but I see this going to 9010 in favor of background.

ニ
ニール#14

この市場が好きなのは、上限がないからだ。人間がループに入らない。バックグラウンドでは、好きなだけトークンを消費できる。対して人間の注意には限界がある。CodexやClaude Codeでトークンを10倍使え、と言われても、もう無理かもしれない。ノートPCの前にいるあいだは、すでにループに入ってコーディングにロックされている。上限がないのは、バックグラウンドや先回りで消費できるトークン量だ。長期では、今年の終わりはバックグラウンドとリアルタイムがだいたい50対50だと思う。ただ、これは90対10でバックグラウンド側に行くと見ている。

EN

What are the sorts of things like what are your favorite examples of something that gets accomplished much better as a background task than as a human in the lip task?

パ
パトリック#15

バックグラウンドのほうが、人間がループに入るタスクよりずっとうまくいく例で、好きなものは何か。

EN

Most deep research, most questions where you want to have a definitive answer over not 100 sources, not a thousand sources, but 10,000 sources or more. If you want to build an authoritative index of information like for example one of our customers parallel web systems seeks to do. They want to build an index over the whole internet and they want to monitor the internet in real time for changes. That is the kind of crazy exabyte scale task that you need a very different kind of intelligence or scale of intelligence to achieve. Deep research is a top category for us and then increasingly we see cyber security following this direction. If you think about, yes, there's so much code you can generate, but there's uh exponentially more ways to break that same code than it is to generate that code. And there are some great customers out there who are working very hard to uh find agents that can break any piece of software and proactively patch them. So when Fable first came out, for example, or Mythos first came out, basically there was this push in the cyber security community to run Fable against every line of code we've ever written and look for bugs in 20 different ways. uh meaning you're looking for both memory errors, you're looking for business logic errors and looking for like network vulnerabilities, all these things. And these are all actually things that you would write specialized agents for. You wouldn't just have Fable look at the source code once, you'd have it actually set up environments where you can pen pentest these applications. And at some point, people started to make this joke that security has become proof of work. When you want secure software, it's really a question of how many dollars did you spend on anthropics APIs trying to break into your software. uh that is the best indication for how secure it is because that's the best tool in the world. And increasingly we found that the frontier of intelligence here is quite jagged. It's not the case that Fable finds a supererset of all bugs in software. You would find some bugs with a very small model that you don't find with the large model. Uh you'd find some bugs with Haiku that you would find with Fable and vice versa. So it encouraged this very diverse approach to sampling and trying to build cyber security agents that break software autonomously such that you can patch them. If you were to get sort of like speculative and imaginative about the sorts of things that longunning very cheap very longunning agents can enable. We talked about some very practical examples deep research um cyber security etc. But if you if you get a little bit dreamier about the use cases new product category that this sort of inference will unlock and I guess the question is just like so what like what if you're maximally successful dream a little bit about what that might enable.

ニ
ニール#16

ほとんどディープリサーチだ。100ソースでも1000ソースでもなく、1万ソース以上を踏まえて確定的な答えが欲しい問い。権威ある情報インデックスを作りたい場合もそうだ。顧客のParallel Web Systemsがまさにそれを狙っている。インターネット全体のインデックスを作り、変化をリアルタイムで監視したい。エクサバイト級の、常識外れなスケールの仕事で、まったく違う種類、あるいはスケールの知能が要る。

ディープリサーチはうちのトップカテゴリだ。その次に、サイバーセキュリティが同じ方向に来ている。生成できるコードは多い。でも同じコードを壊す方法は、生成するより指数関数的に多い。どんなソフトウェアでも壊せて、先回りしてパッチを当てるエージェントを、必死に探している顧客がいる。

たとえばFableが出たとき、あるいはMythosが出たとき、サイバーセキュリティのコミュニティでは、これまで書いた全行にFableを走らせ、20通りのやり方でバグを探す、という動きが起きた。メモリエラー、ビジネスロジックの誤り、ネットワークの脆弱性。どれも、専用エージェントを書く対象だ。Fableにソースを一度見せるだけでは足りない。実際にペンテストできる環境を立てさせる。ある時点から、セキュリティはプルーフ・オブ・ワークになった、という冗談が出始めた。安全なソフトウェアが欲しいなら、要するにAnthropicのAPIに何ドル使って自分のソフトを破ろうとしたか、だ。世界最良のツールだから、それが安全性のいちばんいい指標になる。

しかも知能の最前線はかなりギザギザだ。Fableがソフトウェアの全バグの上位集合を見つける、わけではない。小さいモデルで見つかって、大きいモデルでは見つからないバグもある。Haikuで見つかってFableでは見つからないものも、逆もある。だから多様なサンプリングをして、ソフトウェアを自律的に壊してパッチできるセキュリティエージェントを組む方向に、押し出された。

EN

Yeah absolutely. So I think for individual users what I'm excited about most is this idea of proactive intelligent agents. Um you can imagine Siri that is running in the background all the time to understand what's all the emails you received in a day, all the text messages you receive in a day and and has a much more encyclopedic view of your life and how to be helpful in that life. Right now there's still point solutions and so you have to you end up doing a lot of prompting. Siri is not very proactive. That's something we can fix with abundant abundant inference. If you trust uh the machine enough that it's reliable and also trustworthy as in private um you might even imagine the machine can understand how you interact with it and proactively surface your next action whenever you open your phone. Can we build a good model of what you're going to do next?

パ
パトリック#17

実用的な例はディープリサーチやサイバーセキュリティだった。もう少し夢を見て、すごく安くてすごく長く走るエージェントが解禁する新しい製品カテゴリは何か。最大限うまくいったら、何が可能になる。「だから何だ」の部分を、少し想像してくれ。

EN

My estimation is yes, we totally can.

ニ
ニール#18

個人ユーザーでいちばん興奮しているのは、先回りの知能エージェントだ。Siriが常にバックグラウンドで動き、その日のメールもテキストも全部把握し、自分の人生の百科事典的な見取り図を持って、どう助けるかを分かっている、というイメージ。いまはまだ点のソリューションで、プロンプトをたくさん打つことになる。Siriは先回りしない。豊富な推論があれば、そこは直せる。機械が信頼できて、しかもプライベートだと信じられるなら、自分とのやり取りを理解して、スマホを開いた瞬間に次のアクションを先回りして出すことまで想像できる。

EN

And the key to that is incredibly cheap intelligence.

パ
パトリック#19

次に何をするかの、良いモデルは作れるのか。

EN

You have to be willing to spend tokens without any promise of return. That's the unlock. The long lens view to take on this is that we have we have a form of intelligence that can tackle any verifiable problem. Any verifiable problem means most software. It means a lot of formal like math proofs and similar. And it could also mean scientific discovery. These are all relatively verifiable problems. And all those things currently have a dollar cost attached to them essentially. That's a hidden one. It's like how many tokens could you possibly harness to make this work? And we have actually started to bring it within view a dollar cost for these long horizon tasks that is reasonable. It's not millions, it's thousands and maybe it could be hundreds or even tens of dollars in the near future to have a definitive answer to any scientific question to any research problem. RAMP is the only platform built to make your finance team leaner, faster, and better, saving businesses 5% annually on average, so you can stay focused on growth. RAM customers grow revenue 3.2 two times faster than the average American business. Visa, Verscell, Kerser, Stripe, Notion, 11 Lab, Shopify, and 70,000 other businesses all now run on RAMP. Mine does too, and so should yours. Learn more at ramp.com/invest. OpenAI, Cursor, Anthropic, Perplexity, and Verscell all have something in common. They all use work OS. To achieve enterprise adoption at scale, you have to deliver on core capabilities like SSO, skim, arbback, and audit logs. Instead of spending months building these missionritical capabilities yourself, you can just use Work OS APIs to gain all of them on day zero. That's why so many of the top AI teams you hear about already run on Work OS. Work OS is the fastest way to become enterprise ready and stay focused on what matters most, your product. Visit works.com to get started. Felix by Rogo is a personal finance agent that turns a single prompt into [music] finished client ready work using your firm's own templates, context, and standards. Send Felix an email like, "Take [music] these comments and turn them for me." Or, "Udate my tracker with the context of these emails." And Felix sends back finished [music] PowerPoint decks, Excel models, and sourced research. Felix works the way your team already does, delivering [music] work quickly and accurately around the clock. Learn more at robo.ai/felix. [music]

ニ
ニール#20

作れる、と見ている。鍵は、信じられないほど安い知能だ。リターンの約束なしにトークンを使う覚悟がいる。それがアンロックだ。長い目で見ると、検証可能な問題なら何でも解ける形の知能がある。検証可能とは、ソフトウェアのほとんど、形式的な数学の証明、科学的発見もそこに入る。どれも比較的検証可能な問題で、いまは本質的にドルのコストが付いている。隠れたコストだ。この仕事を成立させるために、どれだけのトークンを集められるか。長時間タスクのドルコストは、ようやく視野に入ってきた。数百万ドルではない。数千ドルで、近い将来は数百ドル、あるいは数十ドルで、どんな科学的な問いにも、どんな研究課題にも確定的な答えが出る、という水準まで行けるかもしれない。

EN

And so, if we dream about that future, we're we then become limited just by the questions that people can ask. Basically,

パ
パトリック#21

その未来を夢見るなら、制約は人が問いを立てられるかどうかだけになる、ということか。

EN

pretty much the questions we can ask. uh the models are on the cusp of basically taking even a high level question and chasing it down every possible follow-up you can have the model essentially take that on its own

ニ
ニール#22

ほぼ、立てられる問いだ。モデルはいま、大まかな問いを受け取って、可能なフォローアップを全部追い切る瀬戸際にいる。モデルがそれを自分で引き受ける。

EN

and the question is what is your token

パ
パトリック#23

問題は、トークン予算がいくらか、だ。

EN

and we will solve the token budget problem

ニ
ニール#24

トークン予算の問題は、僕たちが解く。

EN

what about non-verifiable tasks

パ
パトリック#25

検証できないタスクはどうする。

EN

I put basically the entire category of human taste into that category we have not solved human taste yet and I don't know that it fundamentally can be I'm excited to be surprised here but um we are focused on very quantitative uh problems we leave the quality of writing we the uh the beauty of art to to people.

ニ
ニール#26

人間の趣味・審美、そのカテゴリ全体をそこに入れる。人間の趣味はまだ解けていない。根本的に解けるのかも分からない。驚かされるのは楽しみだが、僕たちは定量的な問題に集中している。文章の質や、芸術の美しさは人に任せる。

EN

All right. Now, let's talk about the uh the very clever stack of solutions that you hope to build

パ
パトリック#27

では、巨大なトークン工場、極端に安い知能の供給源を作るために、君たちが組もうとしている巧妙なソリューションのスタックの話をしよう。他社とは違う切り口で、この課題へのマスタープランを説明してくれ。

EN

ultimately to have this giant token factory, extremely lowcost intelligence supplier of extremely lowcost intelligence.

ニ
ニール#28

ソフトウェア、ハードウェア、電力の三層で考える。いつもソフトウェアから始める。今日のチップ、今日のデータセンターで、効率を上げる余地はどこか。最初にやったのは、LLMのソフトウェアスタック全体を、GPUのピーク効率のまわりに組むことだ。NVIDIAのGPUを使っている。同じチップから、世界の誰よりも多くトークンを絞り出したかった。それはいちばん低い層、プログラミングのカーネルから始まる。それが僕の出自だ。職業人生はずっとGPUとカーネルだ。大学在学中の最初の仕事がNVIDIAで、テンソルコアがチップに載る権利をどう勝ち取ったかを、その場で見た。2016年のことだ。

EN

I think you think about this in terms of level software, hardware, uh and power.

パ
パトリック#29

素人向けに、それは何を意味するのか説明してくれ。

EN

Talk through what your master plan is to approach this challenge that's so different from what others are thinking about doing.

ニ
ニール#30

テンソルコアは、GPU上の専用ユニットで、行列積を加速する。

EN

You know, we always have to start with software. you know where is the opportunity on today's chips with today's data centers to improve efficiency and the first thing we did was we tried to build the entire LM software stack around peak GPU efficiency meaning we're using Nvidia GPUs we wanted to squeeze out more tokens from the same chip than anyone else in the world and that starts with the lowest level of programming kernels it's actually my background I spent my whole life actually uh my whole professional life working on GPUs and kernels in Nvidia was my first job while I was in college and uh I got to see how the tensor cores got to earn their right to be on the chip. This is back in 2016.

パ
パトリック#31

それだけか。テンソルコアの進化には長い歴史がある。それはあとで入る。なぜ行列積がそんなに重要なのか。

EN

Just describe what that means for for the lay person.

ニ
ニール#32

いい問いだ。行列積が計算の原子単位に見える理由を、神の真理として逆から説明できるわけではない。聞いた説明でいちばん近いのは、二つの数の塊を混ぜて、面白い相互作用を起こす、とても簡潔な方法だ、ということだ。言えるのはそこまでだ。線形代数が、データの任意の関係をとてもコンパクトに表せることになるのは、実に都合がいい。

NVIDIAは偉大なグラフィックス企業で、GPUとゲーム向けグラフィックスで長いあいだシェアを握ってきた。2010年代半ばから、追跡していた機械学習の仕事にグラフィックスプロセッサを向ける、一種のスカンクワークスが始まった。NVIDIAにいたとき、上司たちのラボノートを読んだことがある。当時はまだ小さかったICMLやNeurIPSみたいなML会議に行って、論文をメモしている。「ディープラーニングが流行り始めている。面白いのは、院生がゲーム用のNVIDIA GPUで大きなモデルを訓練していることだ。ここを深掘りすべきだ」。2015、2016年までには、少なくともジェンセンに確信があった。この用途は伸びるだけだ。貴重なシリコンのダイ面積を、この芽にどんどん割り当てよう。テンソルコアの最初の版をチップに載せよう。

要するに、画面にピクセルを描くためのゲーム用チップを、行列積ができるように改造する話だ。まだ早かった。シリコン面積を増やしてくれと頼むと、実質グラフィックスチームと争うことになる。チップ企業ではどこでも、その争いはある。設計者はそこをとても慎重に守る。間違った技術に投資したくない。他の機能に回せたはずの機会費用だからだ。歯を食いしばって戦い、第一世代ではダイ面積のごく一部、たぶん5〜10%くらいを取った。当時のコンピュータビジョンの基本演算だった畳み込みを、少し加速するためだ。

そのうえで、チップから性能を絞り切るソフトウェアチームがいた。僕がいたのもそこだ。そこでいちばん学んだのは、NVIDIAの気風だ。社内にはスピード・オブ・ライトという言葉がある。作ったハードウェアはどれも、理論上限まで追う。エンジニアの頭に染みついている。機械ができるなら、自分たちが思うところまで、機械を最前線に押し込む。

EN

So, okay, tensor core is a specialized unit on the GPU that accelerates matrix multiplication.

パ
パトリック#33

スピード・オブ・ライトとは、可能なことの端、ということか。

EN

Simple as that. There's been a long history of how we evolved at tensor core over time that we'll get into.

ニ
ニール#34

可能なことの端だ。そのとおり。この周波数で動いて、サイクルあたりこの回数の乗算ができるはずだ、と思えば、そこに行く。ボトルネックを全部壊して、ピーク性能まで持っていく。だから今でもエンジニアには、100%のスピード・オブ・ライトを追え、と言っている。競合との相対値はどうでもいい。気にするのは絶対値だけだ。チップの上で何ができるか、どうやってそこに到達するか。

EN

And why is matrix multiplication so important?

パ
パトリック#35

NVIDIAにいた章を離れる前に、その文化以外で、考え方を変えたもの、当時の事業の回し方や社風でいちばん印象に残ったものはあるか。

EN

That's a great question. I actually I cannot say that there is a divine truth of the inverse that explains why matrix multiplies seem to be the atomic unit of computation. But uh one way I've heard it described to me is well it's a really succinct way to mix two blocks of numbers together and have them interact in some interesting way. That's as much as I can say about it. It is really convenient that linear algebra turns out to be a very compact representation of arbitrary relationships in data. So Nvidia great graphics company obviously has had market share dominance in GPUs and and gaming graphics for quite some time. And then starting in like the mid2010s they started to actually start these like skunk works projects to make the graphics processor more suitable for machine learning tasks that they were tracking. I remember actually reading some of the like lab notebooks of some of my managers when I was at Nvidia. they would visit these small ML conferences like ICML or NURPS at the time and they would just take note of these papers like oh this deep learning thing seems to be catching on and what's really interesting is that these grad students are using gaming Nvidia GPUs in order to train their large models we should double click on this and figure out what's going on here and by 2015 2016 at least Jensen had the conviction to to kind of double down on hey this usage of our models is only of our chips is only going to grow let's start allocating more and more precious silicon die area to this capability that seems to be emerging. Let's put the first version of tensor cores on the chip. So, we're talking about, you know, taking this gaming chip which is designed for painting pixels on a screen and adapting it to do metric multiplies and it was early and you would be competing against the graphics teams essentially when you ask for more silicon area and any chip company. There's always competition for that. It is something that that the designers guard so carefully. you don't ever want to invest in the wrong technology because that's opportunity cost that you could have allocated to some other functionality. And so we we kind of like fought and tooth and nail and got just a tiny bit of diary maybe like 5 10% something like that for the first generation of these chips uh to get some some amount of acceleration for basic convolutions which were the fundamental operation for computer vision models in the day. Uh, and then we had a software team that was trying to squeeze all the performance we could out of the chip. And I think on that software team, which is where I work, that's what actually taught me the most about um, just the ethos that Nvidia has around they have this term called speed of light. They always chase the speed of light for any piece of hardware that they make. It is so ingrained in every engineer's mind that if the machine can do it, we're going to push the machine to the frontier until it does what we think is.

ニ
ニール#36

NVIDIAの話ならいくらでもある。いくつか話せる。好きなものの一つは、在籍の長さだ。2015、2016年に一緒に働いた人の多くが、今もそこにいる。定着率がすごい。少なくともシリコン側では、キャリアで一緒に働いたなかで最良のエンジニアたちだ。動機も情熱も極端に強い。並列計算という概念を、姿を変えるたびに信じてきた。チップが進化するのを見るのが好きだ。人生の仕事で、その方向では極端に有能だ。

同時に、かなり質素な会社でもある。2008年以降、シリコンバレーの会社は福利厚生を削ったところが多い。無料のランチはない、といった話だ。NVIDIAは一段先まで行った。冷蔵庫に無料の牛乳がなかった。コーヒーに牛乳が欲しければ、毎月1ドルをミルククラブに出して、そのクラブがCostcoの牛乳を冷蔵庫に入れる。はっきり覚えている。セイルでは、そこまではやっていない。

EN

And the speed of light is the edge of what's possible.

ニ
ニール#37

その倹約は会社全体に染みついている。NVIDIAでのあの時期を経ると、ソフトウェアで下のハードウェアをもっと効率よく使うとはどういうことかを、体で覚える。

EN

Before we leave that chapter of your time at NVIDIA, anything else beyond that cultural touch point that really like changed the way you think about things or that stood out the most about how the business ran back then or its culture? I have a ton of stories about Nvidia. We can I can tell you a few of them. Um, one of my favorites is that on the tenure side, a lot of people I worked with in Nvidia in 2015, 2016 are still there today. That company has incredible retention and these are the best engineers uh, frankly on the silicon side at least I've worked with in my whole career. They're extremely extremely motivated and passionate. They've believed in parallel computing as a concept through its various incarnations and have loved seeing the chip evolve. This is their life's work and they're extremely extremely competent in that direction. They're also a very frugal company. Nvidia and all, I guess all the Silicon Valley companies after 2008, they had some cutbacks and like perks. So, no free lunch. Uh, for example, Nvidia took it one step further. There was no free milk in the fridge. So, if you wanted to drink coffee at Nvidia and you wanted some milk, you actually had to chip in a dollar every month to the milk club and the milk club would stock Costco milk in the fridge. And I remember that distinctly. We don't do that at sale, but uh

パ
パトリック#38

そういう経験を得た、と。

EN

it's a it's a frugality that permeates the company. And so coming out of this time there, you get this experience of what it's like to develop more efficient usage of the underlying hardware through software.

ニ
ニール#39

そう。

EN

Yes.

パ
パトリック#40

それを今の環境に結びつけてほしい。

EN

And so so link that to, you know, today's environment.

ニ
ニール#41

GPUは本質的にスループット機械だ。大量の仕事を渡して、演算ユニットをピークまで使い切らせて噛み砕かせるとき、GPUはいちばん機嫌がいい。だがここ数年のAIの使い方は、そうなっていない。いちばん多い形は対話型のチャットボットで、キーボードの前の人に答えをできるだけ速く吐くことが大事になる。ユーザーを待たせるな、できるだけ速くしてほしい、という話と同じだ。

これがGPUにとってはかなり厄介だ。トークンを速く吐き出そうとすると、演算を満載して本領を出させるのが難しい。GPUにはスループット向きかレイテンシ最適化かという、根本的なトレードオフがある。使い方の形がチャットボットだったから、みんなレイテンシ最適化を選んできた。

来年いちばん大きい変化はここだと思う。チャットボットから、先回りするエージェント、バックグラウンドのエージェントへ移る。その世界では、スループット中心のスタックを組む方がよほど筋がいい。

EN

Yeah, absolutely. So, so I think um the GPU is fundamentally a throughput machine. The GPU is happiest when you give it a lot of work to do and let it chew through that work at peak utilization of its compute units. But that's actually not the way that we've taken AI in the last couple years. We've really pushed AI to be an interactive chatbot tool is the most common form of AI usage today. And in that world, you care a lot about actually spitting answers out to the to the person at the keyboard as quickly as possible. To your point about don't make the user wait, I want things as fast as possible. And so that's actually quite interesting for the GPU. It's very difficult to put the GPU in its happy path of being fully compute utilized when you're trying to spit out tokens quickly. There's a fundamental trade-off on the GPU between being uh throughput oriented or latency optimized and everyone has chosen latency optimization because the shape of usage was chatbot oriented. I believe that's the most profound change we're going to see in the next year. We're going to move away from chatbots to more proactive or background agents. And in that world, it makes a lot more sense to build a stack around throughput.

パ
パトリック#42

スループットとレイテンシのトレードオフが壊れない理由を、技術的に説明してくれるか。同じハードで両方はなぜ無理なのか。

EN

Can you explain technically why the trade-off between throughput and latency is unbreakable? Why can't we have both from the same hardware? It's quite foundational in almost every system that you could ever possibly look at. There's always a trade-off between getting a small amount of data through the system as quickly as possible and leaving a lot of buffer uh room for that or trying to run wide and slow like narrow and fast or wide and slow is like a classic trade-off in all computer science. But for GPU specifically, I think there's one thing to focus on which is there's this concept of like batching on the GPU. We want to group many users work together into a batch that we can uh run all at once on the GPU. That's the parallel processing of the GPU. We'd like to have a lot of parallel work to do. The thing is though, you're doing net more work when you run a large batch of compute together. And so you might be filling all the units, but every step along the way as you carry a a batch of work through the GPU, there's more work to be done. And so any individual token or any individual user's request in that batch, it's going to spend a longer time on the GPU being carried with other people's traffic. Maybe the way to say it is um you know, if you want to get downtown and SF, you can take the bus or you can take a private transit. And the private transit is going to have its own direct path as the crow flies or you know, using exactly the roads that you want from point A to point B. a bus, it's going to have to serve many more people and it it has to fundamentally uh do something that works for everyone

ニ
ニール#43

ほとんどあらゆるシステムの根っこにある話だ。少量のデータをできるだけ速く通して、そのためにバッファの余裕を大きく残すのか。細く速く走るのか、広く遅く走るのか。計算機科学の古典的なトレードオフだ。

GPUに即して言うなら、焦点はバッチ処理だ。多数のユーザーの仕事を一つのバッチにまとめ、GPUで一度に走らせたい。それがGPUの並列処理で、並列にできる仕事が多いほどいい。ただし大きなバッチをまとめて走らせると、正味の仕事量は増える。ユニットは全部埋まるかもしれないが、バッチをGPUの中で運ぶ一歩ごとに、やることは増える。だからそのバッチに乗った個々のトークン、個々のユーザーのリクエストは、他人のトラフィックと一緒に運ばれる分、GPU上でより長く過ごす。

たとえ話にするなら、サンフランシスコのダウンタウンに行きたいとする。バスに乗るか、自家用車に乗るかだ。自家用車は、直線に近い自分の経路で、使いたい道だけを通ってA地点からB地点まで行ける。バスはもっと多くの人を乗せる必要があり、全員に通じる動きをせざるを得ない。だから経路は遅くなり、他の人の乗り降りで止まる。バス対自家用車の比喩は、かなり正確だと思う。

EN

and so it takes a slower path and it stops and and waits for other people to get on and off. I think the bus versus car analogy is pretty accurate

パ
パトリック#44

いい比喩だ。だから最適化のステップ1は、NVIDIA GPUの上に可能な限り最高のバスを作ることだ。

EN

and it's a great analogy and so step one for what you're trying to do is like create the best possible bus on top of Nvidia GPUs. Like that's step one of your optimization.

ニ
ニール#45

その通りだ。だから違う並列化の方式を試す。別の例を出すと、NVIDIA GPUで彼らが本当にうまく革新したのが、GPU同士をつなぐNVLinkだ。このNVLinkが非常によくできているので、大きな行列積をもっと速くやりたいとき、半分に切って2枚以上、たとえば最大8枚のNVIDIA GPUにシャードできる。それぞれがその大きな行列積の断片を担当し、最後に結果をリダクションして一つに戻す。演算の最小レイテンシを下げる、非常に良い方法だ。

各GPUの仕事量はたとえば8分の1になるから、より早く終わる。ただし8倍速くはならない。線形には伸びない。サブリニアだ。ハードウェアは8倍使うのに、速度は8倍にならない。4〜5倍くらいだ。ストロングスケーリングは出ない。通信オーバーヘッドがあるからだ。小さいタイルの仕事をさせると、大きいタイルよりGPUの効率が少し落ちる。

最小レイテンシが欲しいなら、速くする手段はこれしかない。できる。だが僕が選ぶ選択ではない。僕ならエキスパート並列やパイプライン並列のような、別の並列化を使う。通信レイテンシをオーバーラップさせて隠す工夫もできる。低レイテンシのサーバでは、その余地が小さい。

EN

That's exactly right. It means we explore things like different parallelism schemes. Maybe that's another example I can give you is um with Nvidia GPUs, one of the things that they've really innovated on and done a great job with is the NVLink uh interconnect between GPUs. And in fact, that NVLink system is so good that you can if you have a large matrix multiply that you want to perform faster. You can actually cut that matrix multiply in half and shard it [clears throat] across two or more up to eight, let's say, Nvidia GPUs and have them all work on pieces of that larger matrix multiply

パ
パトリック#46

NVLinkはレイテンシ性能を上げる技術、という理解でいいか。

EN

and have them connect their results together at the end. Reduce their results back together at the end. And this is a great great way to cut the minimum latency of of an operation. You're each GPU is now doing 1/8 as much work, let's say, uh, and therefore it can finish faster but not eight times faster. It's sublinear scaling. You'll use eight times more hardware, but you won't get eight times the speed. You might get like four to fivex the speed. You're not going to get strong scaling. And this is because of communication overhead. It's because every GPU is going to be a little bit less efficient working on a smaller tile of work than a larger tile of work. And so, it's the only way to speed up if you want the minimum latency possible. You can do that, but it is not the choice I would make. For example, I would prefer to use a different parallelism scheme like expert parallelism or pipeline parallelism. And we may do interesting things to overlap and hide the communication latency in a way that you would have less ability to do that for a low latency server.

ニ
ニール#47

そう。

EN

So is the right way to think about NVLink as a technology which improves latency performance? Yes.

パ
パトリック#48

レイテンシ性能だけ、か。

EN

And only latency performance

ニ
ニール#49

そう捉えていい。それが次に話す、うちが会社として違うやり方をする話につながる。低レイテンシ推論では、NVLinkは必須だと言っていい。

NVIDIAは低レイテンシ推論が非常に強い。ただ、うちは低レイテンシ推論をあまり気にしていない。ではどうなるか。他社がNVLinkをすぐものにするとは、期待して待ってはいない。難しい技術で、スケールさせるのも、本番に載せるのも難しい。

だから他社のチップがあって、基礎的な演算部品が良ければ、行列積はうまくできる。できないのは、仲間同士でその結果を速くやり取りすることだ。ならそのチップは、ドルあたりの演算が非常に良い選択肢として、スタックに入る余地がある。僕がほとんどの場合に最適化しているのは、このチップに何FLOPSあるか、所有と運用で1時間あたりいくらかかるかだ。ドルあたりFLOPSでは、NVIDIAより確実に上に来るチップがある。ただしインターコネクトは弱い。だから僕の仕事は、そのチップを推論に耐えるものにする並列化を見つけることだ。テンソル並列にはならない。それにはNVIDIAがほぼ必須だ。だが他の手法なら、うまくいく。

EN

which will segue into the next segment of what we you know do differently as a company. But yes, NVLink is mandatory I would say for low latency inference.

パ
パトリック#50

レイテンシの話を離れる前に、セレブラスなど、途方もなく速い演算ができる会社についてコメントしてほしい。あのアプローチ、あの会社をどう見るか。超低レイテンシに振ったハードウェアの未来予測は。

EN

So Nvidia is excellent at low latency inference. And I'm telling you that we don't really care that much about low latency inference. So where does that leave us? Well, I think I'm not holding my breath for other companies broadly to figure out NVLink quickly. It's a challenging technology to figure out. It's hard to scale. It's hard to productionize. And so, if I do have some other vendors chip and it is good at the foundational compute components, it can still do metric multiplies really well. It just can't communicate those results across its peers quickly. Well, maybe there's a room for that other chip in my stack as a really really good compute per dollar option. And that's what I actually optimize for in most cases is how many flops does this chip have and how much is it going to cost me per hour to operate to own and operate. Uh and so there are other chips that definitely rank higher than Nvidia on flops per dollar, but they may not have as much interconnect. And so it's my job to figure out what parallelism scheme am I going to use that's going to make this chip suitable for inference. It's not going to be tensor parallelism. Nvidia is basically mandatory for that. But other techniques may work well for me. So before we leave the latency part of the story, can you comment on companies like Cerebras or others that can perform incredibly fast operations? I'm curious like what you think about those approaches, those companies, what might happen in the future. What is your prediction for the future of very low latency focused hardware? Cerebrus Grock uh and a couple others that are coming out of stealth now I think have made a very interesting bet on not just building another GPU but actually building a different kind of accelerator that focuses on a different memory hierarchy. Uh they want to maximize the amount of SRAMM on the chip and use that as very very fast memory for for weights and KV cache. So SRAM versus DRAM there's two ways to make memory for a chip. One is to integrate the memory on the logic die itself. Like meaning you tell TSMC, I want this many megabytes of of storage on my chip. Uh and there's a way to build that. TSMC has a standard cell library you can use and you can just print out a bunch of cells of SRAM. The problem with SRAMM is it takes a lot of area on the silicon die. Um, so if you want to build a large die like let's say the Nvidia Blackwell at 800 mm square. If you made that whole DS RAM, it would be in the maybe like singledigit gigabytes, it's not a crazy amount of of data storage. Compare that to if you're willing to take a different process technology entirely. So not TSMC anymore, but now Micron SKH Highix Samsung. They build DRAM, which is a whole different way to build memory that's more focused on capacitors than transistor cells. So, SRAM, the standard way to build SRAMM is what's called the 6T transistor cell. It's a stable transistor arrangement that allows you to write a bit to it and then it holds that state in that bit regardless of whether you keep applying. Well, you had to apply some power, but uh it it's holding that bit without any sort of like active management. It's static. Now, dynamic RAM, DRAM, it's dynamic because what you do to write some data is you write a charge onto a capacitor and as soon as you write that charge into that capacitor, the charge is dissipating. it's been leaking. And so the dynamic part of DRAM is that you must every 50 milliseconds or so refresh every bit you've written. So you're constantly juggling billions of balls in the air essentially billions of bits have to be managed by a memory controller which is reading and refreshing every bit on the DRM. Now the benefit of that is you can get much much higher density and it's a whole different process technology. There's a ton of different trade-offs. Hence why we split the DM manufacturing into an entirely different company like Micron SKX and Samsung. These are the best companies in the world to do this. They build DM. And if you take DM from those companies and you stack it uh into many layers and you kind of print them or or solder them around the main logic die that you get from Nvidia, you can now get hundreds of gigabytes uh like Blackwell has 288 GB of HPM capacity around the logic die. And the logic die itself maybe only has like 500 megabytes of of SRAM. So it's possibly multiple orders of magnitude, three orders of magnitude difference in density for DRAM versus SRAMM. Okay, so let's go back to Cerebras. What are they doing? Well, they see this problem, there's not really an obvious way to increase SRAM density on the chip. But thing with SRAMM is because it's so physically close to the logic gates that actually do the computation, the arithmetic logic units are right next to the SRAMM that they're going to pull from, the compute units that are doing the matrix multiplies can pull data from SRAMM at just mind-boggling speeds. You know, Serbis quits pabytes per second, 21 pabytes per second further away for scale engine 3. And so compare that to HBM on an Nvidia black wall is u you know 10 terabytes per second or so in that range. So once again, many orders of magnitude difference, more capacity, but proportionally less bandwidth essentially.

ニ
ニール#51

セレブラス、Groq、今ステルスから出てきた他の数社は、単なる別のGPUを作るのではなく、メモリ階層の違うアクセラレータを作る、という面白い賭けをしている。チップ上のSRAMを最大化して、重みとKVキャッシュ用の超高速メモリにしたい。

SRAM対DRAMだ。チップ用のメモリの作り方は二つある。一つはメモリをロジックダイそのものに載せる。TSMCに、このチップに何メガバイトの記憶が欲しいと言う。TSMCのスタンダードセルライブラリでSRAMセルを並べて刷れる。

問題は、SRAMはシリコンダイの面積を食うことだ。NVIDIA Blackwellくらいの大きなダイ、800平方ミリメートルを全部SRAMにしたら、たぶん数GB程度だ。大した容量ではない。

一方、まったく別のプロセスに行くなら、もうTSMCではなくMicron、SK hynix、Samsungだ。彼らはDRAMを作る。トランジスタセルよりコンデンサに寄った、別のメモリの作り方だ。SRAMの標準は6Tトランジスタセルだ。安定したトランジスタ配置で、ビットを書いて、積極的に管理し続けなくてもその状態を保つ。電源は要るが、静的だ。

DRAMはダイナミックだ。データを書くとはコンデンサに電荷を乗せることで、書いた瞬間から電荷は漏れ始める。だからダイナミックなのだ。だいたい50ミリ秒ごとに、書いた全ビットをリフレッシュしなければならない。常に数十億個のボールを空中でジャグリングしているようなもので、数十億ビットをメモリコントローラが読んで書き直す。

利点は密度が桁違いに出せることだ。プロセスも別物で、トレードオフも多い。だからDRAM製造はMicron、SK hynix、Samsungという別会社に分かれている。世界でいちばんこれが上手い会社だ。

そのDRAMを何層にも積み、NVIDIAから来るメインのロジックダイの周りに積層してはんだ付けすると、数百GBになる。Blackwellはロジックダイの周りに288GBのHBMを持つ。ロジックダイ自体のSRAMはたぶん500MB程度だ。DRAMとSRAMでは密度が桁で、3桁違う可能性がある。

セレブラスに戻る。彼らが見ているのは、チップ上でSRAM密度を上げる明白な方法はない、という問題だ。だがSRAMは計算をする論理ゲートのすぐ隣にある。ALUが読むSRAMのすぐ隣にあり、行列積をする演算ユニットは途方もない速度でSRAMからデータを引ける。セレブラスが謳うのは秒間ペタバイトで、ウェハスケールエンジン3では21ペタバイト毎秒だ。NVIDIA BlackwellのHBMはだいたい10テラバイト毎秒だ。また桁違いで、容量は多いが帯域は比例して少ない。

セレブラスがやるのはこうだ。ダイを取れるだけ取る。TSMCが課す800平方ミリメートルのレチクル上限に縛られない。ウェハ全体を使い、スクライブライン越しにすべてのダイを互いにつなぐ。ウェハ全体でSRAMを最大化する。1ウェハあたりたとえば50GBのSRAM。それをパイプラインのように何枚も積むと、最大1TBの超高速メモリになる。

その全部をやるのは、1ウェハあたり21ペタバイト毎秒でSRAMから読むためだ。だから言語モデルを非常に高いトークン毎秒で出せる。Kimiのような巨大モデルのパラメータ全体を、論理コアへ約1ミリ秒で動かせる。だから1000トークン毎秒への道がある。

EN

And so what Cerebrus does is they say that we're going to take as many of these dies as we can. We're not going to limit ourselves to the 800 millimeter u reticle limit, the TSMC 800 square millimeter limit that TSMC imposes on us. We're going to take the entire wafer and have actually every die connect to every other die over scribe lines. And we're just going to try to get as much SRAM as we can on the whole wafer. and we can get to like let's say 50 gigabytes of SRAM per wafer and then we're going to stack many wafers together in a pipeline or similar and now we can have you know up to a terabyte of memory very very fast memory and you do all that work just to get to the ability to read data from SRAMM at yeah 21 pabytes per second per wafer therefore you can now serve these language models at extremely high tokens per second because you can move the entire parameter count of a large model like Kimmy uh you can move all that data in and off the chip or sorry in and off the logic cores in about a millisecond or something like that.

パ
パトリック#52

その市場セグメントの予測は。

EN

So there you go you have a path to a thousand tokens per second

ニ
ニール#53

ハイブリッドな着地になると思う。セレブラスのチップが非常に強いところ、メモリへの超高速アクセスと、もっと容量のあるメモリを持つものを組む必要がある。1兆パラメータのKimiのようなモデルを、大量のセレブラス・ウェハに載せることはできる。だがKVキャッシュはどうにもしにくい。KVキャッシュは人がモデルを使うほど増え、常に動的だ。どれだけ必要か事前には分からない。ユーザーが何人いて、何人に出したいかで決まる。

EN

and so what is your prediction for like that segment of the market? Okay. So I think what happens to them is some hybrid sort of outcome like we we had to pair the Cerebras chip where it's very strong. It's very very good at fast access to memory with something that has more capacity for memory because it's true that you can take a one trillion parameter model like Kimmy and fit it on a large number of cerebrus wafers. But you can't do something about the KB cache very easily. The KV cache is something that grows as people use the model more and that is always dynamic. You don't even know how much KV cache you're going to need. It depends on what your users how many users you have and how many users you want to serve.

パ
パトリック#54

KVキャッシュを、基礎から説明してくれるか。

EN

Can you explain KV cache just like in

ニ
ニール#55

言語モデルに通したトークンは、会話が続く限りコンテキストウィンドウに残る。10万トークン話しても、10万1個目のトークンの時点で、それまでの会話は全部残っている。モデルは過去の会話履歴を参照して、次に何を言うかの予測を良くする。

KVキャッシュはメモリの塊だ。通したトークンごとに表現を保存する。しばしばモデルの重み自体より大きくなる。重みは結晶化した知識で、KVキャッシュは今まさにしている会話の動的な知識だ、と考えるといい。

EN

basic? Yeah. So KB cache whenever you use a language model every token you send through the language model actually uh stays in the context window of the language model for as long as you're having a conversation. So if I we talk for 100,000 tokens, the 100,000th and oneth token is still in the conversation uh behind us and the model is referencing all the past conversation history in order to make better predictions about what the next thing we're going to say is. And so that KV cache is a bunch of memory. Um you have to store a representation for every token that you send through the language model. And it frequently gets to be larger than the weights of the model themselves. You have this like crystallized knowledge in the model weights and you have the dynamic knowledge of the exact conversation we're having in the KB cache is the way I like to think about it.

パ
パトリック#56

長い会話の深みで急に劣化して見えるのは、そういう技術的な問題があるからか。

EN

Yep. And and this is why sometimes people would observe like deep in a conversation things start to degrade because there's some sort of like technical problem.

ニ
ニール#57

KVキャッシュはその点で面白い。過去のすべての正確な表現を保存している。会話で見た情報は全部持っている。だが訓練では、モデルは主に超長文脈の会話では訓練されていない。主に8000トークンや1万6000トークンの会話だ。20万トークンまで持っていくと、その長さでの訓練も多少はあるが、モデルの本業ではない。

フロンティアラボの長年の課題は、1万トークンでの賢さを20万トークンでも同じにすることだ。これはずっと続く戦いになる。100万トークンのコンテキストウィンドウという概念はもう何年もある。Anthropicが最初に100万に届いたと思う。僕はClaude Codeで、100万のはるか手前から /compact を使っている。フル長まで行くのが実際に良いとは思っていない。

EN

Yeah. So the KB cache is quite interesting in that regard. The KB cache is an exact representation of everything that came before. We we store all the information that we've seen in the conversation. However, during training, the model did not get trained primarily on very long context conversations. It got trained primarily on, let's say, 8,000 token conversations or 16,000 token conversations. So, if you take the model to 200,000 tokens, there was some training that happened at that context length, but it's not the model's like core strength. And so, there's there's always been a challenge for the Frontier Labs to figure out how do we make the model exactly as intelligent at 10,000 tokens as we expect them to be at 200,000 tokens. And it's going to be a perennial battle for us. We've had 1 million context windows as a concept for for years now. Enthropic was I think the first to hit the 1 million context window length. I still, you know, use /compact in my cloud code uh well before 1 million context length. I don't think it's actually great to hit the full length.

パ
パトリック#58

では超高速、超低レイテンシのアプローチは、結局この要因に制約される。

EN

And so these extremely fast, extremely low latency approaches ultimately are limited by by this factor.

ニ
ニール#59

そう。重みについては何でもできる。重みの保存では無敵の性能も十分あり得る。だがKVキャッシュは大きな棘になる。

EN

Yes, you can do whatever you want for the weights. It's very possible to have just unbeatable performance on weight storage. However, the KB cache is going to be a big thorn on your side. And so three years from now, five years from now, what role do you think these kinds of chips play? Like what sort of market share do they have in the heterogeneous chip market?

パ
パトリック#60

では3年後、5年後、この種のチップはどんな役割を持つか。ヘテロジニアスなチップ市場でのシェアは。

EN

Crisis and Grock and maybe a couple others, you should think of them as accelerators. What they are really good at is being used in conjunction with an more traditional GPU like device that critically has this offchip memory built in. You want offchip memory for capacity and onchip memory for speed. We want to hybridize these two things. So if you take uh transformers in the limit, you take a transformer to a million context length. What ends up happening is you have this you know computebound stage which is the actual matrix multiplies for the uh what we call the MLP which is where most of the model's knowledge world knowledge is encoded and then you have the attention layer which is where we're kind of dynamically adapting to the current conversation. Attention in the limit is usually memory bound and the MLP in the limit is computebound at large enough batch size. And I would say the original sin of transformers is that you've taken this extremely fundamentally memory bound layer and juxtaposed it right next to a computebound layer. It is very difficult to have a single chip that is good at both compute operations and memory operations. The GPU is quite balanced in this regard, but you have to choose one or the other. Cerebrus has a very fast memory access for something like a matrix multiply and it's really good to host the the MLP the the weights essentially on the Cerebrus chip but the GPU has the capacity to scale to really long context lengths and so you would like to put the uh attention possibly on the GPU and the MLP on the cerebrus chip

ニ
ニール#61

セレブラスとGroq、他の数社は、アクセラレータだと思った方がいい。本領は、オフチップメモリを持つ、より伝統的なGPU的デバイスと組み合わせることだ。容量にはオフチップメモリ、速度にはオンチップメモリが欲しい。この二つをハイブリッドにしたい。

Transformerを極限まで、100万トークン文脈まで持っていくと、こうなる。計算律速の段階がある。MLPの行列積で、モデルの世界知識の大半はここに符号化されている。そしてAttention層があり、今の会話に動的に適応する。極限ではAttentionはだいたいメモリ律速で、バッチが十分大きければMLPは計算律速だ。

Transformerの原罪は、本質的に極端にメモリ律速な層を、計算律速な層のすぐ隣に置いたことだ。計算とメモリの両方に強い単一チップを作るのは非常に難しい。GPUはこの点ではかなりバランスが取れている。それでもどちらかを選ぶことになる。セレブラスは行列積のようなものへのメモリアクセスが非常に速く、MLP、つまり重みを載せるのに向く。GPUは超長文脈までスケールする容量がある。だからAttentionをGPUに、MLPをセレブラスに置きたい。NVIDIAとGroqの間でも、これが起きていると思う。

EN

and I believe this is what's happening with Nvidia and Grock. Can you riff for a minute just on transformers and uh

パ
パトリック#62

Transformerについて少し話してほしい。2017年のあの革新を、深く知らない人向けにも。強みと弱みは何か。AIの未来でも支配的な、あるいは支配的なアーキテクチャの一つであり続けると思うか。

EN

yeah, you've been so good at explaining some of the core concepts just for people that again aren't aren't deeply familiar with what this innovation was in 2017

ニ
ニール#63

教師なしデータから非常にうまく学べるようにした。Transformerが結局やっているのは、任意の系列、どんなデータの並びでもパターンを探すことだ。見出しになる部品がAttentionで、系列の中で今いちばん関係があると思う部分に動的に適応できる。Transformerを一歩進めるたびに、それまで見た入力を再重み付けして、次の予測にいちばん効くのはどれかを決める。任意の系列データの学習に極端に向いている。僕たちが日常的に作るいちばん面白い系列は言語で、それで言語の領域を支配した。

さらに引くと、Transformerが本当にうまくやったのはスケールしたことだ。人間の事前仮定を置かない。系列にパターンがあるなら見つける、うまくいくまでパラメータを投げ込む、と言うだけだ。

コンピュータビジョンの仕事からも恩恵を受けている。CVの課題の一つは、線形モデルやSVMのようなレガシー機械学習の数千パラメータ、あるいは数十万パラメータから抜け出すのが難しかったことだ。ディープラーニングで数千万パラメータになり、CVで巨大と言われたのが1億5000万パラメータくらいだった。今は日常的に数兆パラメータの話をする。数百万から数兆へつないだのがTransformerだ。

EN

like what its strengths and weaknesses are and whether or not you think it will remain the dominant architecture or a dominant architecture for the future of AI. What it did was it it allowed us to learn an unsupervised data really effectively because transformers what they're all about at the end of the day is taking any sequence any arbitrary sequence of data and trying to find patterns in that data and they critically the attention operation which is the headline uh component of transformers. It allows the model to dynamically adapt to what it thinks is the most relevant component of the sequence. every step you take through a transformer, you are essentially like reweing the input that you looked at before and figuring out which is most relevant for your next prediction. And so it it's extremely amenable to

パ
パトリック#64

スケールの重要な単位がデータと計算なら、その二つを増やすのが得意だからTransformerは残る、という話になるか。

EN

uh learning arbitrary sequence data. And the most interesting sequences of data that we produce on a regular basis is language

ニ
ニール#65

データは未解決だ。計算については、そうだ。Transformerはすごいスポンジだ。使える計算を10倍にすると、どこかで対数的な改善が出る。これまでのところスケーリング則は本当にうまくいく。美しいとさえ言える。

Transformerが本当にうまいのは、投げたほとんどどんなデータセットにも伸びることだ。非常に強い汎用学習器だ。Attentionを置き換えようとした他の手法より特に効くのは、任意のペアワイズ関係を表現できることだ。系列のどのトークンも、他のどのトークンにもAttentionできる。系列の中に関係があれば、Transformerは見つける。all-to-all、全トークンが全トークンを見るモデリングが不要な場合もある。だが必要なら、その選択肢がある。その空間をもっと刈り込む、情報のモデリングをもっと選択的にするより良い方法が分かるまでは、Attentionは非常に良い演算だ。

これもCVの頃に学んだ手だ。カルパシーの古い言い方で、新しいデータセットでモデルを訓練したいなら、まず過パラメータにしてデータを過学習しろ、関係をモデル化できること、学習アルゴリズムが動くこと、知識を埋め込めることを証明しろ、だ。過学習できたら圧縮する。圧縮が汎化だ。目の前のデータを暗記したいわけではない。汎化したい。だから過学習したあと、できるだけ小さいパラメータ数に収まる一般パターンを逆向きに探す。

EN

and that's how we got to dominance in the language regime. But uh to zoom out even further, I think what transformers really did well is that they scaled. Transformers make no such human prior. Transformers just say, "Well, there's going to be a pattern in the sequence of data, and if there is a pattern, I'm going to find it. I'm going to throw more and more parameters at this problem until it works."

パ
パトリック#66

データの未来予測と、この物語でのデータの重要さを聞かせてほしい。

EN

Uh, and transformers benefit from a lot of the computer vision work, too. For example, one of [clears throat] the challenges in computer vision was we had a hard time going from hundreds of thousands of parameters, which you get for like linear models like support vector machines or other legacy machine learning models. Those had, you know, on the thousands of parameters. Then we got to deep learning and got to tens of millions of parameters with computer vision. The biggest models were you know around like 150 million parameters was a huge model for computer vision. And now we routinely talk about trillions of parameters and transformers are the link to go from millions to trillions of parameters.

ニ
ニール#67

好きな言い方がある。インターネットはデータに対する一回限りの補助金だった、と。タダでもらった。非常に高品質で、高品質テキストが約30兆トークン。良いテキストの定義を広げると300兆トークン。もうだいたい全部見た。モデルはインターネット全体を何度も見ている。インターネット由来の人間データでやれることは、もうあまり残っていない。

次のフェーズは、僕の考えでは、RL環境のジムを通したモデルの自己改善だ。ランダムなユーザーとAIのやりとりを増やしても、もう得にならない。昔はChatGPTを使う人のインタラクションと、好き嫌いの信号が大事な新しいデータだった。今の議論で好きなのは、僕たちが出している中央値のモデルは、条件なしの一般人のフィードバックよりはるかに進んでいて、その信号はもうほとんど価値がない、というものだ。今欲しいのは専門家の選好だ。モデルはEveryday Joeを超えた。

EN

And so if I think about the important units of scaling being data and compute

パ
パトリック#68

一般人、と。

EN

does it stand a reason then that you think transformers will just stick around because that's the thing that we're good at getting more of those two things. Well, data is an open question, but comput. Yeah, transformers are so they're just such great sponges, you know, like you you you increase the compute available to a transformer by 10x and you'll get you'll get some log improvement somewhere. Uh and and so far the scaling laws really work. They're really quite beautiful. And to [clears throat] the point about I guess what do transformers do really well? They extend to almost any data set you can throw at them. They're extremely powerful general learners. And I think what's especially useful about transformers over other techniques that we've tried to replace attention is transformers represent any pair wise relationship that you want. Any token in the sequence can attend to any other token in the sequence. So if there's any relationship that's in the sequence at all, you're going to find it with transformer. Now it may be the case that you don't need all toall modeling. You don't need every token to look at every other token. But if you need to, transformers give you that option. And until we know a better way to kind of prune that space down uh a better way to kind of have information modeling be more selective attention is a very very good operation. This is another kind of trick that we learned in the computer vision days. Like one of the old Karpathy sayings is that you know if you have a new data set that you want to train a model for. Your first goal should be to overparameterize the the model and try to overfitit the data that you have to prove that there is a relationship that you can model or memorize that your learning algorithm works uh that you can instill knowledge into the model. Once you can overfit then you can compress and the compression is how you get generalization. You don't want to actually memorize the data that you have in front of you. you want to generalize and therefore once you overfit the data set then you can kind of work backwards and try to find the general patterns that fit into the smallest parameter count possible.

ニ
ニール#69

Everyday Joe。その通りだ。だからデータの未来は、モデルに検証可能な難しい課題を与え、隔離されたジムで走らせ、進んだかどうかを測れるようにすることだ。コーディングもこのカテゴリだし、数学もそうだ。だんだん、エージェントにコンピュータを渡して人間の労働者のように動かし、目標への進捗でフィードバックする。その環境がデータになる。あまり独自の見解ではないが、これまでのところ非常に生産的だ。

EN

What's your prediction for the future of data and riff on the importance of data in this whole story? I like the phrase that internet was a onetime subsidy on data. We got it for free. Uh it's extremely high quality about 30 trillion tokens of high quality text. Uh 300 trillion tokens if you take a wider view on what qualifies as good text and we've basically looked at it all already. Models have seen the entire internet many times over at this point. And there is not a whole lot more to be done on human data from the internet. The next phase of data in my mind is model self-improvement through RL environment gyms. Basically, in fact, we don't even benefit from getting more like random user interactions with AI. It used to be that, you know, the the new type of data that we cared about a lot was the interaction data from people using chatbt and giving TetBT signals on what they liked and didn't like. I like the argument now that the median model that we serve is so much more advanced than the kind of un than like a random human uh giving feedback that the signal you get from random human preference or I guess unconditioned human preference is not actually worth anything anymore. You want expert human preference at this point. The model has outgrown everyday

パ
パトリック#70

それは長い間続くのか。インターネットを一つの大きな塊だとすると、これも日の当たる大きな塊で、取り尽くしたらまた次に行く、という話か。

EN

generic Yeah. Everyday Joe. Exactly. So the feature of data to me is giving the model a hard verifiable task and letting it run in this gym where it's kind of isolated and it just has a a problem that it can make progress on and get measurement of whether it made progress on that problem or not. You can imagine coding problems are in this category. math problems are also in this category and um increasingly more and more we have we can just give the agent a computer essentially and have it act like it's a human worker and just give it feedback on whether it's making progress towards the target outcome that environment becomes the data. I think this is not a super differentiated take but uh it's been really really productive from what I've seen so far. And you think that just goes on for a really long period of time or is that another like if I think about the internet as this one big block like this is another big block that will have its you know day in the sun and we'll kind of get it all and and then we'll have to move on to something else.

ニ
ニール#71

もっと根源的だと思う。AGIが欲しいなら、特化知能を積み上げて隙間がなくなるまでやるのがいちばんいい。成立条件はただ一つ、課題が検証可能であることだ。自己採点の仕組みを渡す。それがあれば、好きな課題で自己改善するレシピがある。フロンティアラボの支出の仕方もそれを裏付けている。昔はデータにかなり使っていた。今はRL環境にもっと使う。これらの環境は、検証可能な課題での再帰的自己改善という関係を、まさに捉えている。

EN

I think it's actually more profound than that. Basically the idea is that if you want artificial general intelligence the best way to get there is to just keep stacking specialized intelligences until you have no more gaps to fill. And the test here, the only thing you need to make sure you do to make this work is you must make sure that your task is verifiable. You need to give the model a self-grading system. If you have that, you have the recipe for self-improvement on any task you like. And I think you've seen this held up by the way frontier labs spend. They used to spend much that much on data. Now they spend a lot more on RL environments. And these environments absolutely capture that relationship of recursive self-improvement on a verifiable task.

パ
パトリック#72

脇道に何本か入ったが、既存のハードウェアをもっと効率よく使う、最初の課題に戻ろう。

EN

Okay. Okay. Now, so I like that we've veered off in different little side cars here, but coming back to your initial task of making existing hardware more efficient.

ニ
ニール#73

そう。

EN

Yes.

ニ
ニール#74

ソフトウェアで、ハード層で起きていることをより手元で制御することだ。

EN

Um so, so yeah, just keep going on what you've done so far and what you want to do and then we're going to jump to hardware and then jump to energy finally.

パ
パトリック#75

ここまでやってきたことと、これからやりたいことを続けてくれ。そのあとハードウェアに飛び、最後にエネルギーに行く。

EN

Sounds good. So, yeah, I mentioned kernels. It's surprising people think kernels are done. There are great people like Triau who write excellent kernels and they're they form the bedrock of all of our um modern deep learning is built on flash attention. Modern transformers are built on flash attention. But if you deviate from the happy path at all, if there's a new model that comes out that has a slightly different way to embed positional information like the change of the rope system. Suddenly the kernel that we had is not suitable for this new model and we may have to make a a patch to this kernel. I wouldn't say we're in the phase where we had to invent new kernels from scratch, but having the ability to quickly modify existing GPU kernel, sorry, a kernel, by the way, is a it's a general term for any program you run on the GPU. And so, historically, kernels tend to be put into a library where every kernel has a very very scoped purpose. Typically, you have a kernel for a matrix multiply. You have another kernel for even something as simple as addition. You want to add two tensors together, that's another kernel.

ニ
ニール#76

了解。カーネルの話をしたよね。驚くのは、カーネルはもう終わったと思っている人がいることだ。トライ・ダオのような人が優れたカーネルを書く。それが現代の深層学習の土台で、現代のトランスフォーマーはFlashAttentionの上に乗っている。

ただ、ハッピーパスから少しでも外れると話は変わる。新しいモデルが出て、位置情報の埋め込み方が少し違う、たとえばRoPEの仕組みが変わった、といった場合だ。手元のカーネルは突然そのモデルに合わなくなり、パッチを当てないといけない。ゼロから新しいカーネルを発明する段階だとは言わない。でも既存のGPUカーネルを素早く改造できる能力は要る。

カーネルは、GPU上で走らせるプログラム全般の呼び名だ。歴史的にはライブラリに入れられ、それぞれ目的が非常に狭い。行列積用のカーネルがある。足し算のように単純な演算にも別カーネルがある。テンソルを2つ足す、それも別カーネルだ。

だんだん、それらを融合し始めている。行列積をして、別の行列積の結果に足したいとき、2つを1つのカーネルにする。演算をヒューズする。足し算のためだけにDRAMへ書き出して読み戻すんじゃなく、その場でやってしまえる。

EN

[clears throat]

パ
パトリック#77

なぜまだ人間がこれをやっている?AIが、より効率の良いカーネルを書くのは得意そうに見える。そこに向かっているけど、まだそこまで来ていない、ということか。来ていないなら、なぜ人間がまだやっている?なぜトライ・ダオはあんなに有名なんだ。知っている名前だ。

EN

And then increasingly we've started to fuse those kernels together. So if I do a matrix multiply and then I want to add it to another matrix that I've also multiplied maybe those two become one kernel and I just fuse the operations where instead of writing the data out to DRAM and then reading it back in just to do the addition maybe I can just do this uh easily.

ニ
ニール#78

トライの代わりに話すつもりはない。ただ彼が教えてくれたのは、もう必ずしも手でカーネルを書くべきではない、ということだ。僕らの言い方では、ホワイトボードでカーネルを書く。ホワイトボードに行き、機械が何をすべきだと思うかを記述し、それを自然言語で簡潔にモデルへ渡す。するとモデルが実行に移せる。入力と出力はこう。GPUへの仕事の載せ方はこう。実装する、と。

EN

Why are humans still doing this? Like it seems like the sort of thing that AIs would be exceptionally good at engineering more efficient kernels. Maybe that's where we're going and we're just not quite there yet. But if if we aren't there yet, is that where we're going? If we're not there yet, why why humans still doing this? Why why is Tree out so wellknown? You know, it's a name I know.

パ
パトリック#79

つまり概念設計を人間がやっている。

EN

I don't want to speak for Tree, but what he taught me was uh you shouldn't write kernels by hand anymore necessarily. I like to say we write kernels in the whiteboard. We go to the whiteboard, we describe what we think the machine should be doing, then we succinctly describe that in in natural language to the a model. And then the model is able to do the execution of okay, here is my input and output. here is the strategy of how we want to dispatch this work onto the GPU. I'm gonna go implement this.

ニ
ニール#80

その通り。なぜモデルがここをうまくやれないのか、正確にはわからない。これがうちの堀だとも思っていない。6ヶ月後には、カーネルエンジニアリングがずっと上手いモデルが出ていると思う。ラボ側は、もうかなりのカーネルエンジニアリングを完全自動でやっている、と言うだろう。

EN

So, we're doing the conceptual design.

パ
パトリック#81

ソフトウェアという優位の話だ。ソフトウェアを、基盤ハードウェアのスピード・オブ・ライト、理論上限に最大限近い使い方だと考えると、時間とともに、君たちのような会社の優位にはならなくなっていく。

EN

Exactly. And that I think I'm not sure exactly why models are not superb at doing this. I don't think this is like our remote or anything like that. I'm sure in 6 months time we'll have much better models uh on kernel engineering and I'm sure the labs would tell you that they already do a lot of their kernel engineering uh in a fully automated way. And so software as an edge,

ニ
ニール#82

その通り。MythosやGPT-5.6のような上げ潮が、すべての船を浮かせる。本当にそうだ。カーネルエンジニアリングだけうまくやるモデルを作る、という特化に意味があるとは思っていない。コーディング全体の中でも、いちばん意味のある部分集合ではない。カーネルに限っても、プロンプトに特権的な情報を入れてモデルを誘導するのは有用かもしれない。でも広く言えば、この能力については僕らはみんなフロンティアの下流にいる。

EN

yeah,

パ
パトリック#83

エネルギー史でいつも好きな話がある。原資源、たとえば石炭と、その塊から何パーセントを取り出せるか、の間で振り子が揺れる。エネルギー史の大きな部分は、その数字を10%から95%へ持っていくことだった。今どこにいる?Blackwellをその石炭だと考えて、何パーセントだと思う。今ある1枚を、どれだけ効率よく使えている?

EN

if I think about software as maximally near speed of light, efficient usage of of an underlying piece of hardware

ニ
ニール#84

切り口はいろいろある。GPUがいちばん喜ぶこと、つまり大きな次元の行列積をやっているときには、性能の最適化はかなりうまくいっている。その演算はピーク使用率の70〜80%で走り、律速はソフトウェアではなく電力だ。NVIDIAが掲げるピークFLOPSは少し楽観的で、電力スロットリングがあるからそこに届くことはない。

EN

is is going to trend towards not being an advantage for a company like yours over time.

パ
パトリック#85

熱のせいだ。

EN

That's right. The rising tide of something like Mythos or GBT 5.6 Soul that lifts all boats. It really does. Um I actually don't think there's a point in specializing to say we work on making the model better for just kernel engineering. I think that's actually not not the most meaningful subset of of like coding in general,

ニ
ニール#86

その通り。サーマルだ。70〜80%なら飽和していて、かなり良い。ただ実際には、トランスフォーマーの時間の大半を、そのハッピーパス、つまり大きなバッチの行列積では過ごさない。だから僕らの仕事は、チップの周りにエンジンを組み、GPUに常に大きなバッチの仕事を食わせ続けることだ。

GPUの世界でこの1年いちばん大きな転換は、もうGPUを1個ずつプログラムしない、ということだ。ラック全体、クラスタ全体、データセンター全体で考えるべきだ。NVIDIAも、単体GPUやマザーボードだけでなく、ラックシステムごと仕様として出し始めた。NVL72と呼ぶ。最新チップのGrace Blackwell 300は、72基入りのラックとして出荷される。ラック規模のコンピュータを誰がいちばん効率よくプログラムできるか、公開レースになっている。この形のコンピュートが、効率と速度の両方の未来だと思っている。実際NVIDIAはよくやっている。最低レイテンシが欲しいならそのチップを使うべきだし、最高スループットが欲しいなら、今のところたぶん同じチップを使うべきだ。非常に新しいプログラミングのパラダイムだ。

EN

uh, kernel engineering in particular. Maybe there's some like privilege information you inject into the prompt that's like a useful way to steer the model to be better at writing kernels, but broadly speaking, yes, we're all we're all downstream of the frontier in terms [clears throat] of this capability. I I always love this uh this from the history of energy there there's always this like pendulum between the raw source let's say coal

パ
パトリック#87

よく聞く話として、最高のチップ、たとえばBlackwellの市場は今、麻薬市場みたいだ、と。できるだけ多く手に入れるために奇妙なことが起きている。みんな極端に足りていないから。その比喩に反応してほしい。実際そんな感じか。それから、最先端ではないチップの市場も聞きたい。少し、あるいはそこそこ劣るチップでもいい、と思ったとき、その市場はどうなっている。その世界に入れてくれ。

EN

and then if there's a certain amount of energy available inside of a chunk hunk of coal like what percent of it we can harness and use

ニ
ニール#88

いくつかある。第一に、NVIDIAは自社の全チップについて長期の見方を持っている。Blackwellチップへの需要は巨大だと見ている。他の供給者が過去にやってきたように、ただ値段を吊り上げ、需給曲線がどこかで交わり、みんなは一応ハッピー、ということもできる。でもNVIDIAは、いちばん懐の深い買い手にチップを全部買わせると、その顧客が長期でもっと力を蓄えて自分たちが傷つく、と見ている。今、コンピュートは力そのものだ、と理解している。だから計算資源の割り当てはかなり戦略的だ。これが一つ目。

二つ目は、関係がとても大事だということ。何ヶ月かしかやっていないスタートアップが「Blackwellを1万基、3年か5年借りる」と巨大なレンタル注文を入れてくるのを、誰も喜ばない。金を払える保証がない。今、計算資源へのアクセスを人に納得させるのはかなり難しく、かなり良い関係か、圧倒的な資金力が要る。NVIDIA側でこれを通すのは、希少性があまりに高く、需要が桁外れだからだ。

他のチップについては、劣っているとさえ呼ばない。悪いチップはない。あるのは本当に悪い価格だけだ。 正しい価格なら、どんなチップでも動かす。それが会社の信条の一つだ。

AMDの話をしよう。全体として良いチップだ。課題は、人々がプログラムの仕方をよくわかっていないことだ。カーネルチームの話をしたよね。ハードウェアから性能を絞り出すことに本気だ。NVIDIAは自社チップについてはそれをかなりうまくやる。僕らが絞り出せるアルファは少しある。でも他のチップの方がやることがずっと多い。ベンダーが、NVIDIAほど最初から最良のカーネルを用意しないからだ。僕にとってさらに良いのは、アルファがあることに加えて、AMDはNVIDIAほど良くない、という認識が他の人にあることだ。願ったりだ。このチップを軽視したままでいてくれて、僕が買えるだけ買えるなら、とても嬉しい。

ただ、それはもうあまり本当ではないと思う。AMDは一部の大口買い手の間で、そこそこ人気がある。公開情報では、MetaとOpenAIがAMDチップを大量に買っている。AMDの供給もどんどん割り当てられている。ただ、新しく出てくる会社のロングテールがある。Etched、SambaNova、d-Matrixのような新規は良い。彼らの主な課題はスケールだ。TSMCから十分なウェハ割り当てを取り、市場に出せる量のチップを量産できるか。新しいチップが出たら、できるだけ早く知り、供給のかなりの割合を買えるか評価したい。

EN

and a big part of the history of energy was getting that number from 10% to 95% or whatever

パ
パトリック#89

つまり根本的には裁定だ。注目が薄いチップから性能をずっとうまく引き出せれば、マージンを乗せて売り直せる。良い商売になり得る。

EN

right

ニ
ニール#90

その通り。他のチームがスキル不足でチップを活かせない、という話ではない。僕たちはそこそこ上手い。複数のシリコンアーキテクチャを使い、思いもよらない場所で性能を追い込むチームとしては、世界でも指折りだと思う。

大事なのは、新しいチップに合わせてスタックを組む速さだ。「データセンターの事業者は今あるクラスのチップに縛られていて、新チップの展開は大ごとだ」という慣性を、僕たちはほとんど持っていない。素早く動いてくれる創造的なデータセンターのパートナーがいる。新しいクラスのパートナーも出てきた。何より、挑戦から逃げない。TPUが好きだ、TPUを動かす。Trainiumが好きだ、Trainiumを動かす。簡単に動かなくても、ヘテロジニアスなサービングシステムの中に居場所を作る。どのチップにも比較優位がある。それを見つけて、その方向に絞る。

EN

where are we in that's like if I just think about it at Blackwell or something and Blackwell is the piece of coal

パ
パトリック#91

ハードウェア、データセンター、エネルギーに入る前のインタールードとして聞きたい。投資家層の心配を、君はどう見ている? メモリ株を見る、とか。僕が今いちばん好きなのは、S&P 500に占める半導体の割合のチャートだ。歴史的には2%、3%、4%。今は19%、20%、21%。市場史を学ぶ人間から見ると、短期で狂ったピークをつけて、長期平均に崩れ戻るものがいくつもあった。みんなそれが怖い。

Micron や SK hynix みたいな会社で大金を稼いだ人は多い。でも長期では計算はコモディティで、世界の時価総額の4分の1や5分の1にはならない、と感じて怯えている。それがセットアップだ。不足は巨大だと誰も認める。でも「いつか解決して、資本市場での居場所は普通に戻る」と感じている。このナラティブをどう思う?

EN

like what percent do you think we're at like How how efficiently can we use an existing piece today?

ニ
ニール#92

ひとつ。僕は歴史の学生というより、歴史の一員だ。1997年生まれで、母は2000年に向かうドットコムの熱狂と崩壊の頃、Intelにいた。Ciscoが世界で最も価値のある会社で、Intelがすぐ後ろにいた時代を覚えている。25年前とその対比で今を見ている。

いちばんの違いは、当時のネットワーク機器への投資の多くが投機だったことだ。来ると見込んだユーザー需要は来なかった。トークン消費、広く言えばAI消費が面白いのは、もう投機ではないことだ。人はトークンを、今すぐ価値があるから買う。溜め込まない。すぐ使う。

2年前とも違う。2023年、2024年のHopper世代の供給逼迫は、全部トレーニング向けの支出で、トレーニングは本質的に投機だ。今はみんな、Claude Codeにいくら使えるかの上限を設けている。推論支出の話をして、それが増えると予測する世界はまったく別だ。推論支出は単調に増えると思う。推論支出に投機はない。

EN

There's a lot of different ways to analyze that. I think in some level we are really efficient at optimizing the performance when the GPU is doing the thing that it's most happy doing which is a large dimension matrix will apply that operation runs at you know 70 80% of peak utilization and it's limited not by software but by power. The way Nvidia quotes peak flops is a little optimistic. You never hit that because of power throttling but um

パ
パトリック#93

ハードウェアに戻ろう。チップ、ラック、クラスタと単位を見てきた。データセンターの話がしたい。面白いことをしているパートナーがいると言っていた。今とこれからのデータセンターを、君の目で。推論を全部さばくにはここがクリティカルで、イノベーションも多いし、君もここを見ている。

EN

because of heat.

ニ
ニール#94

会話のテーマに戻ると、トレーニング対推論だ。2年前はトレーニング向き、今日は推論向き。AIスタック全体でいちばん保守的なのはインフラ側だ。データセンター、さらに保守的なのがTSMC、チップのインフラ。

歴史的にAIデータセンターはトレーニング向けに建てられた。トレーニングは推論のスーパーセットで、トレーニング用クラスタは推論にも使えるが、逆は必ずしも成り立たない。差はネットワーキングだ。チップ間の帯域にいくら投資するか、クラスタをどれだけ大きくするか。

データセンターには規模の不経済がある。1つのデータセンターにGPU 10万枚を置くのは、1万枚より、1000枚より、ずっと高くて難しい。今はメガワット、ギガワットの話になる。アメリカでギガワット級のデータセンターを簡単に建てる方法はもうない。100MWですらますます難しい。特別な顧客以外はほぼ不可能。10MWが今日できることのギリギリで、1MWはむしろ豊富だと僕は思う。

市場にはこういう通説がある。電力は合計ではたくさん見つかるが、一点には集まらない。トレーニング用データセンターを建てる人には興味がなかった。全部一箇所にある前提で、データセンター横断のトレーニングは誰もやりたくない。

市場にはラグがある。100MWや10MWのデータセンターを、見つかる限りしぼり取る、というのがまだ多くのデータセンター開発者の態度だ。でも新しい考えの人が出てきた。推論は分散した1MWデータセンターに向く、と。僕たちもまったく同意で、全米の小さな計算プールを買って推論フリートにするのは大歓迎だ。

EN

Yeah, exactly. thermals let's say 70 80% it's saturated it's pretty good

パ
パトリック#95

1MWと10MWの、物理的な大きさを感覚で。

EN

but in practice you don't spend the majority of your time in a transformer in that happy path where you're doing a large batch m matrix multiply

ニ
ニール#96

液冷が入ってから、かなりマニアックになった。1ラックに詰め込める電力密度が異常だ。1MWの計算というと巨大なデータホール、倉庫みたいなものを想像するだろう。今はラック約8本に詰められる。ラック1本は冷蔵庫くらい。8本並べると、それが1MWだ。

EN

and so

パ
パトリック#97

君が助けて作りたい未来は、いろんなチップを一緒に使えることだ、と。

EN

uh our job is to basically build the engine around the chip such that we are feeding the GPU these large batches of work at all times

ニ
ニール#98

そう。

EN

and one of the most profound transitions we've had in the GPU world in the last year has been this moving of you know you don't program one GPU at a time anymore you should think about the whole rack and maybe you should think about the whole cluster, the entire data center at a time. And with Nvidia again, they've started shipping not just a single GPU or a single motherboard, but actually the the whole rack system is something that they prescribe. They call it NVL 72. Uh their latest chip, the Grace Blackwell 300, um that ships as a rack of 72 units. And it is it's an open race to figure out who can program the whole rack scale computer as efficiently as possible. And my belief is that that shape of compute is the future of both efficiency and speed. In fact, Nvidia does a great job of if you want the lowest possible latency, you should be using that chip. And if you want the highest possible throughput, you should probably also be using that chip as of right now.

パ
パトリック#99

買い手として、チップ1枚から最大限を絞る。そのチップをごく小さなデータセンターに載せて、推論だけやる。安すぎるものも含めた寄せ集めの計算と、そこからより多くを食う力、それを小さな単位のデータセンターに出す。イコール、ずっと安い知能。

EN

And it's all comes down to like this is a very new paradigm of programming.

ニ
ニール#100

そう思う。取り方を工夫すれば、安いFLOPSの取り方はたくさんある。僕たちのやり方を一言で言うと、世界のどこでも、どんなチップでも、どんな期間でも買う。この柔軟性と流動性は、今ほかに誰も持っていないと思う。口にしたことを金で裏づける。どんなキャパシティでも取って、フリートで動かす方法を見つける。それが今日の優位の大きな部分だ。長期では、ほかの人が懐疑的なデータセンターに投資して、その優位をもっと作らなければいけない。

1000個の小さなデータセンターの軍勢対、1つの巨大ギガワットデータセンター。何が起きるか。多くの場合、電力の冗長はない。現場のバックアップディーゼルもない。高いから全部切る。ネットワーキングの冗長もないことが多い。電力にアクセスできる施設に置き、電源は単一。光ファイバーは1本埋設する。冗長とフェイルオーバーとSLA付きの3本は引かない。たまに落ちる。中には稼働率95%まで落ちるところが出ても、僕は驚かない。

EN

One of the things you hear is that the market for the best chips, blackwells, let's say,

パ
パトリック#101

それは悪い。

EN

is like a drug market or something right now. Like there's all sorts of fascinating things happening to get as many of them as possible because everyone's so short. Y

ニ
ニール#102

すごく悪い。ほかの誰かには致命的で、ひどすぎる。ギガワット級では生き残れない。

EN

I'd love you to react to that analogy like is that what it feels like

パ
パトリック#103

稼働率95%のデータセンターに、買い手はほぼゼロだ。

EN

but then also to talk about what the market is like for like not the bleeding edge chips like if I if I

ニ
ニール#104

僕がその最初の買い手だ。95%を買う。

EN

am willing to accept a slightly or or moderately inferior chip

パ
パトリック#105

バックグラウンドで動いているなら気にしない、というエンジンの話だから、か。

EN

what's that market like let us into that world

ニ
ニール#106

部分的には。実際は二つある。一つは、かなり頑丈なコントロールプレーンがあること。どの単一データセンターの単一障害でも、ほかと相関していなければ扱える。ワークロードをほかへ移せばいい。障害はある率で起きる。稼働率95%対98%対99%は、僕にはほぼ線形に良い悪いだけだ。

もう一つが、前に言った非同期の部分。長時間ホライズンのエージェントをサーブしている。リクエストが落ちたら、別のGPUを探して載せ直す。エージェントは1時間動いていて、GPUが引き抜かれて障害に当たる。その1ターンだけ、1分、2分、3分、ときには10分のレイテンシが乗る。でも顧客は気にしない、というのが僕の主張だ。エージェントは何時間も動いている。

EN

yeah okay a couple things number one uh yes basically has a long-term view on uh on all their chips they they see this immense demand for the black hole chips and they they can do what other suppliers have done in the past which is like just crank prices and made the market you know supply and demand curves will correct they'll intersect at some point and everyone will be technically happier but Nvidia sees the if they just let the most deep pockets buy all the chips that maybe hurts them in the long term if that customer ends up acrewing a lot of more power they understand that compute is power today and so uh they're quite strategic about how they allocate compute that's the first thought the second thought is that relationships matter a lot nobody wants to have a huge order of of a chip rental come in from this new startup that says, "Oh yeah, I'm going to rent 10,000 black wells for for three years or 5 years." The startup has only been operating for months typically. Who knows they're good for the money. The the way you convince someone to give you access to compute is is quite challenging uh these days and requires some pretty either great relationships or uh just incredible financial backing to make this happen

パ
パトリック#107

寝ている。

EN

on the Nvidia side. And it's all because the scarcity is so high and demand is just off the charts. Now, for other chips, I wouldn't even call them inferior. I I like to say there's no bad chips. There's really bad pricing. And uh I will make any chip work at the right price. That's like kind of one of the ethoses of the company. And let's talk about AMD. AMD, I think great chips overall. The challenge is that people don't um understand how to program them very well. So, you know, I've been talking to you about how we have such a great kernel team. We're so serious about squeezing the performance out of the hardware. Nvidia is pretty good at doing that for their own chips. Frankly, there's some alpha that we can squeeze out, but actually there's a lot more to be done on other chips because the vendor does a little bit less work than Nvidia does to make the best kernels out of the box or or even better for me there is alpha and just like other people have this perception that AMD is not as as good as Nvidia. That's music to my ears. I'm very happy for them to sleep on this chip and for me to buy as much as I can.

ニ
ニール#108

関係ない。たまに1ターンが少し長くなっても構わない。顧客にはこう言う。平均スループットは十分に戦える。ただしP99、99パーセンタイルのレイテンシはコントロールしない。できない。その代わり、勝てない経済性を渡す。バックグラウンドエージェントにはそれが合う。

EN

Now, I think that that's not actually super true anymore. I think AMD is actually uh somewhat popular amongst the some large buyers. Um you know I think publicly Meta and OpenAI have bought a ton of AMD chips and so we're increasingly seeing that uh all the AMD supply is also being allocated but there's a long tale of other companies that are popping up yet net new companies are great uh such as etched or or senova or dmatrix all these companies are popping up and I think the main challenge for them is scale can they actually get enough wafer allocation from TSMC to pump out chips to make it into the market but certainly if there's a new chip on the I'd like to know about it as quickly as possible and evaluate whether we can buy a good fraction of that supply.

パ
パトリック#109

電力というカテゴリの話を。面白いこと、新しいこと、どこへ行くと思う?

EN

And so it's fundamentally an arbitrage for you. Like if you can be much better at eking out performance from chips that have received less attention, you can then resell that at a margin and it could be a great business.

ニ
ニール#110

データセンターに稼働率95%が欲しいと言った。値段次第なら80%でも取れるか。たぶん取れる。それが意味するのは何か。僕はカリフォルニアの息子で、太陽光と風力が好きだ。アメリカでは太陽光と風力は全然使い切れていない。ずっと課題だったのは断続性だ。ベースロードは常時なのに電源が間欠的だから、データセンターには向かない、とまで言われる。どうするか。

実はその問題の解決は、そんなに遠くないと思う。データセンターの停止が日単位、週単位でも耐えられる。風力が吹かず、雲が出て、谷に霧がかかる。データセンターにとって最悪の悪夢である長期停止だ。でもそれはかなり予測できる。起きたら世界の別の場所のキャパシティを呼べばいい。天気をモデルして、オフラインになるタイミングを見て、ワークロードを移す。それでいい。トリックは、誰も触りたがらない電力に手が届くことだ。あの種の停止は、扱うのがあまりに面倒だから。

チップが十分安ければ、たぶんNVIDIAのラックではない。チップが十分安ければ、遊んでいるチップの資本コストも気にしない。

EN

Exactly. Exactly. And I think that it's not the case that everyone else is just, you know, has a skill issue that they can't uh, you know, make these chips work as well. I think we're quite competent in this. I think we're probably one of the best teams in the world to use multiple silicon architectures and and be quite aggressive in chasing down performance in unlikely places. But um yeah, I think it's the speed at which we we're willing to kind of build our stack around a new chip. We don't have a huge amount of incumbency around well our data center providers are only stuck with with this class of chip and it's going to be a huge pain for us to to deploy uh these net new chips. We have some very creative data center partners who are willing to move very quickly and there's a new class of those that we can talk about and most importantly we don't shy away from the challenge. Uh that's frankly a big part of this is just saying yes we love TPUs we're going to make TPUs work. Yes we love tranium we're going to make tranium work and if it doesn't work um easily we're going to find a way to fit it in with the heterogeneous serving system it will have a place every chip has a comparative advantage we have to find that advantage and then squeeze it in that direction. Ju just as an interlude before we get to hardware, data centers, energy, etc. which will be really fun part of the conversation. I'd love you to talk about your perception of the investor classes worry. Yeah.

パ
パトリック#111

このシステム全体を、スカベンジャー戦略だと言っていたね。

EN

Like you look at memory stocks.

ニ
ニール#112

その通り。

EN

Um or my current favorite is you look at the chart that plots the percent of the S&P 500 that's semiconductors.

パ
パトリック#113

その比喩をもう少しほぐしてほしい。

EN

Historically it was like 2 3 4%. Now it's 19 20 21%. And it just sort of looks like if you're a student of market history, you get all these things through time that are sort of reached some crazy near-term peak and then and then collapsed back to long-term norms.

ニ
ニール#114

まずチップを拾い、そのチップのための電力を拾う。どちらの場合も、AnthropicやOpenAIと計算資源の入札でぶつかりたくない。勝てないし、勝ちたくもない。もっと創意を使って、彼らが今日は供給として読めないものを使う。時間をかけて、集約した供給を積み上げる。集中した供給は絶対に手に入らない。手に入るのは集約供給だけだ。そうやって、経済性で誰にも負けない集約工場を作る。

工場を作っている。世界一の製鉄所を作ろうとしている。ただし大型の一体型プラントではなく、ミニミルの集まりとしてだ。

EN

Um, and I'm curious how that has all investors worried.

パ
パトリック#115

垂直統合の度合いで、いくつもの型が想像できる。極端は全部自前だ。資本集約になる。電源を持ち、データセンターを建て、チップを設計し、そのチップを食い尽くすソフトを握り、完成したトークンをエンドユーザーに売る。ユーザーは僕みたいな人間だ。スタックを全部持つ。

逆に、線の引き方はどこでもいい。極端に資本を薄くして何も持たず、全体のコーディネーション・プレーン、バーチャルなスカベンジャーになることもできる。どの型の事業になるか、どう考えている。

EN

Yeah. Uh so you know a lot of people made a lot of money in Micron and Skhinx and companies like this but everyone feels like ah these you know on the long term like compute's a commodity and uh it will not represent a quarter or fifth of the entire market capitalization of the world and and so they're scared and that's the setup. Um everyone acknowledges that like there's a huge shortage but everyone sort of feels like ah we'll figure it out and these things will revert back down to their their normal place in capital markets. I'm curious what you think about about that narrative.

ニ
ニール#116

その問いを受ける自分は、実は二人いる。一人は、毎日きちんと動き、できるだけ持続可能に、できるだけ速く伸ばさないといけない会社のCEO。もう一人は創業者で、こっちは想像力がずっと強くて、この世界が好きでたまらない。創業者の自分は全部やりたい。これが人生そのものだ。チップ、電力、エネルギーのことばかり考えて生きてきた。気になるのはこれしかない。だから当然、野心は最大にしたい。止めたくない。原材料から完成品まで、いちばん効率のいいシステムを作り終わるまで、絶対に止まらない。

EN

One thing, I'm less of a student of history as more of a member of history. I was I was born in 1997 and uh my mom worked at Intel in the 2000 in the run-up to the year 2000 and the the do boom and crash and you know I remember the time where Cisco was the most valuable company in the world and and Intel was close behind. I mean I mostly draw parallels to that period of history from 25 years ago to today. And I think the main difference is that a lot of the investment in networking equipment historically was speculative. We anticipated this future demand for users that never came. And I think what's interesting about token consumption or AI consumption broadly is that it's no longer speculative. People buy tokens because they're immediately valuable to them. You don't hoard tokens, you use them immediately. This is also even different from what we had 2 years ago where there was a supply crunch for hopper generation chips in 2023 2024. Uh in that period it was all training oriented spend and training is inherently speculative. Now it's everyone is instituting caps on how much you can spend on cloud code. It's a very very different world to be talking about inference spend and predicting inference spend to go up. I do think inference spend monotonically increases. Uh there's no speculation on inference spend. Vanta automates security and compliance for over 16,000 fast-moving companies like Ramp, Cursor, and Harvey, keeping them audit ready around the clock. It's the number one Agentic Trust platform, and it now helps companies like yours watch for the risks that show up between audits across your vendors, your AI tools, and your whole [music] environment. Every new tool your team signs up for, every vendor that turns on AI features, [music] is an opportunity for something to go wrong. And most security programs weren't built for AI's pace of growth. The Vant agent works like a 24/7 GRC engineer in the background, finding issues, drafting fixes for you, and cutting vendor assessment time by up to 50%. Whether you're a fast growing startup or a global enterprise, [music] Vanta helps you earn and prove trust. Invest like the best listeners. Get a special offer of $1,000 off Vanta at vanta.com/invest. Ridgeline is the first endto-end system of record with embedded AI for investment [music] management firms running portfolio accounting, reconciliation, reporting, trading, and compliance, all on one unified platform. Firms are moving off legacy technology and onto Ridgeline because of how far ahead Ridgeline's AI features are compared to anything else in investment [music] management software. I've been hearing from a lot of investment managers about AI, and they fall roughly into two camps, with some unsure where to even start, [music] and others convinced they can build their own order management system over just a weekend. The reality is that running an investment firm will always require governance, [music] controls, and a single source of truth for your data. And no amount of AI enthusiasm changes that requirement. If you're serious about your firm's AI strategy, Ridgeline [music] should be part of that conversation. And you can request a demo at ridgeline.ai.

パ
パトリック#117

リアルのファクトリオ(Factorio)をやってるみたいだね。

EN

Coming back now to your take on hardware. And so the unit level is interesting to me like talked about chips, talked about racks, talked about, you know, clusters. I'd love to talk about data centers

ニ
ニール#118

まさにそう。それが感情の答えだ。CEO側はもっと現実的でいないといけない。全部自前で持つ資本は、言った通り正気じゃない。ソフトはレバレッジが高いから、ソフトから始める。でも最終的に、発電を自前で持つのか、電力会社と優れた電力購入契約(PPA)を取れるのか。僕は、歴史的に得意なことを他者に専門でやらせて、そこから規模を取りにいく方に傾いている。翼の下に引き入れる権利を稼げる規模まで行きたい、という感覚だ。

スタックのどこにでも、取りこぼしている効率はあると思う。買う相手は、自分の顧客が誰かを仮定して作っている。その仮定を壊せるなら、余地が出る。かなり楽観的な見方だ。可能なのは、コンピューティング史上最大の計算市場を引き受けようとしているからだ。推論に何十億、何兆ドルもの投資を積み上げていく。その焦点があるから、推論専用のものをたくさん作る意味が出る。可能な場所を全部探すのが自分の仕事だ。自分とパートナーにそれが明らかになったら、パートナーに専用のものを作ってもらう。できなければ、自分で作る。

EN

and you said you've had some interesting partners doing some cool things. Talk us about the the present and future of data centers as you see it.

パ
パトリック#119

ソフト、ハード、エネルギーまで含めたこのシステム全体を引いて、有用な知的トークンを作るうえで、今いちばん非効率な場所を順位づけすると、リストはどうなる。

EN

Y

ニ
ニール#120

計算のスケーリング自体はかなり効率的だ。FLOPSを増やせば、そのFLOPSを使う。FLOPSの使い方は、もうかなり慎重でもある。今のモデルを見ると、10%超の密度のものはほとんどない。起動できるエキスパートのうち、実際に起動しているのが10%という意味だ。フロンティアモデルは1%に近い。すでにかなりスパースだ。MoE側で無駄を出しすぎているとは思わない。長いこと触ってきたし、絞るのは上手い。

下手なのはアテンションと、そのメモリの使い方だ。とくにKVキャッシュは、今かなり無圧縮のまま置いてある。KVキャッシュのエントロピーを見ると、その容量に見合っていない。居場所を稼いでいない。トークンあたり何KBもKVキャッシュに載せている。たぶん1桁、あるいは2桁ずれている。フロンティアラボが何をしているかは分からない。ただDeepSeekは、そこをさらに圧縮する面白い仕事を公開し続けている。進捗もいい。毎年くらいのペースで桁が動くという事実自体が、まだ余地が大きい合図だ。

ここまではミクロだ。もっと引くと、計算資源の動員の仕方がまったくなっていない。世界にはこれだけの計算がある。NVIDIAは今年、Blackwellを500万個流している。全部どこへ行った。常時使われているのか。とてもそうは思えない。どこかで、世界の計算をもっとうまくオーケストレーションする必要がある。難しいのは、計算の多くが日の当たらない私有プールに消えて、GPUが悲しげに遊んでいることだ。物理的に痛い。シリコンと電力が入ったものが、ただ座っている。直したい。世界の計算を共有資源として編成し、もっと密度高く詰めたい。

xAIのクラスタの総FLOPS稼働率を、みんなでからかう。でも世界の残りは、それよりずっと悪い。大量のGPUが倉庫に座っているか、特定顧客に割り当てられた私有プールに座っていて、使われていない。

EN

because this seems like you know obviously a critical thing for being able to serve all this inference is like lots of innovation in in this part of the world and obviously you're focused on it.

パ
パトリック#121

その効率を、かなり直接に攻めている。その方がよほど効く。ファブそのものの未来はどう見る。メモリメーカー、TSMC、Intelその他は、どう容量を増やすのか。アメリカでやるのか。指を鳴らして今あるチップが100倍あれば、トークンはもっと安くなる。この世界の重要な部分だと思うので、見方を聞きたい。

EN

I think one of the themes in our conversation has come back to what is training versus inference like what is the difference between these two? categories and you know what was different about two years ago being training oriented and today being inferenceoriented and I think the most conservative players in the entire AI stack have got to be the infra players whether that's data centers or even more conservative is TSMC the chip infra people uh and so data centers historically were built like AI data centers they were built for training uh and training is the superset workload over inference you can make any training cluster work for inference but maybe not vice versa and what the difference there is networking uh how much do you invest in bandwidth between chips and how large of a cluster do you need? There's actually a diseconomy of scale to to data centers in some way. Like it's way more expensive and difficult to build a um you know 100,000 GPUs in one data center than it is to build 10,000 than it is to build 1,000. And and we now we just talk about you know how many megawatts or gigawatts do you have? And basically there's no way to build a gigawatt data center in the United States easily anymore. Even 100 megawatts is is increasingly hard. It's basically impossible unless you're a very special set of customers. Uh 10 megawatts is probably on the edge of what's possible today and 1 megawatt I would argue is plentiful. So there's this incredible lore on the market where you can find lots of aggregate power but it will not be concentrated and that was not interesting to anyone who's building training uh data centers because you just assume all be in one spot for no one wants to deal with cross data center training.

ニ
ニール#122

面白いのは、全部が互いのバランスで伸びるということだ。指を鳴らして全部を倍にしても、TSMCのボトルネックを直した次に、別のボトルネックにぶつかる。チップを20%増やせば、すぐ次のボトルネックが出る。

ただ興味深いのは、彼らが「必ず守るべきもの」と見なしているものだ。顧客である僕たちが、いつも欲しがる不変条件だと思っているものと、僕がもっと流動的な関係だと思っているものの差。ファブがトレードオフをもっと見せてくれれば、自分に何ができるか、もっと賢い判断ができる。

いちばん面白い例が、どのファブにも、ラインから出てくる最悪チップと最良チップのあいだにかなりの幅があることだ。チップの出来にはばらつきが大きい。TSMCのような会社は、いわゆるプロセスコーナーを必死に締める。最悪チップの特性を、最良チップにできるだけ近づける。それで高い歩留まりが取れる。ただしそのために、工程にたくさんの制御を足している。僕には要らない制御かもしれない。最悪チップの居場所を、自分で見つける気がある。工程管理をそこまで締めなくていい。締めるほど時間もコストも食う。リジェクトをもっと引き取る気もある。

僕らにとっては、ダイのコスト、ダイの供給、電力のコスト、置ける場所、その全体最適化だ。目標は、アメリカ全土の電力供給を劇的に広げて、本来ならデータセンターに居場所を稼げなかったチップにも家を用意することだ。

EN

So the market has some lag in it. I think that the market still assumes that we have to go shake down those 100 megawatt and 10 megawatt data centers wherever we can find them is still the attitude I hear from a lot of data center developers

パ
パトリック#123

自分の会社というシステムの設計の話をしたい。何を学んだか。NVIDIAまわりの面白い教訓は出てきた。この北極星に向けて、チームと事業をどう組んでいるか、文化の中に連れていってほしい。

EN

but increasingly we're seeing a few new thinkers realize that inference is going to be suitable for these distributed 1 megawatt data centers and uh we're we're quite in agreement with that and we are very happy to buy small pools of compute across the United States and use that as our inference fleet. give us a sense of literal physical size of uh 1 megawatt versus 10.

ニ
ニール#124

「極限」で考えることが多い。モデルに取りかかった直後の効率は気にしない。1ヶ月後、6ヶ月後、1年後にどこまで行けるかを見る。相手にする機械の状態を、固定だとは思わない。Blackwellですら、この性能を出すのを止めているボトルネックがあると思えば、よく理解して特性を書いて残す。NVIDIAや仲間に伝えるためでもあり、次に買うチップのためでもある。自分たち、会社にとって長期ではほぼ不変だと思うことを学び、将来の判断に織り込みたい。

かなり協働で動く。いちばん欲しい特質は、良い生徒であり、優れた教師でもある人だ。チームの多くは大学でTAをやっていて、知識をこう分け合うのが好きだった。ホワイトボードはしょっちゅうやる。誰もが教えるものがあり、学ぶものがある、あの大学的な空気が、僕らにはとても大事だ。

EN

Yeah. Well, so this got really wonky with the advent of liquid cooling. Now you can pack insane levels of power density into a single physical rack. Like a megawatt of compute, you you'd imagine this like massive data hall, like a huge warehouse basically. And now you can actually pack that into Yeah. around like around like eight racks worth of compute. Each rack is about the size of a refrigerator. You can just imagine eight of them lined up. Um, yeah, that's a megawatt. And so your view would be that the future that you want to help build is a whole bunch of different chips that can be used together. Yes.

パ
パトリック#125

3年後、機械がもっと仕事を引き受ける環境でも折れない人は、どんな属性か。

EN

That you can buy, you know, you're a buyer to ek out the most per chip.

ニ
ニール#126

好奇心だ。100%好奇心。教えられないものが一つある。性能への愛、機械が動いているマイクロ秒の一つひとつに潜って、その瞬間に何が起きているかを理解したがること。パフォーマンスエンジニアにとっていちばん大事な特質で、僕が探すのもそれだ。AIの経験はたくさん要らない。CUDA経験も要らない。むしろ大きなレッドヘリングだ。CUDAという概念も、GPUという概念も、この5年でかなり変わった。10年の経験を求める意味がない。それは教えたい。でも性能エンジニアリングへの愛は教えられない。探すのはそれだ。

EN

And that those chips can then be coupled in very small

パ
パトリック#127

主要ラボを一つずつ評してほしい。クローズドソースというカテゴリとオープンソースの関係、今何が起きていて、これから何が起きるかも。

EN

data centers

ニ
ニール#128

一言で言うと、ラボは他より3〜6ヶ月先にいるために莫大なプレミアムを払っている。今でもそれはたぶん見合う。OpenAIとAnthropicがやっていることは、完全に理にかなっている。

クローズドとオープンのフロンティアの関係の核心に、蒸留という扱いの難しい話がある。蒸留は盗みだ、フロンティアモデルの出力から何かを奪っている、という感覚がある。別の見方を出したい。たとえ意図がなくても、Anthropicからデータを掻き集めようとしなくても、僕たちがインターネットに出す成果物の、ますます大きな割合がAI生成だ。GitHubだけ見てもいい。この1年に作られたリポジトリの何パーセントが、Claude Codeで作られたと思うか。それを蒸留と呼ぶのか。たぶん、それだけで足りる。GitHub上の、良いと見なせるオープンソースのコード出力だけを使って、フロンティアクラスのモデルを訓練できるとしても、驚きはない。

ユーザーはAIとのやり取りの出力を自分のものとして持ち、それをGitHubに上げる。実際、かなりの人が上げている。その立場を取るなら、潜在的な蒸留は長いあいだ残る。情報やモデル能力の拡散を根本から止めるのは、不可能だと思う。起きる。問題は速さだけだ。

EN

to just do inference. And that those two steps of a whole bunch of random compute, some of which is cheaper than it should be,

パ
パトリック#129

では次の問いは、スケーリング則と改善則が永遠に、あるいはかなり長く続くかどうかだ。続くなら、3〜6ヶ月先にいる価値はある。続く限りその優位は残り、オープンソースの安いトークンに対して巨大なプレミアムを取れる。この見方で合っているか。

EN

your ability to eat more out of it, and then small units of expression in a data center

ニ
ニール#130

ありうる。ただ、3〜6ヶ月先にいるプレミアムが、そんなに長く続くかは分からない。企業の導入は、3〜6ヶ月の速さでは動かない。多くの企業はまだOpus 4.6やOpus 4.7にいる。最先端をすぐに採用しない。変更を一つ出すだけでも、疑問がたくさん出る。まだ表面を引っ掻いている段階なので、この競争の勝者を決める方法はないと思う。そもそも、決着がつく競争だとも思っていない。ずっと続く過程だ。オープンソースが消えることも、根本的にはない。一人のリーダーが抜けて真空ができれば、新しいリーダーが入る。インセンティブが大きすぎるし、追い風も多い。フロンティアクラスのモデルを訓練するのは、毎日簡単になっている。

EN

equals way cheaper intelligence. I certainly think so. Yes, there's a lot of ways to access cheaper flops if you're able to be creative with what you take. And so, one of the ways that I describe what we do is we will buy any chip anywhere in the world for any duration of time. That is a level of flexibility and liquidity that I think no one else has right now. Uh we're very aggressive about putting our money where our mouth is and we will we will really take any capacity uh and find a way to make it work in our fleet. And that is a big part of our advantage today and long term we had to create more of that advantage by investing in these data centers that other people are going to be skeptical of because you know what's going to happen when you set up these like this army of a thousand small data centers versus the one the one big gigawatt data center. Well, few things. You're not going to have power redundancy frank quite often. You're not going to have backup diesel generators on site. Those are all very expensive. We cut all that overhead. We're not even going to have redundant networking in a lot of cases. We're going to put these in facilities where we have good access to power, a single source of power and we're going to trench one line of fiber to these data centers, but we're not going to have like three lines of fiber with redundancy and failover and SLAs's. It's just going to go down sometimes. In fact, I won't be surprised if some of them get down to like 95% uptime,

パ
パトリック#131

では望む未来のバランスは何か。クローズドとオープン。スタックを握っている強みで全部をやるモデル会社。Anthropicはそれができる。新しいGoogleみたいになり得る。Google自身がそれをやることもできる。どんな未来を望む。

EN

which is bad.

ニ
ニール#132

豊富なトークンと、多様なハーネスが欲しい。みんなに自分のハーネスを作ってほしい。どの会社も、どのユーザーですら、エージェントを自分のものにしてほしい。その水準のカスタマイズと能力までは、もうそんなに遠くない。人に知能の主権を持ってほしい。その知能のカスタマイズは、たぶん重みのファインチューニングではなく、もっと文脈内学習(in-context learning)の側になる。技術的な細部だ。

ただ、この豊富さの未来の入力は、要するに安いトークンだ。自分の仕事は、人間に可能な限りトークンを安くすることだ。それを達成する。使えるスタックのすべての層でやる。供給側のレバーが好きだ。あらゆるチップを使う。あらゆる電源を使う。アメリカでこれに向いている土地は、全部使う。その見返りとして、人は豊富な知能を持つとはどういうことかを探る動機を持つ。

今でもエージェントを、相談料の高い人間みたいに扱っている。難しい問いのときだけ聞け、と。知能の捉え方として、それは違う。機械が考えられるのはすごい。できるだけ多くの手に、できるだけ多くの人に渡すべきだ。

EN

Very bad. That's fatal, atrocious for anyone else. survive in a in a big a big gig

パ
パトリック#133

この未来を現実にしようとしている席は、かなり独特だ。同じように詳しくて関心の高い友人たちと比べて、いちばん外れている見方は何か。頭が三つある人間を見るみたいな顔をされるアイデアは。

EN

you'd have basically zero buyers for a data center that has 95% up time

ニ
ニール#134

チップの話がほとんどだ。カスタムチップを作ると言うと、「何が違うの」と聞かれる。要するに、HBM不足を迂回して、フラッシュなど別のメモリへの極端なオフロードに寄せることだ。この考えはかなり気に入っている。チームはみんな知っている。KVキャッシュをフラッシュへ、もっと本格的にオフロードするために、モデルアーキテクチャを何から変える必要があるか。ずっと太鼓を叩いている。ホワイトボードも、ずっとそれだ。

推論のコミュニティの中でも、秒間1〜10トークンで出す前提でシステムを組むと何ができるかについて、かなり外れた見方がある。それが僕らの北極星だ。もっと広く言うと、人はどうやって1日1兆トークンを消費するのか。彼らがそれをできる世界を作りたい。

EN

I'm that first buyer I will buy 95%

パ
パトリック#135

1兆トークンって、どれくらいか感覚を合わせてほしい。

EN

up time

ニ
ニール#136

1兆トークンは、OpenAIの価格だと、GPT-5.5や5.6でも少なくとも500万ドルだ。

EN

and the reason for that is because of this background engine thing that if there's things running in the background you don't care

パ
パトリック#137

ドルで見るのがいちばんいい指標だな。いまの価格で、1人1日500万ドルかかる消費が起きる世界は、どういう世界だ。

EN

partially uh it's actually two things one is that we have a really robust control plane that is going to be fine handling any single failure in any single data center as long as it's not correlated with other data centers and I can just move the workload somewhere else I'm cool with that um the failures happen at some rate and I am basically linearly happy with a data center that's 95% uptime versus 98% versus 99%. It's just linearly good or bad for me.

ニ
ニール#138

トークン単価を、少なくとも3〜6桁は改善する必要がある。5000ドルまで落とせば、顧客は出てくる。実際、あるサイズのモデルでは、1兆トークンが数万ドル規模で測れるところまで来ている。1ジョブで回す想像がつく。

EN

Mhm.

パ
パトリック#139

平均的な人は、そんなことはできないし、やらない、と心配にはならないか。今、自分の脳でもやっていない。世界に、知能への需要はそんなにない、という話だ。

EN

Now, you do need that async piece that I mentioned of, you know, we serve these long horizon agents because what happens when a request fails is that I'm going to have to go find a new GPU to put that request on. And that means that for that single turn of the agent's work, you know, it's working for an hour, but then it hits a roadblock because it's GPU got pulled away. At that moment in time, that agent is going to experience maybe like an extra minute or two or three, maybe even 10 of latency. But my argument is that my customers don't care because their agent was running for hours.

ニ
ニール#140

それは絶対に信じない。世界に知能への需要は、常にある。課題は、その知能への入り口だ。プロダクトのコミュニティの仕事だ。自分はプロダクトの人間ではないから、最高のビジョンを持っているとは言えない。

EN

They're sleeping.

パ
パトリック#141

その人たちを可能にしたい、と。

EN

Doesn't matter. [laughter] It doesn't matter if like a single turn occasionally becomes uh a little bit longer. Yeah. So we tell our customers, look, our average throughput is going to be very competitive, but our P99, our 99th percentile latency, it's not going to be controlled. It cannot be. And in return, I'll give you unbeatable economics.

ニ
ニール#142

その人たちを可能にしたい。「無料枠のユーザーには使わせられない」「この量のトークンは出せない」という感覚で、絶対に止まってほしくない。顧客からずっと聞いている。直したい。

EN

And I think that's the right fit for background agents.

パ
パトリック#143

逆の問いも聞きたい。自分がいちばん狂っていると思う話ではなく、コンセンサスで間違っていると思うものは。

EN

Talk about power as a category. What are you seeing that's interesting, innovative, where do you think this goes?

ニ
ニール#144

何度も戻ってくるのが、NVIDIAの話だ。短期ではNVIDIAに強気だ。彼らに逆らう賭けはするな。いつも自分を作り直す。ただ、人を驚かせるのはこれだ。HopperからBlackwell、Rubinと、同じ土俵で比べる。BF16乗算の性能/ワットは、そんなに伸びていない。さらに一段、TSMCに行く。TSMCの5nmと4nmと3nmと2nmを見ても、これらのチップの性能/ワットは劇的には変わらない。

その帰結として、人は地政学で頭がおかしくなる。何らかの理由でTSMCを失ったらどうなる、と。コントラリアンな見方を言うと、そんなに悪くない。供給は確実にショックを受ける。ただ、西側が持っている最良のプロセス、たとえばIntelは、そんなに後ろではない。最悪でも、性能/ワットが2倍悪い程度だ。チップ戦争まわりの議論を追っていると想像するほど、差は大きくない。

EN

Okay, so I said I want 95% uptime on my on my data centers. Could I even take 80% up time at the right price? Probably. Um, and what does that mean? Well, I'm a son of California. I love solar and wind. I think solar and wind power is way undertapped in the United States. And the challenge has always been this intermittency. You would even consider solar and wind unsuitable for data centers because you have a persistent base load and an intermittent power source. What are you going to do? Well, I think we're actually not that far from solving that problem. I am totally capable of tolerating a outage for my data center that's measured in in even days or weeks which is like the worst case nightmare scenario for a data center is that we're going to have a long-term outage because the wind is in blow and the clouds are in the sky fog is hanging over the valley for some time. That's the worst case scenario. It's in fact highly predictable and I can just call in capacity in some other place of the world whenever that happens. uh I'll just model the weather and figure out when my data center is going to be offline, move my data my workload somewhere else and it's fine. The trick is that it's going to give me better access to power that no one else is going to touch because it is so annoying to deal with that kind of outage.

パ
パトリック#145

この道筋に乗っていないAIの世界で、いちばん興味があるものは何か。このシステム全体の部品として、自分で手を出す話ではないもの。

EN

And if my chips are cheap enough, they're probably not going to be Nvidia racks. And if my chips are cheap enough, I don't mind the capital cost of having idle chips.

ニ
ニール#146

僕らは完全にモデルの下流にいる。アーキテクチャを決めるのはモデル側だ。OpenAIやAnthropicに口を出すことは、ほとんどない。僕にとって都合のいい方向に進んでくれと祈るしかない。あるいは、彼らがどこへ行くかをできるだけ予測して、サービングのアーキテクチャをそれに合わせて組む。ソフトの選択も、ハードの選択も、彼らが持っている。

ある意味いちばん面白いゲームは、また機械が考えることの深さに戻る。スパースアテンションかデンシアテンションかを決めることが、どれだけ結果を変えるか。データ型を変えることが、どれだけ結果を変えるか。BF16で訓練していたのが、今はFP8やFP4、より低い精度で訓練できる。恣意的な選択に見える。だが、どのチップが使えるか、ハードをどう組むべきか、計算の未来をどう見るかに、深い含意がある。

EN

Yeah. So I've heard you describe this entire system as like a scavenger strategy.

パ
パトリック#147

新しい計算スタートアップを作りたい起業家が100人部屋にいるとする。チップ、システム、ラックなど、ハードを作りたい人たちだ。世界は全部試すし、それは良いことだ。何かは当たる。ただ、この先の世界で事業をうまくやるために、どう会社の向きを決めればいいか。技術の賭けそのものではなく、会社の型として、何を助言する。

EN

That's right.

ニ
ニール#148

全部、サプライチェーンのボトルネックだ。まず僕か投資家を説得してほしい。現代のチップ供給を決める3〜5個のボトルネックを理解している、と。TSMCのウェーハ容量、HBM容量、先端パッケージング。四つ目は電力かもしれない。電力をどこから取るか。ラックをどう組むか。この四つのボトルネックそれぞれに、どう迂回するかの良い答えが欲しい。結局は全部アービトラージだからだ。

チップを作るのは、NVIDIAが簡単には変えられない選択をしたと思っているからだ。それは本当だ。NVIDIAは、変えにくい選択をたくさんしている。完璧ではない。よくバランスが取れているだけだ。だから尖れ。何かを選んで、「HBM不足の影響を、彼らは過小評価している。僕らは別の方向に全力で押す」と言え。余談だが、攻めるならたぶんそれがいちばんいい。

EN

Is is that Yeah. Unpack that analogy a little Well, first we scavenge chips and then we scavenge power for those chips. The idea is in both cases I do not want to be in bidding against Anthropic or Open AI for compute capacity. I'm not going to win against them and I don't want to. I want to be more creative and use the supply that they don't find legible today. And over time I amass enough aggregate supply. I'm never going to get concentrated supply. I will only get aggregate supply. And over time I build my aggregate factory that is unbeatable in economics.

パ
パトリック#149

なぜ。

EN

We are building a factory. We're trying to build the best steel factory in the world. Uh but it will come through mini mills not through large monolithic steel plants.

ニ
ニール#150

メモリのファブを一気に増やす簡単な方法がない。しばらくかかる。ボイシの連中(Micron)は、循環的な巨大CapExが嫌いだ。

EN

And and if I imagine the different versions of this like how vertically integrated you can be.

パ
パトリック#151

何度も火傷してきたからな。でもその不足があるなら、世界はシステムの他の部分を効率化する方向で迂回するんじゃないか。

EN

Yeah.

ニ
ニール#152

僕は逆で、他のものが全部高くなると思う。iPhoneはメモリを削る。iPhoneの価格は上がる。それで付き合う。

EN

One version would be the extreme would be you own everything. So that it's a very capital inensive business. You own the power you know source. You build the data centers. You design your own chips. You control the software that eats the most out of those chips. And you sell the end finish token

パ
パトリック#153

なぜNVIDIAは最後まで行って、トークンを売らないのか。

EN

to your user like your user is me

ニ
ニール#154

NVIDIAはここが本当に賢い。顧客と競合しない。すべてを長期で見る。じゃあなぜNeoCloudから始めないのか。裏口でコンピュートを売らないのか。NVIDIAは上手い。ジェンセンは友人を億万長者にするのが上手い。CoreWeaveを何十億ドル企業にした。その善意を壊す必要がない。NeoCloudと推論プロバイダの多様なコミュニティを作り、NVIDIAの需要を作るために競い合わせたい。誰か一人が垂直統合したり、AMDや他の選択肢に走っても、その席を埋めたがっている人があと3人いる。買い手同士が競争するのは、彼にとって最高だ。

EN

and you just own the whole stack.

パ
パトリック#155

いつもの締めの質問だ。誰かにしてもらった、いちばん優しいことは何だった。

EN

But you can imagine many other permutations of the business

ニ
ニール#156

いちばん優しいこと。すぐ浮かぶのは、長年のメンターたちだ。スケジュールを大きく割いて、ほぼ個人的な関心にして、相手が何かを理解するようにする人は稀だ。教える。あるいは、理解の寸前まで来ていると思っていることを、一線を越えるまで押し込む。先に話したNVIDIAの人たちは、性能エンジニアリングへの愛を植え付けてくれた。大学の教授たちもそうだ。2年生のときのアドバイザーを覚えている。すごくせっかちな学生で、オフィスアワーに行ってこう言った。「AIチップを作りたい。やりたいことは分かっている。ネットワークやOSみたいな基礎科目で時間を無駄にしているのはなぜだ」と。彼は僕を見て、スタック全体を広げてくれた。パズルの一片一片を理解することの深さ、美しさを見せた。システムの一部に集中しようとする道筋を受け取って、こう言った。ゲートレベルのシリコンから、インターネット規模の優れたサービスを作るところまで、スタック全体を理解できる人は本当に稀だ。一生かけて、その水準の理解に到達する人であれ、と。

本当に稀な特質だ。その水準の専門性を追うのは高貴だと思う。それは今でもかなり残っている。

EN

where you know you could whatever you draw the line anywhere you could be incredibly capital light own nothing and just be like the coordination plane

パ
パトリック#157

ニール、最高の会話だった。時間をありがとう。

EN

across all this stuff the virtual scavenger how do you think about that question of like which which type of these businesses to be You're doing real life factorial basically. we have to be more pragmatic. I think that the capital we're we're looking at for owning everything is like you said, it's insane. Yeah, software has high leverage, so we have to start with software, but ultimately, you know, do we own power generation or can we get great power purchase agreements with uh utilities? I'm more inclined to pursue like letting other people specialize in the things that they're historically good at and then see if we can get to the scale. I think of it as like I want to get to the scale where I earn the right to take this under our wing. I absolutely think that there's efficiencies to be gained everywhere in the stack. If you can break the assumption that people I would be buying from, they made assumptions about who their customers would be. And I maybe break those assumptions. It's a pretty optimistic view. Uh I think it's only possible because we're actually trying to underwrite the largest market for compute in the history of computing. we're actually going to build so many billions, trillions of dollars of investment into inference. Uh, and because of that focus, it makes sense to build a lot of things that are custom for inference. And it's my job to seek all the places where that's possible. And then as as they become obvious to me and my and my partners, I will get my partners to build custom things for me. And if they can't do it for me, I will do it myself. Yeah. I think compute scaling is actually like very efficient. Uh as in like you give me more flops and I will use more flops. And I would say we're actually fairly judicious already with our use of flops. Uh if you look at a modern model, there are very few models that are more than 10% dense, meaning 10% of the possible number of experts you can activate are activated. And I think the frontier models are closer to like 1%. So fairly sparse already. I don't think that we're wasting too much on the MOE side. People have been working with for quite some time. They're pretty good at squeezing. Where we are not good is attention and its use of memory. Specifically, the KV cache is quite uncompressed right now. I think if you look at the entropy in a KV cache, it's nowhere near it's not earning its keep. Like we're storing many kilobytes of data in the KV cache per token. Um, and that's probably off by an order of magnitude or two. And I I don't know what the Frontier Labs do, but Deep Seek certainly publishes really interesting work to compress that further and further. And they're making good progress. And I think the fact that they're able to make order magnitude progress here every year or so signals that there's a lot more room to go. I guess this all on the micro scale. If you zoom out further, I think that we actually don't marshall our compute effectively at all. Like we have all this compute in the world. Nvidia is pumping out 5 million Blackwell chips this year. Where are they all going? Are they all being used at all all the time? I certainly doubt it. I think that at some level we just need better orchestration of compute across the world. Uh this is very difficult to do because a lot of the comput disappears into private pools of compute that will never see the light of day and those GPUs sit very sadly idle. Uh it's actually it pains me physically to see that those GPUs are just you know silicon and power went into that and it's just sitting idle and I want to fix that. how we or organize and orchestrate the world's compute as a shared resource and and pack it more efficiently. I would estimate that, you know, we all make fun of XAI for having, you know, some challenges with total flop utilization on its clusters, but um the reality for the rest of the world is it's far far worse. A ton of GPUs just sit in warehouses or sit in private pools allocated to a specific customer um just don't get utilized. That's way more effective. Yeah. Um, what about fabs? Like what do you think is the future of fabs themselves? Like I think everyone is wondering be able to how will they expand capacity basically? Will we do it here in the US? Yeah. Well, it's interesting. Everything grows in balance with each other, right? If we snap our fingers and double all those things, you might fix a TSMC bottleneck that there you're just going to run into another bottleneck. You make 20% more chips, then you have another bottleneck immediately. I will say though, it is interesting what they consider to be a mustd deliver. uh like what they consider to be like an invariant that their customers me are always going to want versus what I think of as like a more fluid relationship. I think that if a fab exposes more of their trade-offs to me, I'm able to make more intelligent decisions about what I think I can I can do. Can we talk about how you designed the system of your own business? But like bring me into the culture and how you structure a team and a business where this is the northstar. Curiosity. It's 100% curiosity. You know the one thing I cannot teach is love for performance, love for uh digging into every microscond that the machine is working and understanding what's happening on the machine at that time. That to me is the most important trait for a performance engineer and it's what I look for. I don't look for lots of AI experience. I don't look for, you know, CUDA experience at all. That's actually a huge red herring. I mean, CUDA as a concept or GP as a concept have evolved so much in the last 5 years. There's no point asking for 10 years of experience. I want to teach that, but I cannot teach the love for performance engineering. That is what I seek. in a line. I would say the labs pay an immense premium to be 3 to 6 months ahead of of everything else. Uh and I think that's probably still worth it. I think it makes perfect sense for open and anthropic to do what they do. You know there's a sensitive topic around distillation which I think is part a very core piece of the relationship between closed and open frontier. And you know I'd like to offer an alternative view on that which is there is the sense that distillation is theft that you are taking something from the frontier models when you distill on their outputs. And in fact, even if that's not your intent, even if you don't ever try to, you know, scrape data from anthropic, one thing I'll offer is that an increasingly large percentage of the artifacts we put out on the internet are AI generated. Even if you just look at GitHub alone, you know, what percentage of repos created in the last year do we think were created by cloud code? Um, do we consider that to be distillation? Because that's probably all we need. I would not be surprised if you could train a fable glass model only on the outputs of code you consider good on GitHub that's open source. And certainly if we take the position that users own the outputs of their interaction with AI and they choose to put that up on GitHub, which a lot of them do, we're going to have latent distillation for a long time. It seems fundamentally impossible for me. Like I I don't think it's fundamentally possible to prevent the diffusion of of information or model capabilities. It will happen. The question is just how fast. And so then the question becomes, do scaling and improvement laws hold forever or for a really long period of time? And if they do, then there's value to being three and six months ahead and that will just last as long as it lasts and they can charge a huge premium for those tokens relative to a very cheap open source token. Is that the right way to think about it? And so your hope of what the future looks like is what like what balance between closed and open, you know, what balance between model companies doing everything because they have the advantage of owning the stack or whatever. You know, Enthropic can do that, you know, is like the new Google could Google just do that or something. What do you hope the future looks like? every company every user even make the agent your your own. Uh I think we're we're very not that far away from that level of customization and capability. I want people to own their intelligence and I want that intelligence to be customized probably not through weight fine-tuning but probably through more in context learning. That's a more technical detail. But the underlying input to this abundance future is about is basically cheap tokens. My job is to make the tokens as cheap as humanly possible. I will achieve that and I will do it through every layer in the stack available to me. I love the supply side levers. I will use every chip. I'll use every source of power and I will use every piece of land in the United States that's you know suitable for this. And in return, people will have the incentive to explore what it's like to have abundant intelligence. We still treat the agent as a person that is expensive to consult and you should ask them when you have a hard question. That's not the way to think about intelligence. It's incredible that the machine can think and we should try to get that into as many hands as as many people as possible. Most of the ideas on chips, I would say. You know, when I talk about building custom chips and they ask me, "Oh, so what's different?" Basically, it's it's about sidestepping the HPM shortage and focusing on more extreme offload to other forms of memory such as flash. Um, I'm quite passionate about that idea. Everyone on my team knows that I keep banging the drum around like what would we have to change about the model architecture to make offloading KB cache to flash work at a much greater level. And um I'm whiteboarding that all the time. That's like in the community of like inference people. you know we have some divergent views on what you can do if you design a system around serving at you know one to 10 tokens per second which is our whole north star more broadly I think there is this like larger sense around you know what do you do how do people consume a trillion tokens per day like that's the world we want to create the capability for them to do that a trillion tokens well okay at openi pricing that's at least $5 million at the very least for 5.5 or 5.6 six. Yeah, I think the dollars was probably the most good metric. Yeah. Yeah. So, what's the world in which we consume what currently costs $5 million per person per day? Are you at all worried that just like the average person just can't and won't do that like doesn't do that now with their own brain? Like there actually isn't that much demand for intelligence in the world. You want to enable those people. What about the inverse question? Not what you think is craziest, but like what consensus thing you think is wrong? so the consequence of this is people lose their minds over geopolitics like what what happen if we lost access to DMC for any reason. And um my contrarian take is that it wouldn't be that bad. Supply would take a shock for sure, but the best processes that we have in the west uh like Intel not that far behind at worst like maybe 2x uh worse performance per watt and the gap is just far smaller than than you would make it out to be if you talk if you follow like the chipboard dialogue. What else is happening in the AI world that is not in your path? Meaning it's not like a component of this whole system that you would end up doing something in that interests you most. if you had a 100 entrepreneurs in a room, all of whom wanted to create some new compute startup. um and let's say they were specifically wanted to make hardware chips or systems or racks or whatever. because it seems like we're going to try everything and that will be great for the world. You know, some stuff will work. But if you had to give them advice on how to orient their business to be successful in this coming world, what advice would you give them? It's all about the bottlenecks on supply chain. So, you need to first convince me or convince an investor that you understand the like three to five bottlenecks that dictate modern chip supply. There's TSMC wafer capacity, there's HPM capacity, and there's um like advanced packaging, and maybe a fourth one would be power. Like, where will you get the power? How will you build these racks? Uh and I I want to hear like you should have a great answer to each of those four bottlenecks and how you're going to work around them because it's all arbitrage at the end of the day. You're building a chip because you think that Nvidia has made some choices that are difficult for them to change, which is true. Nvidia makes a lot of choices that are difficult for them to change. They're not perfect. They're just really well balanced. And so, you want to be spiky. You want to pick something and say, I think they've underpriced the impact of how short we're going to be on HBM. We're going to push really hard in this other direction instead. Which, you know, as a as an aside, I do think is probably the thing to attack most. so it's going to be a while until we They they've been burned on that many times. I think they're gonna make everything else more expensive. Think that iPhones will cut their memory. iPhones are going to go up in price and um we're just going to deal with it. Why doesn't Nvidia go all the way to the end and sell tokens? Do you think It's great to have competition amongst his buyers. The kindest thing I mean I my immediate first thought is like all the mentors that I've had over the years. It's a rare person who takes a lot of time out of their their schedule and um and you know makes it like their personal interest essentially to to make sure that you understand something that uh or teach you something or or like ingrain some value in you that they think that you're on the cusp of understanding but just push you over the line for understanding. uh a lot of the people in Nvidia that I mentioned earlier who instilled that like love of performance engineering in me but also my professors in college who I remember like my adviser in like sophomore year I was very impatient student so I show up at his office hours and say like I I want to build AI chips I know what I want to do why am I wasting time taking all these like other basic classes and networking and you know operating systems and he just looked at me and said like you know he laid out basically like the whole stack and showed me the depth of or the beauty of like understanding every piece in the puzzle like he he took my entire path of like trying to focus on one piece of the the system and said that you know it's so rare that someone can actually understand the entire stack from the gate level silicon all the way to building a great internet scale service and you know you should aspire to be someone who over the course of your lifetime achieves that level of understanding. Thank you so much for having me.

ニ
ニール#158

呼んでくれてありがとう。

EN

right you know there's actually two parts of me to receive that question. One is the CEO of a company that needs to work every day and and grow as sustainable and and as quickly as it possibly can. The other is the founder and the founder is much more imaginative and and just loves this stuff. The founder in me wants to do everything. This is my entire life. I spent my entire life thinking about chips, power, energy. Like all I care about is this stuff. So of course I want to be maximally ambitious. I don't want to stop ever. I will never stop until I have built the most efficient system from soup to nuts. Very much so. very much so. So that's like the emotional from the hard answer. On the CEO side, I think If you had to just zoom out on this entire system, software, hardware, energy, etc., and stackrank the places that you think that we are the most inefficient today at producing useful intelligent tokens. What does that list look like? You're attacking the efficiency of that very directly. will the memory companies will TSMC will Intel and others Yeah. Um, yeah. Riffon fabrication of chips themselves. Like if we could just snap our fingers and have 100 times the chips, you know, in the stock today, uh, we'd probably have way cheaper way cheaper tokens. So yeah, that that seems like an important part of the universe to hear your view on. One of the most interesting examples here is that any fab has a lot of spread in their like worst chip that comes out of the production line and the best chip that comes out of the production line. There's a lot of variance in how chips are made. Uh and then the question is like you know if you have a company like TSMC they work very very hard to tighten what we call these process corners. We want to keep the worst chip as close in characterization to the best chip and they get a great lens to make that possible. But that means that they are adding a lot of controls in the process that maybe I don't need. Maybe I'm actually willing to find a place for that worst chip. You don't need to tighten the process control as much which takes more time and cost. Uh maybe I'm willing to take a lot more rejects. And I think for us it's like a more holistic optimization around there's you know cost of the dies, supply of the dies and then the cost of power and places we can put them. And my whole goal is to actually so dramatically expand the supply of of power uh across the United States that I have a home for a lot of chips that otherwise would not have earned earned their place in a data center. What lessons have you learned? You talked about some of interesting Nvidia lessons. Yeah. I think there's a lot of um you know in the limit thinking we don't worry about the immediate nature of like when we start working on a model the efficiency is not going to be very good. Uh but we we we think about like where we could end up in in like a month or or six months or a year's time. We don't accept the state of the of the machines we work on as fixed like even something like the blackwell chip if we think that there's some bottleneck that is holding us back from achieving this performance. I mean it's very important to me that we we understand and characterize that very well and write it down so we can both a tell Nvidia about it or friends and also to basically keep this in mind for future chips that we buy. We want to learn things that are what we think are um essentially like invariant for us or the company uh long term and and kind of fold that into future decisions that we make. We're very collaborative. I think one of the most important traits that we look for are people who either who are both good students and great teachers. Um a lot of our people on the team were TAs in college and and loved the experience of of sharing knowledge in this way. uh we we do whiteboard sessions all the time and I think the collegial environment where everyone has something to teach and something to learn is is extremely important for us. What are the attributes of people that you would want to hire that you think will be resilient to you know the work environment 3 years from now when more stuff is handled by machines? Can you give your assessment of the major labs one by one, but also then the relationship of like closed source as a category to open source and like what you think is happening and will happen I think it's possible. I don't know that the premium for being 3 to six months ahead is going to last that long. I mean, if you look at like enterprise deployments, uh, they don't move at 3 to six month speed. A lot of enterprises are probably still on like 46, Opus 46 or Opus 47. They don't they don't adopt the bleeding edge rapidly. There's a lot of questions that people have around rolling out any change at all. And I think we're just so early in scratching the surface that um I don't think there's any way to call a winner in this race and certainly I don't even think this is a race that can be decided ever. There's always it's a continual process and fundamentally I don't think open source ever goes away. If there's a vacuum because one leader steps out, a new leader will step in. There's too much incentive and too much there's a lot of tailwinds too. It's just it gets easier every day to treat to train a frontier class model. I want abundant tokens and diverse harnesses. I want everyone to build their own harness and and You sit in such a unique seat and you have such a unique perspective on like what you're trying to do to make this feature a reality. What do you think are your most like divergent views of the world versus your friends who are really well informed and interested in this stuff? Like what what make your what ideas of yours make your friends look at you like you have three heads? what's a trillion tokens like ground us in how much that is Yeah. It's millions of dollars. Yeah. I mean, we were asking for at least at least um you know, three to six orders of magnitude improvement in cost per token. Uh get that into 5,000. You probably have some customers. And in fact, I would argue that we're for some size of model, we are approaching a trillion tokens being measured in, you know, tens of thousands of dollars. And that's something that you can imagine running for a single job. I never will believe in that. There is always demand for intelligence in the world. I think that the way in the on-ramps to that intelligence are our challenge as a product u you know community. I'm not a product person so I cannot say I had the best vision. I want to enable those people. I want them to never be held back by the sense that, oh, I my free tier users cannot use or I can't afford to give them this many tokens. And I hear that from my customers all the time. Um, we want to fix that. One of the things I keep coming back to is this question of Nvidia. I am bullish on Nvidia in the short term. And you know, Nvidia, you should never bet against them. They're always going to reinvent themselves. But like fundamentally I think one thing that surprises people is when I tell them that hey if you look at you know Hopper to Blackwell to Reuben and you compare like for like like what is the performance per watt of Bloat 16 multiply it hasn't improved all that much or or even you take that one step further go to TSMC if you look at TSMC 5 nanometer versus four versus three versus two the performance per watt on these chips doesn't change like a dramatic amount Well, we're fully downstream of models, right? So the model people get to decide how to design their their architectures and I have only like very light I mean I don't have any input to open AAI or anthropic but um I can only pray that they go in the direction that is a minimal to me and the ch like or I have to like do my best to predict where I think they're going to go and build my serving architecture accordingly. Both software and hardware choices they have. I think the most interesting game in some ways to play like once again this is going back to like the profoundity of the machine thinking and how consequential it is to decide to use something like sparse attention versus dense attention or um how consequential it is to like use a different data type like we were training in B16 but now we can train in FP8 or FP4 lower precision data types that is just an arbitrary choice it feels like but it has profound implications for what chips I can use and and you know how I should build my hardware think about the future of compute Y What advice would you give them on like how to orient their companies or like the type of company, not the specific choice they're making on a tech tech bed or something like this Why? There's no easy way to bring on a lot more fabs of memory and those guys have been Yeah. Yeah. The boys in Boise don't uh don't love huge capex for for cyclical. But conceivably like because of that shortage, the world is just going to route around it by making everything else in the system more efficient. Nvidia is really smart about this? They don't compete with their customers. Nvidia takes the long view on everything. Um, why don't they even start with the Neocloud? Why don't they just sell computer out the back door? Well, Nvidia is really good. Jensen is really good at making his friends billionaires. He's made Cororeweave a billion dollar company, many billion dollar company. And there's no need for him to kind of uh destroy that goodwill. like he wants to create a diverse community of NeoClouds and inference providers who are all jockeying to create demand for Nvidia such that if any one of them decides to I don't know vertically integrate or go with AMD or any other option he's got three more people ready to hungry to fill that position. My favorite closing question for everyone is what is the kindest thing that anyone's ever done for you? It is such a rare rare trait and um you know it that level of expertise is so noble to chase and I think and that stays with me quite a bit. Not a common but an advisory. Neil, amazing conversation. Thanks so [music] much for your time. You know how small advantages compound over time? That's true in investing and just as true in how you run your company. [music] Your spending system is your capital allocation strategy. Ramp makes it smarter by default. Better data, better decisions, better economics over time. See how at ramp.com/invest. As your business grows, Vanta scales with you, automating compliance and giving you a single source of truth for security and risk. Learn more at vanta.com/invest. [music] The best AI and software companies from OpenAI to cursor to perplexity. Use work OS to become enterprise ready overnight, not in months. Visit works.com [music] to skip the unglamorous infrastructure work and focus on your product. Ridgeline is redefining asset management technology as a true partner, not just a software vendor. They've helped firms 5x and scale, enabling faster growth, smarter operations, and [music] a competitive edge. Visit ridgeland.ai to see what they can unlock for you.

FAQ

なぜ英語が「自動字幕ベース」なの? 公式の文字起こしは?

このページの英語は、YouTube が生成した automatic captions(ASR=音声認識) を yt-dlp で取得したものです。人手で整えた公式トランスクリプトではありません。 公式サイト Colossus(Join Colossus)にはエピソードページと導入文がありますが、全文トランスクリプトはログイン壁の向こうです。 参照: https://colossus.com/episode/from-transistor-to-token/

自動字幕の精度はどれくらい? どこが壊れやすい?

一般に YouTube 自動字幕の精度は条件で大きく変わります。 - 条件が良いと 85–95% 前後まで上がる例もある - アクセント・早口・専門用語・固有名詞・複数話者 で一気に落ちる このエピソードは技術インタビューで、チップ名・モデル名・並列化手法が多く、名前の誤変換が起きやすいタイプです。

この記事の英語で実際に見つかった/起きやすい誤変換は?

自動字幕原文に出やすい典型例です(日本語本文側では多くを修正済み)。

| 字幕に出た/出やすい表記 | おそらく正しい表記 | |---|---| | Sal research / sale | Sail Research | | Cerebrus | Cerebras | | Grock(チップの話) | Groq(イーロンの Grok とは別) | | base 10 | Baseten | | Triau | Tri Dao(FlashAttention) | | Opus 45 | Opus 4.5 | | Kimmy | Kimi | | SKH Highix | SK hynix | | longunning | long-running | | KB cache | KV cache | | Enthropic | Anthropic | | cloud code | Claude Code | | codeex | Codex | | sandboxes / sale boxes | sandboxes(長時間エージェント用VM) | | Fable / Mythos | 字幕の固有名。製品の一次確認は動画側 |

つまり英語欄は「音声の生ログ」に近く、学習用の対照テキストとしては有用でも、引用の一次ソースとしては動画そのものを優先してください。

日本語と英語は1対1で完全対応している?

完全対応ではありません。 - 日本語は読みやすさのため発言をまとめ・整文している - 英語は自動字幕の切れ目が多く、発言数が日本語より細かい場合がある - ページでは話者(パトリック/ニール)を手がかりに、順番で英語を紐づけている そのため「この日本語のこの1文 = この英語のこの1文」とは限らず、だいたい同じ話題の塊として読むのが正しい使い方です。

上部トグル(英語原文 / 少し大きく)は何のため?

- 英語原文 ON/OFF … まず日本語だけ読みたい人向け。英語は約9.5pxの補助情報 - 英語を少し大きく … 約11px前後。対照しながら精読したいとき用 英語をメインにすると誤変換に引きずられやすいので、デフォルトは「日本語メイン・英語は極小」にしています。

B. 人物・会社・番組

この回は Senra Show ではない?

はい。ホストは Patrick O'Shaughnessy(パトリック・オショネシー)。番組は Invest Like The Best(Colossus / EP.488)。原題は YouTube では *Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper*、公式ページでは *From Transistor to Token*。 公開: 2026-08-25、約1時間23分。 動画: https://www.youtube.com/watch?v=uyzqxIoiobU

Neil Movva / Sail Research とは?

Neil Movva(ニール・モヴァ) は元 NVIDIA エンジニアで、テンソルコア初期(2015–16頃)にGPUカーネル側で働いた、と本編で語る。Apple でもチップ関連に関わった、という報道もある。 Sail Research(セイル) はオープンソースLLMの推論(inference)をAPIで提供する会社。ニールはこれを トークン工場(token factory) と呼ぶ。長時間エージェント向けのクラウドVM(sandboxes)もホストする。 報道ベースでは、2026年にステルス解除、シード+シリーズAで約 8000万ドル、評価額約 4.5億ドル。Kleiner Perkins がシリーズAをリード、Sequoia / Redpoint 等も参加、とされる。金額はIRではなく報道・口頭ベース。

Patrick O'Shaughnessy とは?

資産運用会社 Positive Sum のCEOで、Invest Like The Best のホスト。創業者・投資家インタビューが中心。

C. 本編の核になる技術・戦略

トークン工場 / 1000倍安くなる、とは?

ニールのノーススターは「業界で最も安いトークン単価」。10倍安くなると新しい製品カテゴリになる、という考え方で、トークンについてもそれを狙う。1000倍は会社の北極星であり、達成済みの監査数字ではない。 本編後半では、OpenAI価格で 1兆トークン ≒ 少なくとも500万ドル、それを 5000ドル まで持っていく(3〜6桁の改善)という言い方も出る。口頭のオーダー感。

低レイテンシ(チャット)vs スループット(バックグラウンド)

過去2年の推論インフラは、対話型チャットボット向けに 低レイテンシ(すぐ答えが返る)を最適化してきた。その方向に強く引っ張った顧客例として Cursor が挙がる。競合例として Baseten や Fireworks。 ニールの主張は、これから伸びるのは 何時間〜何日も裏で動くエージェント。待たされる人間がいなければ、100 tok/s である必要はなく、10 tok/s で十分、そのぶん効率が上がる。 比喩: 「最良のレイテンシは、レイテンシがないこと」。朝起きたら仕事が終わっている。 バス(大量・遅い・みんなを乗せる)対 自家用車(速い・少人数)も、GPUのバッチ処理の比喩。

テスト時計算スケーリング(test-time compute)

エージェントに時間を与えるほど答えが良くなる、という考え。本編では「理論は約2年前、実際に賭けられたのは昨年末の Opus 4.5 あたり」と語られる。口頭では、年末時点でバックグラウンドとリアルタイムが 50/50、数年で 90/10 へ、という見通し。

スピード・オブ・ライト

NVIDIA社内の文化として語られる言葉。チップが理論上出せる上限性能まで詰める、という意味。相対的な競合比較より、絶対値でピークを取りにいく。

NVLink / テンソル並列

NVIDIAのGPU間インターコネクト。大きな行列積を複数GPUに分割して 最小レイテンシ を下げるのに効く。ニールの立場では、低レイテンシ推論にはほぼ必須だが、自分たちはそこを最優先しない。代わりにエキスパート並列やパイプライン並列など、インターコネクトが弱いチップでも回る方式を探す。 「悪いチップはない。悪い価格があるだけ」がハードウェア側のスカベンジャー戦略。

SRAM vs DRAM / セレブラス / Groq / KVキャッシュ

- SRAM: ロジックダイ上の高速メモリ。近いので帯域は極端に速いが、密度が低い。 - DRAM / HBM: Micron / SK hynix / Samsung が作る高密度メモリ。Blackwell 級だとロジック周囲に百GB超(本編では 288GB 級の話)。 セレブラスはウェハスケールでSRAMを積み、毎秒ペタバイト級の帯域で高tok/sを出す、という説明。弱点として、会話ごとに膨らむ KVキャッシュ(動的メモリ。重みより大きくなり得る)を挙げ、将来は「MLPを高速メモリチップ、Attentionを従来GPU」みたいなハイブリッド、という予測。

トランスフォーマーの原罪

Attention(メモリ律速)と MLP(計算律速)を隣に置いてしまったこと。1種類のチップで両方の最適は難しい、という整理。

データの未来

インターネットは「一度きりの補助金」。高品質テキスト約 30兆トークン(広く取ると300兆)をモデルはもう何度も見ている。次は人間のランダムな好みフィードバックではなく、検証可能な課題(コーディング・数学など)の RL環境(ジム) で自己改善する、という話。

スカベンジャー戦略(電力・データセンター)

大手が欲しがらないものを拾う。 - チップ: NVIDIA以外(AMD、評価対象として他社チップ)を価格次第で買う - 電力・拠点: 1MW級(液冷ラック約8本、冷蔵庫サイズ、という口頭説明)でも買う。稼働率 95% でも、制御面でワークロードを移せれば致命傷ではない - 太陽光・風力の断続も、天気を読んで事前に移せるなら利点になり得る

オープン vs クローズド / 蒸留

顧客がモデルの主権(weights を持ち、誰にも取り上げられない)を欲しがるのが追い風。インターネット上の成果物の多くがAI由来になり、GitHub等に上がれば 意図せず蒸留が起きる、完全に止めるのは難しい、という見立て。

地政学・TSMCへのコントラリアン

「TSMCが止まったら世界が終わる」言説への反論として、西側(Intel等)は最悪でも性能/ワットで数倍悪い程度、5nm→3nm→2nmの伸びも劇的ではない、という口頭。これは話者の見解であり、コンセンサスではない。

1兆トークンの需要はあるのか

「知能への需要は常にある」。無料枠ユーザーにトークンを出せない、がボトルネックになってほしくない、というのがニールの希望。人がチャットで消費できる量には上限があるが、バックグラウンドエージェントにはその上限がない。

D. このHTMLの使い方・CRIT

CRIT で質問するには?

見た目用は HTML preview、行コメント向きは article-bilingual.md です。 - わかりにくい段落を選んでコメント/質問 - エージェントが返信できる 「この用語がわからない」「英語と日本語が食い違って見える」なども歓迎です。FAQに足りない点があれば追記できます。

引用や学習に使うときの注意は?

一次情報は動画(YouTube / Colossus のエピソードページ)です。このHTMLは学習用の再構成です。 - 固有名詞・数字は話者の口頭ベース - 英語欄は自動字幕由来の誤変換を含む - 日本語は整文・意訳を含む(逐語ではない) - 調達額・評価額は報道ベース 公の場で引用するなら、可能ならタイムスタンプ付きで動画を当たってください。