What rebuilding AlphaGo teaches us about self-play, RL, and future of LLMs - Eric Jang
What rebuilding AlphaGo teaches us about self-play, RL, and future of LLMs - Eric Jang
Summary
Eric Jang (formerly VP of AI at 1X Technologies, before that at Google) rebuilds AlphaGo from scratch on a $10,000 Prime Intellect grant and walks Dwarkesh through every primitive of intelligence the original system embodied: Monte Carlo Tree Search, the policy/value network, self-play, and the way these three components interact. The pedagogical arc starts with rules of Go, builds up MCTS with PUCT scoring (a variant of UCB that adds a policy prior), explains how a neural network is added to short-circuit deep tree search by giving a value estimate at the leaves and a policy prior at the branches, and culminates in self-play: the network plays itself, MCTS produces “improved” labels by searching beyond the network’s own intuition, and the network is retrained on those labels.
The middle of the conversation is the part that matters most for current AI: why does MCTS work so well for Go but not (yet) for LLMs? Jang’s answer hinges on credit assignment and what he calls the “supervised learning signal” property of AlphaGo. In AlphaGo, every training step is just supervised learning on improved labels — the MCTS policy is strictly better than the network policy at any state, so the network just hill-climbs forever on a beautiful signal it can never run out of. In LLM RL, by contrast, you only get a sparse reward at the end of a long trajectory, and naive policy gradient has to figure out which of the 100,000+ tokens in the trajectory actually got you the right answer. The variance is enormous; the information rate is much lower than the supervised learning case. Jang argues policy gradient RL is roughly bits-per-flop limited, and AlphaGo sidesteps this because MCTS gives strictly better targets every move.
The final stretch covers Jang’s experience using Cursor + Claude Opus in an “autoresearch” loop to do the research itself — a peek at what it’s like to use frontier LLMs as research engineers. He found Claude excellent at executing experiments, optimizing hyperparameters, and grinding metrics, but still bad at choosing what to investigate next or escaping research dead ends. He has a Claude skill called “experiment runner.” He notes that he caught the model rabbit-holing on dead-end ideas (like his own off-policy MCTS relabeling experiment) and couldn’t get it to step back and reflect laterally. This shapes his answer to the intelligence-explosion question: yes, the inner loop is automatable; the outer loop — knowing which problem to attack next — is still the bottleneck.
Highlights
”AlphaGo’s core conceptual breakthrough was using neural nets to make this search problem tractable”
“Computer scientists for many years thought that Go was not a tractable problem this century because the branching factor and the depth of the tree are just too large. So AlphaGo’s kind of core conceptual breakthrough was using neural nets to make this search problem tractable.” — Eric Jang, 12:02
Clip command
yt-dlp --download-sections "*12:02-13:00" "https://www.youtube.com/watch?v=X_ZVSPcZhtw" --force-keyframes-at-cuts --merge-output-format mp4 -o "alphago-tractable.mp4"
”The maximum power of a neural network is at the edge of chaos”
“You have the maximum power of a neural network at the edge of chaos.” — Eric Jang, 1:23:41
Clip command
yt-dlp --download-sections "*1:22:30-1:24:30" "https://www.youtube.com/watch?v=X_ZVSPcZhtw" --force-keyframes-at-cuts --merge-output-format mp4 -o "edge-of-chaos.mp4"
”$10K to rebuild AlphaGo”
“I got a donation from Prime Intellect for about 10k and then I spent maybe the first 4k doing exploratory research and then about 3k on the final run.” — Eric Jang, on the total cost of rebuilding AlphaGo in 2026, 1:55:40
Clip command
yt-dlp --download-sections "*1:55:10-1:57:00" "https://www.youtube.com/watch?v=X_ZVSPcZhtw" --force-keyframes-at-cuts --merge-output-format mp4 -o "10k-alphago.mp4"
”Why AlphaGo is so beautiful: you never start at 0% success rate”
“The major reason is that you never have to initialize at a zero percent success rate and solve the exploration problem of how to get a non-zero success rate… It’s just supervised learning on a value classification as well as a policy KL minimization. So it’s just a supervised learning problem on improved labels. The training is very stable — you can train as big a network as you want.” — Eric Jang, on why MCTS-based self-play is more sample-efficient than naive RL, 2:19:58
Clip command
yt-dlp --download-sections "*2:19:46-2:22:00" "https://www.youtube.com/watch?v=X_ZVSPcZhtw" --force-keyframes-at-cuts --merge-output-format mp4 -o "alphago-beautiful-supervised.mp4"
”In many RL problems you spend all the time at zero”
“Once you’re here, you have something, but you actually in many RL problems spend all the time here. So there’s a sort of question of like how do you initialize so you’re at least not at zero but at a non-zero pass rate.” — Eric Jang, on the brutal exploration problem in policy-gradient RL, 2:18:13
Clip command
yt-dlp --download-sections "*2:17:02-2:19:10" "https://www.youtube.com/watch?v=X_ZVSPcZhtw" --force-keyframes-at-cuts --merge-output-format mp4 -o "rl-stuck-at-zero.mp4"
”LLMs are bad at lateral thinking when stuck in a research dead end”
“Occasionally I’ll rabbit-hole down a track like this off-policy MCTS relabeling, do a few experiments and then realize it’s a dead end. And they don’t seem to be able to kind of step back and do the lateral thinking of like, wait a minute, this track doesn’t really matter.” — Eric Jang, on Claude Opus as a research engineer, 2:25:08
Clip command
yt-dlp --download-sections "*2:25:08-2:26:30" "https://www.youtube.com/watch?v=X_ZVSPcZhtw" --force-keyframes-at-cuts --merge-output-format mp4 -o "lateral-thinking-llms.mp4"
Key Points
- Go’s branching factor (~361) and depth make brute-force search intractable (12:02) - This is why computer scientists thought Go was unsolvable for the century
- Trump-Taylor scoring rules are unambiguous for computers; Chinese rules require value-function agreement between humans (3:00) - The two scoring philosophies treat dead-stone resolution differently
- MCTS data structure: each node tracks N (visits), Q_A (mean action value), P_A (prior probability) (14:13) - These three quantities are what gets propagated
- PUCT vs UCB (16:31) - PUCT adds a policy prior P_A; UCB doesn’t. Both have an exploration bonus rewarding unvisited actions
- The neural network’s job: predict policy (good moves) + value (will I win?) (2:14) - These short-circuit the deep tree search at leaves and branches
- ResNets still outperform transformers on Go due to local inductive bias (33:00) - Jang has tried hard; KataGo’s global feature aggregation tricks help
- AlphaGo Lee: initialized with supervised learning from expert human play (39:06) - AlphaGo Zero famously skipped this
- The Lorenz attractor analogy for Go (1:20:44) - Both are deterministic, both are sensitive to initial conditions, but the macroscopic outcome (who wins) is predictable
- The cipher / neural network parallel (1:23:02) - Both rely on repeated mixing of information for power
- Replay buffer requires being on-policy to avoid DAgger-style compounding errors (2:01:26) - Why on-policy vs off-policy RL matters
- Policy gradient RL has bits-per-flop limit (2:12:00) - Bits per flop = samples per flop × bits per sample. Sparse rewards = low bits per sample
- The “stuck at zero” failure mode of policy gradient RL (2:18:13) - If your policy never samples the success path, you never get gradient signal. Initialization is everything
- AlphaGo is a stable supervised learning algorithm (2:19:58) - No TD error, no dynamic programming — just KL minimization to improved MCTS labels
- Why MCTS doesn’t work for LLMs (1:47:00) - In Go the MCTS policy is strictly better than the raw network. In LLMs, getting strictly better targets without a value function is much harder; reasoning models do something search-like without explicit trees
- The CPUCT square-root term breaks for ~100K-vocab LLMs (1:48:56) - sqrt(N)/(1+N_a) explodes with huge action spaces
- Self-play matchups have very high variance (1:26:13) - 51-49 wins in 100 games could just be noise; advantage estimation is the fix
- The exploration problem is the bottleneck (2:18:18) - How do you initialize at non-zero pass rate? This is where supervised pretraining matters
- Total cost: $10K from Prime Intellect (1:55:40) - The compute required to catch up is much smaller than the compute required to be first
- AI as a research engineer: great at experiments, weak at hypothesis-jumping (2:25:08) - The thing Jang couldn’t get Claude to do is recognize when a line of research is a dead end and pivot
- Outer-loop research taste vs inner-loop research execution (2:30:18) - Inner loop is automatable; outer loop (choosing the right question) still seems uniquely human-shaped
- Bitter lesson taste: knowing how to throw compute at the problem (2:30:57) - The research-taste skill for executing on the bitter lesson is non-trivial
- Off-policy Q-target relabeling in hindsight (the “daydreaming” analogy) (2:09:01) - Re-evaluating old trajectories under a better policy is computationally expensive but valuable
Mentions
Companies
- 1X Technologies (0:00) - Eric Jang’s most recent role (VP of AI)
- Google / DeepMind (0:00) - Where Jang was previously a Senior Research Scientist; original AlphaGo home
- Prime Intellect (1:55:40) - Donated the ~$10K compute that funded Jang’s rebuild
- Cursor (0:30) - Sponsor; used to build the flashcards pipeline and run the autoresearch loop
- Jane Street (0:55) - Sponsor; data center deep-dive tour with Ron Minsky
- Anthropic / Claude (2:22:54) - Opus & 3.5 used as the research engineer
Products & Technologies
- AlphaGo (Lee version) (39:06) - The supervised-initialized network that beat Lee Sedol
- AlphaGo Zero (2:00:24) - The tabula rasa version; spent ~30 hours catching up to supervised baseline
- AlphaZero / MuZero (1:49:58) - Successor architectures
- KataGo (33:22) - The open-source successor that introduced global feature pooling
- MCTS (Monte Carlo Tree Search) (8:16) - The core search algorithm
- PUCT (16:31) - Variant of UCB with policy prior; AlphaGo’s exploration formula
- UCB (Upper Confidence Bound) (16:31) - Classic bandit-arm exploration rule
- PPO / VMPO (1:42:03) - Model-free policy-gradient algorithms
- TD learning / Q-learning (2:09:00) - Temporal difference value propagation
- DAgger (Dataset Aggregation) (2:01:26) - Imitation learning with on-policy correction
- ResNet vs Transformer (33:00) - The local-vs-global inductive bias choice
- Trump-Taylor / Chinese scoring rules (3:00) - The two Go scoring philosophies
- Lorenz attractor (1:21:51) - Used as analogy for Go’s deterministic chaos
- Cipher / avalanche property (2:14:00) - Drawing parallel between cryptographic mixing and neural network forward passes
- Reiner Pope’s “As Rocks Make Things” blog post (2:36:24) - Recommended further reading on the cryptography/NN parallel
People
- Eric Jang (0:00) - Former VP of AI at 1X, ex-Google; built this AlphaGo-from-scratch project
- Andy Jones (1:51:03) - 2021 paper on test-time scaling laws Dwarkesh cites
- Reiner Pope (2:36:24) - Referenced multiple times for the cryptography / neural network mixing analogy
- Joshua Stoldikson (1:24:00) - Cited for research on “edge of chaos” neural networks
- Lee Sedol (implicit) - The pro Go player AlphaGo Lee defeated
Surprising Quotes
“AlphaGo’s kind of core conceptual breakthrough was using neural nets to make this search problem tractable.” — Eric Jang, 12:15
“You have the maximum power of a neural network at the edge of chaos.” — Eric Jang, 1:23:41
“The compute required to be the first to do something is always much larger than the compute it takes to catch up.” — Eric Jang, on rebuilding AlphaGo for $10K, 1:56:04
“At the end of the day, you’re just saying like, I have some improved labels, let’s retrain my supervised model on these targets.” — Eric Jang, summarizing why AlphaGo is so stable, 2:21:08
“Automated scientific research is one of the most exciting skills that the frontier labs are working on right now.” — Eric Jang, 2:22:32
Transcript
Dwarkesh Patel: 0:00 Today I’m here with Eric Jang, who was most recently Vice President of AI at 1X Technologies. Before that, Senior Research Scientist at Google.
Eric Jang: 0:34 I like making things, and AlphaGo and Go AI is one of those things that really got me into the field.
Dwarkesh Patel: 1:38 If you plot out how much compute it took to build various iterations of strong Go bots over the years, you can see this is the most compute-efficient way.
Dwarkesh Patel: 2:23 I guess we should first discuss how Go works.
Eric Jang: 2:25 How does the game work? Go is a very simple one that can be implemented quickly and easily in companies. There’s what’s called Trump-Taylor rules. Trump-Taylor rules are designed to be completely unambiguous for Go.
Dwarkesh Patel: 3:34 All right. I’m basically playing randomly here but trying to get around your stones.
Eric Jang: 3:42 This move exposes one empty neighbor for your white stone — akin to a check in chess.
Eric Jang: 4:06 This one is surrounded on three sides. You’re at threat of losing that stone.
Dwarkesh Patel: 4:12 Now you can see that I’m starting to pressure you because by putting a stone here.
Eric Jang: 4:23 Yes. If you think through what happens if you respond here, you can search into the future.
Dwarkesh Patel: 4:31 You have a lot of confidence in my abilities but I’m guessing you’d put the black here.
Eric Jang: 4:36 That’s right. I would capture all three of these stones.
Eric Jang: 4:42 In Go it’s actually okay to let an opponent capture some stones if it allows you to position to capture more.
Eric Jang: 5:36 It’s actually black because I have surrounded this whole area.
Dwarkesh Patel: 5:49 When the final score is tallied, would these ones also count as being in…
Eric Jang: 5:55 Great question. This is where different rule sets have different ways of scoring.
Eric Jang: 6:55 Once two humans’ so-called value function agrees on a consensus, then Chinese rules resolve.
Eric Jang: 7:03 In Trump-Taylor scoring, it’s perfectly unambiguous so it can be decided algorithmically by a computer.
Eric Jang: 7:50 So that is a very big difference in how computer Go scores things and how humans score things.
Dwarkesh Patel: 8:01 How does the game end?
Eric Jang: 8:02 The game ends when either a player chooses to resign or both players pass consecutively.
Dwarkesh Patel: 8:11 Cool. So that’s the rules. Now help me crack this with AI.
Eric Jang: 8:16 Let’s start with intuition about the underlying search process used to make moves.
Eric Jang: 9:29 The AI can choose, let’s just pick three possible random moves, and which move is good is what we want to figure out.
Eric Jang: 10:31 Like 300 depth of the tree.
Eric Jang: 10:35 If we keep expanding possible moves, in this move the AI is going and then here the human would go.
Dwarkesh Patel: 11:22 What do you mean by merging of children?
Eric Jang: 11:24 Both arrived at the same spot but through different paths. So this child node can be thought of as one node.
Dwarkesh Patel: 11:51 And it starts at 361 but decreases by one each time.
Eric Jang: 11:56 The branching factor decreases by one each time. Yes.
Dwarkesh Patel: 11:59 But this is a very, very, very large tree.
Eric Jang: 12:02 Yes. This is also why computer scientists for many years thought that Go was not a tractable problem this century because the branching factor and depth of the tree are just too large.
Eric Jang: 12:15 So AlphaGo’s kind of core conceptual breakthrough was using neural nets to make this search problem tractable.
Dwarkesh Patel: 12:17 Go is actually a deterministic game.
Eric Jang: 14:13 We’ll call this an action. So one thing easy to trip on is if you come from robotics, the state we result based on the action.
Eric Jang: 14:59 Q represents the mean action value of this action.
Eric Jang: 15:00 I’ll use a subscript A to denote this corresponds to taking a specific action.
Eric Jang: 15:16 We’re going to also store the probability of taking this action.
Eric Jang: 15:46 This is the basic data structure to implement a tree. In AlphaGo, they use a slightly different action selection criteria called PUCT.
Eric Jang: 16:31 The equation forms are actually pretty similar. These are both scoring criteria.
Eric Jang: 17:11 In both UCB and PUCT, there’s this term that basically rewards taking actions you haven’t taken before.
Dwarkesh Patel: 17:54 Just to make sure I’m understanding it, maybe I can put it in my own words.
Dwarkesh Patel: 18:00 Focus on UCB. What we’re saying here, you can think of it conceptually as two different things: the Q term and this exploration bonus.
Eric Jang: 18:39 Correct.
Dwarkesh Patel: 18:41 And the Q is basically representing: what is the probability that I’ll win this game from here?
Eric Jang: 19:40 Yes. The motivation for UCB was an algorithm where if you don’t know the payoff of the arms, you can still get a regret bound.
Eric Jang: 20:15 One small clarification: we talked a little bit about simulations and probabilities and so forth.
Eric Jang: 21:00 What is the expected action value under the random distribution induced by some random search process?
Eric Jang: 21:12 That’s where P of action comes in. If we assume a very naive uniform random search, this is just a valid integral you can take but it’s very slow.
Eric Jang: 22:33 We can assign a value U to a terminal leaf node, but how do we assign values for intermediate nodes?
Eric Jang: 23:23 The weighted average could be dependent on the sampling distribution.
Eric Jang: 24:00 And only a few actions give you high value, so the search in practice is still very expensive.
Dwarkesh Patel: 24:22 Your explanation about the states where it’s obvious to a human who’s going to win, but not obvious to the computer.
Eric Jang: 24:49 We talked about this U value being your final resolution of whether you won or lost.
Eric Jang: 25:45 This is remarkable. If you think about the beauty of something like this — it’s a neural network in a tree.
Dwarkesh Patel: 26:30 Makes sense.
Eric Jang: 27:00 The important intuition at a high level: classically for game-tree search the depth was the bottleneck.
Dwarkesh Patel: 27:44 So we take this idea that humans can glance at a board and instantly predict whether we win.
Eric Jang: 28:17 Let’s go back to how this play-out works. We’ve only talked about making one move.
Dwarkesh Patel: 29:44 So let’s talk about the neural network part of this.
Dwarkesh Patel: 30:00 Once a move is made, a human makes a move, then the AI looks at this and runs simulations to figure out the best move.
Eric Jang: 30:30 Correct.
Dwarkesh Patel: 30:41 You discard all of that, then the next player makes a move.
Eric Jang: 30:46 One small addendum: you don’t discard all of that, you keep one thing behind.
Dwarkesh Patel: 30:47 [Cursor sponsor segment - flashcards pipeline]
Eric Jang: 31:53 Now we have a basic intuition of how moves are made with search. We’re going to talk about how neural networks can speed this up.
Eric Jang: 32:23 I’m going to draw a one-dimensional flattened move distribution, but this is really like a square grid.
Eric Jang: 33:00 ResNets work for small data regimes — my experience is that ResNets still kind of outperform transformers for Go.
Dwarkesh Patel: 33:10 Oh really, why is that?
Eric Jang: 33:11 They provide the inductive bias of local convolutions. Transformers start to outperform residual convolutions with bigger data.
Eric Jang: 33:22 One interesting finding from the KataGo paper was that they found it quite useful to pool together global features.
Dwarkesh Patel: 33:40 What does it mean to aggregate global features?
Eric Jang: 33:44 If you have a 19 by 19 Go board with battles going on — you need information from far away to predict the local outcome.
Eric Jang: 34:52 I’ve tried very hard to make transformers work for this problem because I was curious if transformers would present some advantage.
Eric Jang: 35:58 Go is a perfect information game and in perfect information games, there does exist a Nash equilibrium.
Eric Jang: 36:25 That is the design choice that most Go agents, AlphaGo chose to do, which in hindsight turned out to work very well.
Eric Jang: 37:01 Returning back to the neural network, the architecture is not super important. You can get it to work with transformers.
Eric Jang: 39:00 The outcomes of games given the board state, and we’re also going to train this to predict good moves.
Eric Jang: 39:06 The OG AlphaGo paper, called AlphaGo Lee, initialized this network with a supervised learning data set of expert human play.
Dwarkesh Patel: 41:25 I didn’t understand the significance of this way of thinking about values to expert data.
Eric Jang: 41:32 It is not relevant to the expert data. It’s true for any data you train on.
Eric Jang: 1:18:00 So this is sort of similar to what we covered earlier on Q learning.
Eric Jang: 1:19:02 The question we should be asking ourselves is, we’ve been formulating solutions to NP-hard problems.
Dwarkesh Patel: 1:20:08 I feel totally not qualified on the computational complexity to comment on this, but I wonder if the broader macroscopic outcome is what we care about.
Eric Jang: 1:20:44 Here’s an analogy to weather that might be relevant in Go. The problem of “here is our current board state, what is the exact board state in the future” is extremely sensitive to initial conditions.
Eric Jang: 1:21:21 This captures a lot of possibilities. There’s this more macroscopic quantity that we really care about.
Eric Jang: 1:21:51 If you start anywhere on the Lorenz attractor, you don’t know where you’re going to end up, but you do know that the system stays on the attractor.
Dwarkesh Patel: 1:22:15 Contrast that to something like a hash function which is incredibly dependent on initial conditions but doesn’t have a macroscopic pattern.
Eric Jang: 1:22:30 Intuitively that seems correct.
Eric Jang: 1:22:47 If they were able to do that, then you could prove P is not equal to NP.
Dwarkesh Patel: 1:22:51 In fact, we know there is structure in many cryptographic protocols.
Dwarkesh Patel: 1:23:02 Reiner has a very interesting blog post where he talks about how if you look at a high level neural network forward pass, it looks very similar to a block cipher.
Eric Jang: 1:23:41 You have the maximum power of a neural network at the edge of chaos.
Eric Jang: 1:24:00 I think there’s some research papers from Joshua Stoldikson on this.
Dwarkesh Patel: 1:24:04 Yeah, there’s something quite fundamental about chaos that is not just hopeless noise, it’s just complex structure.
Eric Jang: 1:24:31 It is crucially not saying we’re going to increase the probability of winning directly.
Eric Jang: 1:25:33 What are some alternatives we could do to train self-play?
Eric Jang: 1:26:13 What ends up happening is, let’s say you have a chain of actions that led to a win and you have a matchup between two agents.
Eric Jang: 1:26:34 Let’s say you play a hundred games and each game lasts 300 moves. You’re doing policy gradient.
Eric Jang: 1:27:00 So let’s say 51 games policy A wins and then 49 games policy B wins. This is just due to random luck.
Eric Jang: 1:30:00 Whether you won or lost. In the case where you lost, you just don’t train — your gradient is zero.
Eric Jang: 1:30:18 The trouble is this is very high variance.
Eric Jang: 1:31:14 Let’s actually map this to an LLM case, and we can answer why LLMs only do one-step RL instead of multi-step RL.
Eric Jang: 1:32:01 Log probability of the whole sequence is equal to the sum of probabilities of individual tokens.
Eric Jang: 1:34:29 This is sort of the very basic form, but this is still a contributor to variance.
Dwarkesh Patel: 1:34:59 If you applied that, the only thing you could do is eliminate 49 of the games.
Eric Jang: 1:35:14 Actually the optimal case is to discard all of these moves and only get a gradient on that single move that you got better at.
Dwarkesh Patel: 1:35:22 But how would you do that?
Eric Jang: 1:35:25 This is a pretty tricky problem in practice. This is where advantage estimation happens in reinforcement learning.
Eric Jang: 1:36:10 This is where in RL people use things like TD learning to better approximate the quality function.
Eric Jang: 1:36:24 Ideally what you want to do in RL is push up the actions that make you better than the average and push down the ones below.
Eric Jang: 1:37:07 This model-free RL setting is trying to solve a credit assignment problem.
Eric Jang: 1:37:41 What happens if you don’t have the ability to easily search a tree? In Go it’s a perfectly observable game.
Eric Jang: 1:39:00 Assuming we have a good value function, the search will give us a better result than our initial guess.
Eric Jang: 1:39:18 You fix your opponent. Let’s say you’re currently training Pi A against a strong opponent Pi B.
Eric Jang: 1:39:33 You treat this as a classic model-free RL algorithm where your goal is just to beat this guy.
Eric Jang: 1:40:25 Once you have a good policy that you trained with PPO or SAC.
Eric Jang: 1:41:22 Instead of MCTS teacher, you use the model-free RL.
Dwarkesh Patel: 1:41:55 Just to make sure — if you win a game against this other policy, you reinforce all the actions on that trajectory.
Eric Jang: 1:42:00 Reinforce all the actions.
Eric Jang: 1:42:03 Here you can use a number of algorithms like PPO, VMPO, Q-learning.
Eric Jang: 1:42:17 But there is an interesting connection from MCTS and Q-learning.
Eric Jang: 1:42:46 In model-free algorithms, there’s often a component of estimating a Q-value.
Eric Jang: 1:43:13 Q(S, A) is backed up as r plus some discount factor times the max over Q of your next step.
Eric Jang: 1:43:41 The best action you can take at this state is equal to the reward you take.
Eric Jang: 1:43:58 Once I know the Q-value of this action, I can use that to compute the optimal policy.
Dwarkesh Patel: 1:44:08 So when earlier I was like “Hey, why are we training policy, why don’t we just train value alone?” That is what this is.
Eric Jang: 1:44:14 This is an algorithm for recovering value estimates of intermediate steps when you don’t have the ability to do forward search.
Eric Jang: 1:44:30 The intuition is the same — knowing something about the Q-value here can tell you about the Q-value upstream.
Eric Jang: 1:44:51 Q-learning or approximate dynamic programming propagates what you know about the future back to the present.
Eric Jang: 1:45:00 In this case, you’re planning over trajectories your agent hasn’t actually taken.
Dwarkesh Patel: 1:45:46 This is very interesting. And then to unify this with our discussion of LLMs.
Dwarkesh Patel: 1:46:15 Up-weight all the tokens in a trajectory that might or might not have led to a result.
Eric Jang: 1:46:55 There was some research from Google in 2023, 2024, where they tried to apply tree structures to reasoning.
Eric Jang: 1:48:00 The jury is still out on how the final instantiation of reasoning for LLMs would look like.
Dwarkesh Patel: 1:48:08 Don’t LLMs sort of natively learn to do MCTS where they’ll try an approach and be like “oh that doesn’t work, let’s back up”?
Eric Jang: 1:48:22 Certainly I think the LLMs manage to do something that looks like real human reasoning without having to do an explicit tree.
Dwarkesh Patel: 1:48:43 Just to make sure I understand the crux — the breadth from the number of legal actions being wider and the depth different in LLMs.
Eric Jang: 1:48:56 Here’s an example where LLMs break down. The PUCT rule involves square root of N over 1 plus N_A. In an LLM, the vocabulary is huge.
Dwarkesh Patel: 1:49:31 The crux comes down to the fact that in Go you know the MCTS is almost certainly better than your current policy.
Eric Jang: 1:49:58 “No way” is a strong word. Lots of people have thought about how to apply MCTS or its successors like MuZero.
Dwarkesh Patel: 1:51:03 In 2021 Andy [Jones] wrote a paper on test-time scaling laws.
Eric Jang: 1:51:55 Test-time scaling and reasoning and how it interacts with model size are quite profound.
Dwarkesh Patel: 1:53:58 Say more — you’re saying scaling laws did not work or there was no scaling laws pattern you could see in your Go bot?
Eric Jang: 1:54:05 A mistake I made initially when I had bugs around how MCTS labeling was working.
Dwarkesh Patel: 1:55:10 Speaking of compute, you can look at these charts of compute used to train the best AI model in the world over time.
Eric Jang: 1:55:40 I got a donation from Prime Intellect for about 10k and then I spent maybe the first 4k doing exploratory research and then about 3k on the final run.
Dwarkesh Patel: 1:55:57 Is there a sense that they did a bad job training it if you can do it in 10k now?
Eric Jang: 1:56:04 The compute required to be the first to do something is always much larger than the compute it takes to catch up.
Eric Jang: 1:58:18 The first AlphaGo had lots of compute and didn’t need to worry too much about efficiency.
Eric Jang: 2:00:00 The core thing is how can you get as quickly as possible to some strong opponents?
Eric Jang: 2:00:24 AlphaGo Zero — the first 30 hours or so were spent basically catching up to the supervised learning baseline.
Dwarkesh Patel: 2:00:58 Why is it okay to have a replay buffer in AlphaGo? Because every time I talk to an AI researcher, off-policy is bad.
Eric Jang: 2:01:26 This gets into off-policy versus on-policy reinforcement learning.
Eric Jang: 2:02:08 You’re basically supervising them to take good actions on states you would never achieve.
Eric Jang: 2:02:46 In a DAgger-style setup, what your optimal training data looks like.
Eric Jang: 2:03:00 The state and then you win here. These are your optimal policy actions.
Dwarkesh Patel: 2:03:13 But why isn’t this a fully general argument for off-policy training?
Eric Jang: 2:03:19 This is actually why you want to do off-policy training sometimes. You don’t want a compounding error.
Eric Jang: 2:04:13 You always want to be able to correct back to your winning conditions.
Eric Jang: 2:04:31 If we don’t have any of this data and we’re going to just generate from scratch.
Eric Jang: 2:05:08 As part of this project, I tried an experiment where I took a bunch of trajectories and saturated the buffer.
Eric Jang: 2:06:00 These are all off-policy offline trajectories.
Dwarkesh Patel: 2:07:41 Wait, so where is the trainer?
Eric Jang: 2:07:43 The trainer is, you try to minimize Q(s,a) and Q-target.
Dwarkesh Patel: 2:07:46 Can you explain the whole setup again? At a high level.
Eric Jang: 2:07:51 You have your off-policy data that came from various policies. You’re constantly pushing transitions.
Eric Jang: 2:08:32 You send back the Q-target to this transition.
Dwarkesh Patel: 2:08:52 In the background you’re just like, let me think through how valuable were all these actions actually.
Eric Jang: 2:08:56 In a more optimal policy where you’re trying to maximize, what is the Q-target?
Dwarkesh Patel: 2:09:01 It’s like basically daydreaming.
Eric Jang: 2:09:02 Exactly. You’re going back in hindsight and being like “given what I’ve seen now, was this action good?”
Eric Jang: 2:10:13 The trainer would be just predict the MCTS label as possible.
Dwarkesh Patel: 2:10:46 Or sorry why have they converged to that?
Eric Jang: 2:10:47 It’s just more stable.
Eric Jang: 2:10:49 You might use the off-policy Q as a way to do advantage computation.
Dwarkesh Patel: 2:11:22 I’m reminded now of our earlier conversation of why MCTS is so favorable as compared to reinforce or policy gradient.
Eric Jang: 2:12:00 I wrote a blog post a few months ago about how RL, at least policy gradient RL, has a kind of bits-per-flop limit.
Eric Jang: 2:12:54 Bits per flop is samples per flop times bits per sample.
Dwarkesh Patel: 2:15:00 You’d learn negative log p, p being pass rate, bits once you get this label. Whereas in RL, if you’re just randomly guessing.
Eric Jang: 2:15:25 What’s also tough is that the distribution you’re sampling under is your policy’s distribution.
Eric Jang: 2:15:32 If your policy has no chance of sampling blue, then you will never get a signal.
Dwarkesh Patel: 2:17:02 However, the problem is you spend most of training in this regime, the low pass rate regime.
Eric Jang: 2:17:58 Yeah, and arguably…
Dwarkesh Patel: 2:18:00 You spend all your time here, potentially never even getting a single success.
Eric Jang: 2:18:06 Once you’re here, it’s not at all obvious how you get to here.
Eric Jang: 2:18:13 Once you’re here, you have something, but you actually in many RL problems spend all the time here.
Eric Jang: 2:18:18 There’s a sort of question of how do you initialize so you’re at least at a non-zero pass rate.
Eric Jang: 2:18:25 One more thing — there’s the question of how surprising is the label.
Eric Jang: 2:19:02 It would just be the entropy of this distribution.
Eric Jang: 2:19:10 This is also why AlphaGo is quite beautiful. In AlphaGo, you don’t train the policy network to imitate the MCTS action — you do soft KL.
Dwarkesh Patel: 2:19:29 Earlier I was stumbling around — why is this ability to do iterative search where you don’t necessarily need to play out everything so valuable?
Dwarkesh Patel: 2:19:47 I don’t know a formal way to think about this. Why is AlphaGo an elegant RL algorithm?
Eric Jang: 2:19:58 The major reason is that you never have to initialize at a zero percent success rate and solve the exploration problem of how to get a non-zero success rate. This allows you to hill climb this beautiful, supervised learning signal — if you look at the actual implementation of AlphaGo, every step of the way, there’s no TD error learning or dynamic programming, at least explicitly. It’s just supervised learning on a value classification as well as a policy KL minimization. So it’s just a supervised learning problem on improved labels. The training is very stable — you can train as big a network as you want.
Eric Jang: 2:21:00 Everything will just go stably, the infrastructure’s very simple to implement.
Eric Jang: 2:21:08 At the end of the day, you’re just saying like, I have some improved labels, let’s retrain my supervised model on these targets.
Eric Jang: 2:21:15 You’re always in this beautiful regime where you’re just trying to improve the policy, rather than escape this exploration problem.
Eric Jang: 2:21:26 If you draw the win rate of an MCTS policy versus the raw network — you’re never in a situation where the MCTS is giving you no signal.
Dwarkesh Patel: 2:21:55 Okay, that’s a great way to explain it.
Dwarkesh Patel: 2:22:01 Cool. Maybe we sit down and I ask some questions about automated research.
Eric Jang: 2:22:05 Sounds good.
Dwarkesh Patel: 2:22:06 One thing I really wanted to talk to you about is that you did a bunch of the research for this project through this kind of autoresearch loop.
Eric Jang: 2:22:32 Automated scientific research is one of the most exciting skills that the frontier labs are working on right now.
Eric Jang: 2:22:54 I mostly used Opus and 3.5 throughout this work. What works is that the models can do a lot of straightforward execution.
Eric Jang: 2:23:30 The really cool thing automated coding can do now is that it can search a much more open-ended set of options.
Eric Jang: 2:24:00 The ability to just grind a performance metric — this can squeeze out quite a lot of performance.
Eric Jang: 2:24:16 Fantastic now at executing any experiment. I have a Claude skill called experiment runner.
Eric Jang: 2:24:43 That’s what works quite well today. It’s also kind of fragile.
Eric Jang: 2:25:08 Occasionally I’ll rabbit-hole down a track like this off-policy MCTS relabeling, do a few experiments and then realize it’s a dead end. They don’t seem to be able to step back and do the lateral thinking of like “wait a minute, this track doesn’t really matter.”
Eric Jang: 2:25:54 With Mithos class models or Mithos++ models coming online, maybe this just completely changes.
Eric Jang: 2:26:16 One of the motivations for setting up this Go environment was that Go captures a lot of very interesting research dynamics.
Eric Jang: 2:26:37 The inner loop involves research engineering around distributed systems, predicting whether your idea will work.
Dwarkesh Patel: 2:26:56 Or automating AI research.
Eric Jang: 2:26:58 Or automating AI research. Which is the real crux.
Dwarkesh Patel: 2:27:00 The scary/incredible thing of just making AIs making future versions of AIs.
Eric Jang: 2:27:10 There’s a lot of deeper questions one could tackle.
Dwarkesh Patel: 2:27:41 There’s questions on the inner loop and outer loop. On the inner loop, how stackable are these improvements?
Eric Jang: 2:28:46 As in the case of the success story for deep learning, you can think about this as a decades-long idea.
Eric Jang: 2:30:00 At almost any level of the stack they can think locally, as well as step back and think in broad steps.
Dwarkesh Patel: 2:30:18 The other question is how stackable local improvements are in the attempt to get to a better result on the outer loop.
Eric Jang: 2:30:57 The research taste for executing well on the bitter lesson is that you need to know how to scale.
Dwarkesh Patel: 2:32:31 How about the outer loop — how verifiable for making AI smarter?
Eric Jang: 2:33:00 There’s not an automated procedure one can easily imagine of knowing which paper is the scaling laws paper.
Dwarkesh Patel: 2:33:33 We improve on the things we can measure.
Eric Jang: 2:33:51 I’m going to give a non-rigorous argument, but one I intuitively believe — DeepMind has historically done well at this.
Dwarkesh Patel: 2:34:48 I don’t know — isn’t the issue with — hasn’t the issue historically, until Gemini 3 or whatever, been…
Eric Jang: 2:35:14 The jury is still out.
Dwarkesh Patel: 2:35:56 Cool. Okay, we should let people know how they can find out more about this project.
Eric Jang: 2:36:09 My website is evjang.com. There’s a blog post that kind of links to an interactive version of this tutorial.
Dwarkesh Patel: 2:36:24 I also highly recommend people check out this blog post “As Rocks Make Things” which we touched on some of the ideas in this episode.
Eric Jang: 2:36:39 Exactly right.
Dwarkesh Patel: 2:36:41 I highly recommend people check out that blog post as well.
Eric Jang: 2:36:43 I encourage the audience to think about the relationship between thinking and Go.
Dwarkesh Patel: 2:37:15 Cool. Awesome. Eric, thanks for doing this.
Eric Jang: 2:37:17 It’s an honor to be on the podcast.
