Models, Research & Prompt ·
4 hour interview with Yao Sun Woo
English translation of the Chinese original. This version is generated for international readers and may be refined over time.
English translation of the Chinese original. This version is generated for international readers and may be refined over time.
Date: 2026-05-14
I saw Zhang Xiaoxiao's interview with Yao Sun-woo's podcast, which yielded a great deal.
4 Hours of time, and there are not many who can listen well now.
So I put the point out.
It has to be said that this type of podcast, though long, does learn a lot, and it is a form that the foreigners like.
Important clarification of the guest's identity
Silicon Valley, AI Circle, ** Two researchers from Qinghua who graduated from the same grade, all called Shunyu Yao**, are often confused in the Chinese media:
| Yo Sun Rain (other) | ** Yao Sun-woo** | |
|---|---|---|
| Undergraduate | Qinghua Yaoban (computer) | Qinghua Physics Department (Kykoban/Physics Division) |
| Doctor. | Princeton(NLP) | Stanford (theory physics) |
| Representative | Rect, Tree of Thoughts, Second Half of AI | Non-Hermitian Skin Effect (non-Eurmy Skin effect), Scramblon Theory |
| Patth's OpenAppendium. |
Photo by Yao Sun Woo, used with permission
- 2015-2019 Undergraduate, Principal Scholarship, Qinghua Physics Department + Ip Sonsat Physics Award
- During the undergraduate period 3 top issues (2 PRL + 1 PRB) The first author, in collaboration with King Qinghua, proposed ** New approaches to the theory of non-Empire systems** 2019-2024 Doctor of Theoretical and Mathematic Physics, Stanford University, mentor ** Douglas Stanford** and ** Stephen Shenker** on quantum field theory and quantum gravitational dynamics
- After briefly joining Berkeley as a doctor. October 2024
- ** 19 September 2025** Separation from Anthropic, ** 29 September** Add Google DeepMind, Senior Staff Research Scientific
- Participation in Gemini 3, Gemini 3 Deep Think, Gemini 3. 1 Pro
Subtitles of "Shunwoo" and "Anthropic" have been translated as "anthropological/humanist/humanist/human factor/ ape science/Antropik" and "Treasure/Telestar" as Gemini.
One, two Shunyu Yao (01: 26)
Yao Sun-woo introduced another Yao Sun-rain: There is some overlap in our main career paths, so it may seem difficult to separate us." He stressed that the greatest difference between the two was: ** The other was in computer science from the beginning, and he was of physical origin, only "in a sense"**.
The two Qinghua undergraduates went to Princeton, and the other to Stanford -- "It's strange that the whole world thinks that Stanford is a CS holy place and Princeton is a physical sanctuary, and we happen to be the opposite."
- The two men met every few weeks in Silicon Valley, mainly by "showing" - walking, eating, playing poker.
- For another Yao Sun Yu, ** "AI into the second half"**, Yao Sun-woo says: I don't know what the first half and the second half mean. I don't get it."
- His own stage definition is: ** "You're starting to be less worried about one thing, and AI can't be sure whether it's clearly defined in itself, which is the biggest change." ** A year ago, Anthropic was concerned that OpenAI's reasoning capacity would not be followed; now, no one in Gemini, Openai, Anthropic is really worried about "not catching up" - ** It's hard to figure out what to do.**
- Models are homogenized and commodified, and the gap on paper (benchmark) has narrowed to 1-2 percentage points, ** "Most of them are noises, not signals",** and the real difference is only in the actual user experience: Claude's tool is the strongest, Codex has recently been calmed down, Gemini's daily reasoning is better and intelligent coding is still being pursued.
# II. Competition and flight (07: 15)
** Product judgement on OpenClaw**:
- "The inner circle is not really tense. It's more tense outside than inside." In his view, when OpenClaw did not prove anything new - Claude 4. 5 Opus released, the tool was already ahead of OpenAI and Gemini 3, except that nobody was packaging it.
- Manus was acquired by Meta (note: the acquisition was cancelled after the cataloguing system) and OpenClaw was acquired by OpenAi, which suggests that the "package layer" is not currently free of the control of Models - ** "not fast enough to escape"**.
- Wrapper has two ways to survive: ** "Being fast enough"** (Cursor playing) - sufficient user intelligence before model companies react and train their own models. He said that Cursor's relationship with Anthropic has reached a very delicate stage, and Cursor's training his Composer, from close partners to rivals. ** "The market is so small that model companies can't look at it"** (Midjourney playing) - a sub-market such as "demeaning Gemini's dignity".
- Asked if Lovett counted, "I think they had a chance."
- ** Projections for 2026**: Models should achieve ** "training with limited context, use with infinite context"** - The model side and you constantly interact with each other to judge, to discard unimportant information, to become a real personal assistant. This year, it will certainly be possible, but there are many technical paths to be tested.
- With regard to Meta's acquisition of Manus: He's "not fully aware" that the best thing to guess is to get a ** strong Asian product team**, "China is more talented at the end of the product than the United States"; but why can't Meta make it? He didn't think clearly.
Three, "Pre-train does not end" (25: 22)
This is one of his most anti-mainstream judgments.
- ** "In the first quarter of 2026, the pace of model improvement has not slowed at all."**
- He refuses to measure growth by benchmark: "Benchmark is defined in [0, 100], the closer 100 growth is certainly slower, but this does not mean that the growth felt by users is slower. From 70% to 75% could be more valuable than 50% to 60%."
- His judgment is based on the researcher's sense of being: ** Models are becoming easier to learn** - one thing that used to take a lot of effort to model the church, and now, as long as the problem is clearly defined + the data/environment are right, the model is almost "automatic".
- ** Pre-training has been getting stronger over the past few months**." A few months ago, a lot of people said that Scaling Law hit the wall. My experience was that I didn't hit the wall.
- Why does anyone think they hit the wall? He gave three possibilities and pointed to the third most common: I think the paradigm itself is over.
- Dismissal of data etc. ** "They have bugs in their own work, but do not realize - I have observed that the vast majority of the `wall crashes' fall into this category." ** A bug is often much more advanced than a fancy skill.
- It's supposed to be ** a mental problem**: if you believe the problem can be solved, you will systematically do a digestion experiment -- "Gemini and Anthropic are doing well in this matter."
- Current main driver: ** Data and calculus** (the two are strongly related). The algorithm is more like a phased leap forward (e. g. Transformer), followed by gradual improvement.
- "In a relatively clear paradigm (pre-train / post-train), the primary driver is data and algorithm. Multi-modular generation algorithms, which are also confiscated, remain ** scientific issues**; however, natural language generation is no longer a scientific issue but engineering issues."
- "If I estimate a time line, there will be progress in the next four months. But no one in the AI field can predict four months from now." Who's excited?" ** People who make products are excited about OpenClaw, people who make models are excited about model progress**. People in Anthropic and Gemini think more: "Ai will replace us soon, what should we do next - not worry about the wall."
Fourth, Coding outbreak (35: 08)
Why is programming the fastest in the year and a half? He sees two main structural advantages:
- reward signal (reward signal) is clearly defined: SWE missions are naturally measurable and a match of input and output is successful.
- ** The data base is naturally available**: GitHub has deposited high-quality volume codes for decades and the environment is easily constructed.
From a product point of view, the code is also unique: ** The code style of a good programmer is highly similar** (simplified, well structured, easy to expand, abstract), so no recommended algorithms like social/play-like algorithms need to be adapted to the taste of each user - which greatly simplifys the product pattern.
- ** More than 90 per cent of his own code output is generated by models** (conservatively estimated, 99 per cent may actually be); but he spends a lot of time reviewing the review code." After AI's support, the most important thing becomes how it is designed, how it is given the right form.
- Asked Google not to use Claude Code: "You almost lost my job on this question -- Google not to use Claude Code. 20-50 times more efficient than a year and a half ago, but he worked longer:" Because there's more ideas to try, it's gonna take hours for colleagues to figure out a file, and now ask Claude or Gemini for five seconds." "Google is no longer coast along." No one touches fish in GenAI unless you lose interest in technology." ** He starts checking mail and night experiments every day at 9 o'clock, arrives at the office at 10 o'clock, does it when he's single until 10 to 11 o'clock, and his wife takes it home.
- Next Coding level flashpoint? ** "If I could see it, I would have gone to start a business long ago."** ( laughs) The market is not big enough for any other direction than programming - AIG markets are limited to "people only 24 hours a day"; the most likely big market candidate is ** interactive education,** but far less than programming.**
- With regard to the future of programmers: AI will eventually replace programmers, but gradually; ** "AI is a highly centralized technology that empowers a few and deprives the majority of their unique value";** The final conclusion of traditional software works may be ** "one in 1, 000 people who do all their work, with 100 times their wages"**." One in a thousand is only a metaphor, or perhaps one in a thousand or a hundred thousand.
- The group of programmers that survived: ** highly skilled (fully unnecessary) ** understand their position in large organizations + have a strong planning capacity** (can distribute complex things into small pieces to different AI).
- IA research is itself a gold rush or a scientific revolution? ** "All". He said that training AI product managers was not likely at this time - because there was no objective criterion for what was a good product, ** the feedback was too vague.
V, Seedance (50: 10)
Evaluation of byte beat Seedance:
- "may put pressure on DeepMind's multi-modular team, but not ** paradigm level**. The byte has been relatively strong in multi-modular generation, mainly with good data and detail."
- The reason for the guess is ** data**, because the multi-modular algorithm level is not fundamentally new; but he "did not work bytes, only guess".
- An evaluation of Wu Yonghui, who jumped from Google to the byte: "I haven't seen his past code submission and lead the project, and he's one of the few I've seen before ** but with a particularly advanced technical capacity,** and I'm not yet in a position to evaluate him."
- The Sino-American model gap: the past year and a half ** are clearly narrowing**, but whether it will disappear completely or even go back, "is a pending issue".
- "China is at a clear disadvantage in real terms, but this disadvantage has given rise to interesting things - ** China Models are very good at distilling from other models**."
#vi, "hard" and "soft" (54. 30)
In response, Dario Amodei recently publicly accused three Chinese companies of distilling their models:
- ** "The distillation itself is an open secret."**
- He's got two types of distillation:
- ** "brute-force distillation" : Take the token created by Claude directly to enforce your model. ** "Commercially immoral, intellectually stupid - it's an admission that you don't even know what you're going to do, you can only imitate others, and look at the benchmark numbers well." -** "smart distillation": using other models as assistants in their own data, or other models as evaluator. ** "Commercial Grey, but technically interesting... Chinese laboratories may be pioneers in the field of multi-agent training: This is the real multi-agent if they integrate many different companies with very different models of language distribution into a unified training system."
- Calling (later decorated): may have been done before the hard steaming of a family" and then gradually softening of the steam; ** "The least steaming is byte, and its model is still very unique."**
- About bean buns:
- "Bean bags are certainly not as smart as Gemini or Claude. But it's really the best voice generation in the world.
- Why don't American companies do that? ** "Data problem + differences among user groups. Americans are more concerned about productivity, and Chinese people have so many "life questions" to ask "peasets." My own life is boring, and there's nothing interesting about life: everyday technical questions for Gemini."**
- Bean bag phone: "It's a good idea, but I don't know how much the technology costs -- ** You can't let the model book you a high iron ticket, and the last one costs more than the ticket itself, which is unacceptable. ** "I don't know.
- Apple AI Strategy:" ** Looks like I don't care, but I care too much, but if I care too much and I can't do it, I'm stupid. Face problem. ** "I don't know.
VII, Robot (1: 04: 07)
- I saw the show late in spring and went to Amazon to search for the price of human robots, "a lot cheaper than I thought," reflecting the advantage of the Chinese hardware chain.
- But on the software side: "The robotic model is still in the age of ** character engineering** - to give a set scene, to do RL optimization for this scene, everyone knows what to do, but it's not broad-based."
- ** "Whether or not there is a generalization capability is actually an AI watershed in many directions." ** It was easy to get a single scene for certainty, but it was possible a decade ago; the language model crossed this threshold only after Transformer/GPT - "all capabilities can be fully enhanced by training at one level". The robot is far from here.
- Visited Google DeepMind's own robotic lab and Physical Information:" Laboratories are much more interesting than language model laboratories -- language model laboratories are like offices, robotic laboratories are really robotics that go to various shelves to pick up things."
- The robot is currently ** not even in the GPT-1 phase , like multi-modular generation, ** has not found scale.
Eight, gambling in Underdog (1: 08: 45) - growth experience
Born in Ningxia, a city born of coal mines, and in Shanghai from primary to high school. Personally, "I always like to do things I'm not good at." ** "I don't know.
** Key Life Choice - High School Choice**: He could have been admitted to the regular classes of the four top schools in Shanghai (Shanghai, China, China, China, China, Japan, Japan) but abandoned in order to enter ** a "slight" high school contest ** - ** "The barefoot is worth a try."**
The physical competition failed to join the National Collective Team (which did not receive the warranty) and later failed to pass the Qinghua examination. But fate changed: ** during the senior summer camp, when he heard that Qinghua had an independent enrolment for Beijing students, he texted Qinghua's admission teacher on the spot** - "Why don't you take the Beijing students' exam and not Shanghai students?" - to get an examination and sign the first grade down, and finally to accept Qinghua.
** "Bolder." If you don't fight, you'll never get it. Even if you fight, you don't get it. But if you don't fight, you won't get it."**
The assessment of parents: "It's good that Chinese parents can get their kids to talk, ** I usually just inform them.** The best thing for my parents is that they choose not to interfere when they cannot understand what I am doing."
Personality: "Don't try to stop me from doing what I want, and I will do what I don't want to do, but you won't do what I don't want." And, "I'm more competitive with myself and I'm less willing to compete with others - of course, if you care, I must be better than you."
#ix, non-Earmi systems and quantum physics (1: 19: 44)
The choice of condensation theory is the arrangement of fate. ** "Two thirds of the students in Kikot never do physics."**
The undergraduate teacher was Zhong Wang** (subtitled "King Zhong"), who was young and had few students. The Dr. Wang's mentor is ** Shoucheng Zhang** (subtitled "Screen/Sun City", Stanford's famous conglomerate physicist, died in 2018). Teacher Wang doesn't talk much, but he's good at seeing things clearly."
** The popular description of the work on the non-Eurmy system** (his own progress note: I don't want to hear to skip):
- Basic assumptions of quantum mechanics: Isolation system evolution is described by Hamiltonian.
- The vast majority of reality is not isolated systems (and environmental exchange of particles/energy), ** the corresponding Hamiltonian is non-Earmi**.
- When they first studied the phenomenon of open quantum systems, they found that ** the results of the analysis of calculations (cyclical boundary conditions) and numerical calculations (open border conditions) were completely incorrect.**
- Later it was discovered that the basic paradigm of the Euphrome system - ** Blohbo hypothesis** - in the non-Emmium system ** completely collapsed**. The energy element of the non-Eurmy system ** will all accumulate at the system boundary** (i. e. later known as Non-Hermitian Skin Effect, non-Eurmy Skin effect).
- They set up a set of frameworks describing the basic indicia and dynamics of the open border non-Ethymic systems - this is the work of ** paradigm shift**.
** Why didn't you do it?**
- "The paradigm shift is hard to catch, it's already caught once and it doesn't want to catch again."
- ** "This is a weakness of humanity - I always want to challenge things I don't know."**
- Now look back, "If we keep doing it, it'll be the most important job in this direction, and I'll be better known, more quoted and better teaching; but scientific careers will become less exciting."
- So the Ph. D. stage went to ** Theoretical High Energy Physics** (Quantum Fields and Quantum Gravity), the two directions were "near connected".
Rethinking the "Challenge Hard": ** "It's a challenge to yourself, it's a self-abuse"**, "It's a psychological problem if a person is abused for the sole purpose of being abused; but it is worth it if it is to gain information, experience and competence."
The greatest gain of science physics: ** "Throught things clearly, read them in depth, and don't believe in pure theory." ** - Because the non-Emi discovery itself resulted from "inconsistency between numerical calculations and theory, the problem was found only through in-depth tracing".
X, high energy physics (1: 36: 27)
Recognition of the Ph. D. phase"** does not contribute to the world**
- ** "High-energy physics has evolved to the point where experiments have not kept pace with theory."** There are no objective criteria for evaluation, depending on the subjective judgement of several seniors in the field". -** "The life of a human being is not long enough to waste time serving the elderly."**
- The most important lesson the doctor learned in five years: ** "Doing things with relative objective evaluation criteria"** or ** "Doing things that affect the world"**.
- Self-evaluation: "In truth, my doctoral thesis does not go well, but it has little impact on the world. I'm personally very unhappy, but I'm not so bad as to let people say I'm lazy. ** You can meet all external expectations, but you can't fool yourself.**"
- Meet the small circle standards = train a model: ** "If you get into that small circle, you know what the evaluation criteria are, it's easy to do it, even if you don't agree with them."**
- Two or three months after the doctor actually left Berkeley. I told them I might have to do it, and they said it wasn't urgent to keep the job."
#xi. Physics and AI (1: 43: 09)
The advantage of the physicist doing AI:
- Hard-skilled help is really rare.
- Real help in ** character/temporal**: exploring the nature, systematization (both experimental and theoretical).
- But ** "This is not unique to physics - it is also true of people with CS, chemical and biological backgrounds."**
- Anthropic, a person of special multi-physical origin. - Two of the co-founders had a physical background, so they recruited them. But by the time I joined, that inertia was over."
** About AI being a black box**:
- ** "Everything is black, even physics." ** We don't know the most micro-level dynamics.
- Language models are not yet understood at the neurosurgery level (except for the Anthropic Interpretability team, which can be done on very small networks).
- But Scaling Law is already the law of experience** - "The boundaries of the law of experience and the law of science are blurred. The law of thermodynamics was initially also an empirical law, and it became a scientific law with the understanding of micromechanical mechanisms. The future, Scaling Law, may also evolve."
- ** The word "intellectual emergence" is not scientific in itself** - "It is more subjective to me. There's only one real quality change: ** technically capable of scale up, fully upgrading all capabilities**. This is my only definition of "emerging."
** Why ultimately choose AI instead of quantum calculations?**
- Both give opportunities to young people, but quantum calculations ** bottlenecks are on the experimental platform** - "That's something I'm not good at, that's not what I'm interested in".
- AI is more like ** "17th century thermodynamics"** - people didn't even know what "hot" was at that time (and believe in flammables), but that did not prevent experimenting, summing up the rules of experience such as the first law, the second law, the Clapeyron equation, and eventually promoting heat machines to change the world.
- ** "Theoretical physics is far from experimental physics, far from theoretical physics to AI. AI, for me, is a numerical experiment -- there's an idea, a design experiment to verify the nature and to do physical numerical calculations.**
- Awe of experimental physics: "Everyone knows how to build an optical platform, someone can get out, someone can't get out for six years - I don't understand, it's quite mysterious."
XII Training Claude 3. 7 and 4. 5 (1: 52: 32) in Anthropic
Induction
- ** August-September 2024,** Contacted through a former colleague Anthropic (the first manager is also a theoretical physical background).
- During the same period, OpenAI and DeepMind - ** "DeepMind was too slow and finally Anthropic."** OpenAI did not find the right place.
- I went through all the self-learning courses before the interview, and handwritten the nanoGPT of Andrej Karpathy.
- Two teams approached him. ** He chose a more uncertain RL orientation.**
The state of Anthropic
- The entire company, 700-800 people, he joined a group of 10-11 people, almost the entire later RL team.
- ** First impression of Anthropic**: "Enforcement is very strong, relative to top-down companies; there is no concealing between people, and the atmosphere is very good - because it is small and everyone knows it."
- ** Anthropic, why top-down? ** Because ** technology decision makers are the co-founders of the company** (Jared Kaplan and Sam McCandlich) and Dario has enough trust in them." Other companies could not -- Ilya was there, and OpenAI might have, but he lost his decision-making power for some reason and left."
- He worked most with Jared Kaplan.
- The group of co-founders of Anthropic ** "No one has ever left"**, "they are a group of people who really fought side by side - the Scaling Law paper, the GPT-3 papers are all co-authors (Jared, Sam, Dario, Tom Brown, Benjamin Mann, etc.)." - This is the basis of mutual trust that many companies cannot afford.
Claude 3. 5 → 3. 6 → 3. 7
- ** "Claude 3. 5new is called 3. 6 because Anthropic had no early ability to produce - both models were called a name (3. 5)** and were then forced to accept an external call 3. 6. So the actual product line is 3. 5 → 3. 5 new (=3. 6) → 3. 7."
- Claude 3 was found on Twitter to be better coded than GPT-4; ** "This is a signal source for Anthropic bets on programming, but it may be randomly tested at first - purely for technical reasons - from the bottom up, then from all-in."**
- 3. 7 is the watershed of post-Anthropic training: formerly post-training is a patch; 3. 7 is not really large-scale RL.
- ** "When I joined, you already knew what to do, but you didn't know exactly how to do it."** August-September 2024, o1 had not been published, only OpenAI had a mysterious project called Strawberry.
- The real secret (the part he can talk about in public): ** "Make things easier than everyone." ** RL's simplest algorithms are politics and there are many complex algorithms that can cause infra-problems; ** How to trade-off these details are the real ones.**
- One of his important observations: ** "Many of them are useless." ** The differences between the sampler and the numerical of the different companies depend on each infra, so "you may not be useful by copying other algorithms -- algorithms are part of the system." And that's why I don't like to answer people's questions about how Anthropic / Gemini does -- answers can mislead them."
3. 7 → 4. 5
- By the time he left, Anthropic had been close to 2, 000 (more than double when he joined). ** "I caught up with the end of a small company" ** - three or four months later, the company suddenly got bigger, the culture started to disarray, and some came from outside to bring conflict with the original culture".
- ** He doesn't like people ** ** ** "I think `ideas are cheap'**. The real hard part is immunisation. I don't like people who spend most of their days in Slack talking about Grand Princes -- nothing.
Reason for separation
- Main cause: ** want to learn something different.**"Anthropic is very focused, only linguistically related, not multi-modular, less low-level engineering and infra -- I want to learn this."
- ** About 40% Reason: I don't agree with Dario's anti-China stance**." As a CEO, he can think of it, but pushing it so extreme is a very emotional reaction.
- 40 per cent is not the primary cause, but it is not irrelevant, let alone ** "the reason for holding shareholders"**. (chuckles)
- Perceptions of the future of Anthropic (when leaving): ** pessimism** - "API sales token is a bad business, price wars come, only Google wins." But then it proved too pessimistic, and Anthropic did very well at the product level.
- Asked if he regretted it: "Not really. My motive is to learn something in a different place."
- The birth of Claude Code: ** "It was almost the rare moment of modern times, and it also showed individual heroism."** Founder ** Boris Cherny** (subtitled "Boris Cherny" ) was just trying to make it work for himself and his colleagues, and eventually became the product." It is likely to be an interactive transformation product at the same level as the tremors."
About "heroes are over"
This is one of the core points of the interview:
** "Personal heroism may have passed in the field of linguistic modelling -- that's after that moment."**
** "Now that we're all surfers, it's essentially surfers, not your surfers."**
** "There are no heroes, sometimes even the old heroes are a little stupid."**
** "My contribution to any model, my statement will always be: I'm not so important to that thing; I'm more fortunate to have the opportunity to join an important project and do something at that time."**
In particular, he pointed out that the programming success of Anthropic was indeed a "corporate heroism" (bold enough to gamble), but every technical detail within the model was collective.
** Criticism of AI Safety** (very sharp):
- Anthropic was created for the purpose of AI safety, but also to train the front model - Anthropic's own explanation was that "the strongest model must be made to move forward the safety agenda". This idea is very naive - ** It now seems impossible. The more likely result is that everyone has a powerful front-line model and nobody can stop anything."
- The true institutional analogy is ** nuclear weapons**: ** multiple holdings, mutual deterrence** ** - "The self-regulation of a company by its own legislation is uncontrollable" - it only regulates itself, but self-regulation is tantamount to unregulated.
- Yes, Anthropic interpretative team: only interesting progress has been made on very thin, small networks, and the practical language model has not yet reached the level of neurosurgery.
#13, "AI is by nature simple" (2: 35: 03)
** Core proposition**: ** "AI is by nature simple". ** (he stressed it was status, not conclusion)
Explanation:
- ** Because you can do experiments**. Compared to physics (energy scales limit experimental data), AI is not bound to do anything - it takes time to expand and prepare infra, but there are no fundamental difficulties.
- ** "AI doesn't feel like hitting a wall, not because the method is exhausted, but because there's too many ideas to try."**
- In the future ** 6-12 months** AI will start ** doing its own experiment** ** ** ** ** ** not just writing code, but ** running experiment ** analysis ** new assumption → design new code ** running new experiment** **, this chain will close.
XIV, training Gemini 3 (2: 41: 10) at Google DeepMind
Reason for joining DeepMind
- Oppose the inertia of "researchers leaving the factory to join the factory" - he went the other way, ** because he wanted to learn more and more.**
- ** "If you really want to put an idea into the final product model, Google may be a very bad place; but if you want to study freedom, a broad vision, there is no better second place in the world than Gemini."**
- At the time of accession (end of September, 2025), Gemini-Gemini 2. 5 made the generation realize that Google is working on it.
- He was dug in for personal contact, two-way choice.
- Why didn't you go to OpenAI? To put it bluntly, there are not many people who can really do things like Gemini, even fewer than Anthropic."** (laughter) Internal political struggles are beginning to emerge.
- XAI: ** "I don't get it."** "The people who were in contact have gone, and I don't know how they are."
Gemini 3 turning point
- ** Gemini 3 and Nano Banana overlap is the real turning point**: Nano Banana has brought many new users to Gemini App, Gemini 3 to keep them. Only Gemini 3 is not enough - when the market share is less than 10%, the models are spread slowly.
- Gemini's current market share may be around ** 20%** (he has not been accurately verified).
- ** "Openai saved Google's life from the outside."** If ChatGPT really swallowed up the search, Google would be finished; but Openai did" to make Google realize what was important, but failed to swallow the search, and to reverse Google.
- Why didn't Chatbot swallow the search? Search has a lot of "very stupid" needs - "I'll just check where to buy rice, where to order it, and I don't want to wait till the chat robot finally gives us a link." The Chatbot form has not reached its end.
- ** "What is the ultimate form of chat robots? After all these years, there's only one chat box. I really feel stupid." ** - ** "A product manager is needed to unlock the entire capacity of the model."** (laughter)
What happened inside Google
- Externally see model performance jump; internal is ** organizational logic begins to be clear**:
- There's a clear framework for the pre-training phase -- who's in charge of which Node is very clear.
- Google ** engineering management capacity is extremely strong**, pre-training has entered Google's comfort zone, can ** controllably know that the next generation will not be bad, or even predict how much.**
- Anthropic walks from the top to the bottom; Google is still relatively from the bottom, but more from the top than from the past.
- ** "Cultural works"** - Large companies and start-ups are different in nature.
- Google Killer: ** "Find a very simple form of product expression that everyone looks the same and then relentlessly crush you at the technical level, you can't compete."** Search is a typical example.
- OpenAI location: ** "No one's position is secure."** Chatbot is the ultimate form of super app? - "I have no rational answer at all, but it feels like it's not over."
- ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** But I really think that bitbot is stupid, and why is this the ultimate form?"**
Google's Hero
- Heroes in the backstage: ** Sergey Brin** ("As a matter of great decision, he will end up on board"). ** Koray Kavukcuoglu** (Google Deepmind CTO / Senior Vice President of Google).
- ** Demis Hassabis** is more scientific (Isomorphic Labs, etc.), and Gemini has seen Koray most frequently on a daily basis.
XV. Technological prediction and organization (3: 01: 28)
pre-training vs post-training
- There is little difference in the nature of pure technology - ** the greatest difference is in the distribution of data**: pre-training ** wide** (no need for Quality particularly high); post-training ** narrow and fine** (quality is extremely demanding).
- Different types of laboratory organization:
- Anthropic / Gemini: pre-train and post-train in two teams.
- OpenAI: More chaos - the first three teams (pre-train, RL "Strawberry", post-train) and their post-train** are themselves a product team**, "The people of the model are also involved in the product".
On the next paradigm shift
- Presumably ** it's not** a paradigm shift, but two things that are particularly valuable to Google: ** Machine Learning Code (ML coding)**: To enable AI to accelerate AI's own research closed loop - Google is the most complete platform for AI research (hardware + connection + model), which is of great value to Google.
- ** Long-term planning / Long-term (long-term)**: everyone feels important.
- To achieve direction:
- Pre-training side: sparse attention (deep attention, DeepSeek and academia are doing it).
- Post-training side: like Cursor ** External context management** (let the model choose which part to keep or throw away).
- ** Both are essentially the same** - KV Cache of the context token is also a weight.
- "** 10, 000 individuals are defined as 10, 000 `world models'**. Gemini's world model is more like end-to-end training (the conditions produce the next moment scenario); it's a different thing to do with Lee flying -- I don't know what their labs are doing."
- Continual learning is no different than long horizon.
- He ** The main focus is on post-training methods** (pre-training is not a formal job).
- ** "Gemini's techniques on long context really surprised me."** (LAUGHING)
AI Talent Scarcity Challenge
- High pay is because people feel scarce, but** "may not be that scarce - it's not difficult to train a person, it's just that you need an environment to do this. In the past, few people had had such opportunities and the market was therefore relatively scarce. On the other hand, it's probably too much for some people, and you love myths in particular."**
He designed the interview
- The candidate is requested ** 24 hours to do a RL project from zero** - choose models, data, algorithms and discuss with him for one hour.
- Two objectives: Seeing candidates ** The ability to work with AI** (the code itself is now no longer scarce)** There is a trap: if life is completely left to AI itself, the discussion will be exposed in an hour**. "**24 hours limit depends on whether he values this opportunity - whether he can stay up late. People who don't care can't survive this 24 hours. ** "I've got some dark little ideas in it."
# Engineering vs science
"Google pre-training has now become ** Project ** - Top-down, clear nodes, assessable. This is Google's strength."
- There is greater uncertainty about post-training, which remains bottom-up and individual attempts at different approaches.
Core principles of the organization
- ** "Standing systems + individual heroes do not shine"** and ** "Can allow individual heroes to shine but the system is fragile"**
- He prefers the former - ** "An example of system instability is OpenAI: One person leaves, the whole structure collapses."**
- To himself: ** "Researchers must be considered as a whole, otherwise not a good researcher. In academia, it's 'one person eats his whole family'; in a company you're responsible to the company - two completely different mentalities.**
- He admits, "I may not be able to pull my face -- since I signed the contract, I don't think it makes any sense not to do it."
TPU vs GPU
- Large-scale commercial deployments ** There are no differences of excellence**. Open source ecology GPU is better, but this is not a bottleneck for mass deployment.
- Different design concepts:
- ** GPU (especially the Hopper generation)**: NVLink bandwidth is extremely high in monopod, but less in pd (8).
- ** TPU**: Abandoning two or two interconnections between cards, using 3D Torus ** to form more cards into one large rack, each linked only to three immediate neighbours. If the compiler/score is well written, ** Total memory capacity is greater and communication bottlenecks are less.
- Shortcomings of TPU: ** Small-scale inflexibility, poor utility**.
Evaluation of xAI
Short, sharp: ** "I don't understand. They've always been quite volatile."**
XVI, the triumph of collectivism (3: 23: 33)
To the new lab tide
- Recently a bunch of new AI laboratories in Silicon Valley:" ** Most new laboratories will collapse. ** "I don't know.
- Thinking Machines keeps coming up with something new; but some new labs -- ** "I have no idea what they're trying to do, and the founders have actually left the field for a long time."**
Sino-American route split
- ** Central America has split.** China's consumption side:
- "China can come up with a very complex and seemingly unnatural product structure, allowing profits to roll snowballs - you can see the video without US$ 0. 2, but with secret advertising, live broadcasting, electricians."
- The American way of playing -- "Property software: I'll write you codes, 150 costs, 200 sold to you, 50, that's it."
- ** "Meta should just copy bytes -- it can't locate itself, it can't make a consumer bytes better. But there's a positive feedback cycle in the United States over the last decade: B2B is too easy to make money, and no one wants to study how to make consumption money."**
AI Artificial Myth
- ** "When I got into this business, the era of personal heroism was over -- so no hero."**
- ** "No one old Den is your relative -- so you think he's stupid, he's stupid, he can just say he's stupid, whatever."**
- Why do you say that? "I have no mentors in this industry, no old friends, of course I spray anyone I want." "This is an area of objectivity - ** How you do in this field is an objective evaluation criterion and will eventually be respected.** As long as you have a good point of view, and are not in the wrong, there is no need to worry too much about who gets offended by the view."
- Why? The most important trait of this industry is to rely on merit, to do well and to be responsible for what they do." ** In physics, he's seen a lot smarter people than himself (e. g. his doctoral mentor, Douglas Stanford) -- "Where is he? Where do you need me?"
- ** An assessment of the old AI "heroes"** (calling later, but the clue is clear):
- XXX (a person who sees him in vague terms) ** "I think he's always been stupid" ** ** in Pauli's words, he can't even be miscalculated, because what he says is not clearly defined - I hate this vague person most, and it doesn't make sense. ** "I don't know.
- He's willing to admit the hero:
- ** Haldane** (Holtan, founder of Condensation Physics) - "He first proposed Haldane Model and a fraction of Haldane House-related things, decades before the whole field became clear, but he could feel that it was important and was moving forward."
- ** Geoffrey Hinton** -- ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** ** This may be a hero character."
- Transformer Collective (Noam Shazeer, Ashish Vaswani, Niki Parmar, etc.) - "This may be a hero collective."
- Attitude to "Old Den" (in Chinese)
- ** "Most of the seniors are actually good - people get old in two ways: one with high moral standing, no more stabbing, true guidance to young people, and the other with little knowledge of what they are talking about and a special love for stabbing and fingering. Growing old doesn't necessarily mean getting old. ** "I don't know.
- He was not as straightforward as he was from the beginning: "The student age was more restrained, but it was later discovered that restraint was not good for himself or for others. After entering AI, it becomes more direct -- nothing stops me, and the field is objective enough."
For young people
- Pure language model direction: ** "The Blue Sea isn't the Blue Sea anymore. I caught up to the last bus."**
- But AI is a very large area - ** multimodular generation, robotics, and the use of AI to solve practical scientific problems** (e. g. quantum control) is still the Blue Sea.
- ** "To be young enough, it may not be right to do the hottest thing now; to do what nobody does, it may be better."**
About your future
- Not in Google.
- "I'm still going to challenge myself, to torture myself, just to find something worth tormenting myself."
- It's not likely to jump any more; it's not like to do AI for Physics -- "Many people are already doing it, not many more me, not many less me."
- Current top priority: push ML coding and long horizon to relative stability.
recommended
- A book that changes life: ** "I don't." ** Recently read ** Hideki Yukawa, 1949 Nobel Prize in Physics, an autobito "Tabito"** - "See a very successful scientist later, true struggle when he was young."
- Leisure reading: ** From the New World** (You Jisuke, Japanese novel).
- Favorite place: ** Hawaii**.
- Food: Sushi.
- He thinks the most influential AI paper: ** Seq2Seq**
- Scaling Law paper (Jared Kaplan et al., OpenAI) - "Although the specific method may not be entirely correct, it is the first paper to introduce this systematic research approach into the field, and it is essential."
- MBTI?
The last question is: "What's the key bet?"
** "Long horizon. (long distance)"**
Supplement: several cross-checks and background notes
** Contrast of the reasons for the departure of Anthropic**: Yao Sun-woo's statement in his personal blog (alfredyao. github. io) coincided with the interview - stressing that "it does not want his experience to be limited by a specific laboratory, especially since the core research is rarely published". During the interview, he directly said ** about 40 per cent of the opposition to Dario's anti-China position**, which is also cross-referenced in his blog and in public reports like 36kr, Shinjimoto, etc. 2. Reliability of the participating model: 36kr report confirms his involvement in Claude 3. 7 (regional work) and Claude 4 family (RL Uniteds); Gemini 3 Deep Think ' s involvement is also confirmed in Google Home Office announcements. ** Non-Eurmy Spectrum Effects**: The "cycle/open border results he described in the interview are totally out of tune and all the symptoms are accumulated at the border" is the core finding of the PRL paper Edge States and Topological Industries of Non-Hermitian Systems (Yao & Wang 2018) that fully matches his description - the `King Zhong Wang' in his subtitle is ** King Zhong Wang**, ** Zhang Zhang Zheng / Suu City is the first Shuchueng Zhang**. ** Dr. Douglas Stanford and Stephen Shenker are top high-energy/quant gravitationalists for Theoretic Physics, who specifically stated in the interview that Douglas Stanford "is much smarter than me" - a genuine awe. ** "Claude 3. 6 is, in fact, 3. 5 new": this is consistent with the history of the Anthropic official name, and the outside community does call itself "3. 6" because Claude 3. 5 produced two versions. ** Between the time of cataloguing (2026, March) and the time of publication (2026, May) has occurred: Meta's acquisition of Manus was cancelled, Cursor may be acquired by SpaceX, XAI may be incorporated into SpaceX - the relevant expression in the text remains recorded at the time of recording, and the interview's guests' chute against XAI (which has been quite volatile) has been taken seriously.
A quick look at the core view
| Dimensions | Yao Sun-woo's judgment |
|---|---|
| ** Pre-training** | Far from it, it's been getting stronger over the past few months; it's probably the code that hit the wall. |
| ** Post-training** | The real scale starts with Claude 3. 7; the key is that data distribution is narrow and precise. |
| Coding | The outbreak originated from a clear reward signal + GitHub Data Base; is already the only large-scale success scenario for AI-native |
| ** Robot / Multimodular Generation** | It's not even GPT-1 yet. It's still in character engineering. |
| Chatbot form | Stupid. Far from being the ultimate form. |
| Wrapper Survival | They either grow fast enough, or the market is small enough; otherwise they're all bought. |
| AI Security | Anthropic's "the strongest model to speak" is naive; the real institutional analogy is a nuclear-weapon-based deterrent. |
| ** Distillation** | Hard steam is shameful and stupid; soft steam is a pioneer in multi-agent training, and it's technically interesting. |
| ** Organization** | System solid > personal hero flash; OpenAI is the reverse |
| ** Heroes** | Language modelling is over; it's surfers now. It's basically that wave. |
| AI essence | Simple -- because it's possible to experiment, it's just math and infra, it's not a problem. |
| ** To young people** | Language model Blue Sea is over; nobody does anything. |
| ** Personal style** | It's direct, it's squirm, it's "Old Den's not your relative" and it's not vague. |