cxuan-ai-labs

Industry, Business & Companies ·

At the invitation of the organizing committee, a meeting was held in Katsumura, Beijing.

English translation of the Chinese original. This version is generated for international readers and may be refined over time.

Interactive reading view ↗

English | Chinese Original

English translation of the Chinese original. This version is generated for international readers and may be refined over time.

Date: 2026-05-23

At the invitation of the organizing committee, a meeting took place in Katsumura, Beijing.

The whole meeting sounds like the national budget is coming up.

Say a few scenes heard something.

This road map has been launched directly. It is clear that China is Fellow Liao, and for Agenic AI, the measure of chips, memory bandwidth, memory capacity, interconnective IO bandwidth is four indicators, with different scenarios of priority. The real decision on the ability to exceed nodes is interconnection. The 950 chip is a clustering of a pile of chips with super-high bandwidth, with one core objective - the cost of reasoning down.

There's an interesting detail. The MoE model's EP communication is the lifeline of a time extension, with the age of Age Agent pressing to milliseconds, but the EP's All-to-All communication particles are fined to a single package of 7 to 14KB, and the frequency rises with the number of specialists in square scales, and the traditional network is simply not functioning. The way to do this is to put the EP communication in the Scale Up domain, with small particles in the Load & Store synonyms, and then the DMA. This is not a stack of parameters, but a system-level optimization at the communications level.

KV Cache also used a knife. In the Agent scenario, the number of frequency surges increased 50 to 100 times, the serial length was pulled from 4k to nearly 1 trillion, and Cache hit 95 per cent. High hit rates are good, but the cost of KV Cache goes up. In order to provide a direct UB port to the SSU unit, China has provided NPU with KVCache to save the intermediate layers of the storage system and the file system, with a bandwidth scale. Such work cannot be done without hardware.

There's no ambiguity here. Hu Xinwei is straightforward, and the substrate is no longer designed for training models only, but is reshaped by the Agent load. Supernode TB-level bandwidth, 100-nab extension, global memory unified site, rollback dry to 10 ms, and Agent mission success directly above 10%. The three technologies in the accelerated base of communications are also practical: a 20 per cent drop in Lingua SSL, a 40 per cent drop in Transparency UBSocket without changing the source code and a 90 per cent drop in shared TP communications memory. The translation is to use all the accumulated communications over the years.

And then the fun came.

The Assembly was open and, on the evening of 22 May, DeepSeek announced that the V4-Pro model API price had been permanently reduced to a quarter of the original price as of 31 May. Enter tokens 0. 025 per million of caches, output six. Not promotion, permanent.

Look at this rhythm. The cost of decomposition of the hypernodes in the hardware layer, the direct link of SSU, and the EP communication to ms. The sandbox infrastructure, which is high-density low-duration sandbox infrastructure, is used in the generic computing layer, with Agent mission costs being hit downwards. DeepSeek then went straight through the model price.

It's not a coincidence. When the hardware layer breaks out the cost space, the application layer is willing to press the price on the floor. In turn, the more people are used, the more the hardware becomes.

Throughout the scene, the greatest feeling was that the wave of domestic computing was not following single-card parameters, but was a system-level performance with cluster efficiency, communication time, bandwidth optimization with KV Cache, and Agent load. It's a spiral that's moving.

DeepSeek, the wave of permanent price reductions, is the signal that the cycle begins to turn.

That's it. AI has three words: affordable. It's a national line. It's coming through.

The country's macho, so everyone can use cheap models.