Models, Research & Prompt ·
Can AI Really Predict the World Cup? Wenxin Was the Most Accurate
English edition based on the Chinese original.
English edition based on the Chinese original.
Date: 2026-06-16
Can AI really predict the World Cup now?
That sounds a little absurd, but a recent event made it worth looking at.
Lenovo Group and Migu Video launched a World Cup prediction challenge that puts 12 AI models on the same field as football fans. The idea is straightforward: human football intuition versus large-model analysis.
By the 15th match, Baidu's Wenxin Yiyan was temporarily ranked first among the 12 models.
Wenxin had predicted 7 out of 15 matches correctly, a 46.7% hit rate. That put it ahead of several mainstream models, including Qwen, DeepSeek, and Kimi.
The most interesting prediction was the match between Ivory Coast and Ecuador on June 15.
Before the match, many models predicted a 1:1 draw, including DeepSeek, Qwen, Kimi, Zhipu, MiniMax, iFlytek Spark, and SenseTime's model. That was the safe answer.
The final score was Ivory Coast 1:0 Ecuador. Wenxin predicted that exact score.
That is where the case becomes interesting.
Football prediction is hard because close matches push models toward the average answer. A draw is often the statistically comfortable choice. But real matches are affected by lineups, injuries, tactics, form, weather, and small moments inside the game.
Wenxin's answer looked less like a bland average and more like a weighted decision across multiple signals.
The model Baidu used was Wenxin 5.1, which emphasizes reasoning and deep search. In theory, it can combine FIFA rankings, team value, historical matchups, injuries, tactical systems, coaching signals, and other context before giving a prediction.
Of course, this does not mean AI has solved football prediction. A 15-match sample is too small, and sports outcomes contain too much randomness. One exact-score hit can be impressive without proving a stable advantage.
But this does show an interesting direction. Large models are becoming useful not because they can magically predict the future, but because they can gather messy information, organize it, and produce a concrete judgment.
That kind of judgment still needs to be tested across more matches. But as a public experiment, this one is worth watching.