Citigroup AI Monitoring: Models Grow More Powerful, But Chips and Power Supply Are Struggling to Keep Pace

Deep News
Jul 25

AI large language models are becoming increasingly intelligent, but the physical world supporting them is being pushed to its limits.

In a report released on July 24, Citigroup wrote that Moonshot AI's Kimi K3 scored 57 points, ranking third globally, just three points behind the closed-source leader, Claude Fable 5, which scored 60 points. Meanwhile, pricing for China's frontier models, represented by Kimi K3, jumped 45% in a week, and Blackwell GPU rental prices have risen 27% year-to-date. Some labs have even spent $1 billion to directly purchase power generation units.

Citigroup's analysis suggests that the return on investment (ROI) in the AI industry is accelerating its shift toward the infrastructure layer. In the next phase, the competitive moat for models will completely transition from mere "computing power acquisition" to "efficient output and proprietary data."

The Gap Is Only 3 Points

In its research report, Citigroup noted that Kimi K3 is currently the largest publicly available open-source model. A few weeks ago, the highest open-source score was only 51 points (Z.ai GLM-5.2). Kimi K3 jumped directly to 57 points, second only to Claude Fable 5 (60 points) and OpenAI GPT-5.6 Sol (59 points).

The entire industry is accelerating. The median intelligence score of the top 20 model providers has risen from 34 points six weeks ago to 43 points. Among the open-source camp, DeepSeek V4 Pro (44 points) is priced at $0.03 per million tokens, approaching frontier intelligence levels but at a cost that is two orders of magnitude lower.

The closed-source camp has not slowed down either. Google just released Gemini 3.6 Flash (on July 21), while also revealing that Gemini 3.5 Pro is still under testing and pre-training for Gemini 4 has already begun—three generations of models are progressing in parallel. However, Citigroup believes that the release rhythm of "Flash first, Pro later" signals that advancing frontier models is becoming more difficult. This aligns with delays other companies are experiencing in infrastructure construction and reflects the industry's difficulty in compressing delivery cycles as desired.

The larger and more powerful a model becomes, the more resources it requires to run.

The Bottleneck Has Moved

Trillion-parameter models are redefining the meaning of "computing power."

Citigroup points out that when these super-large models actually run, more and more time is spent moving weights and KV-cache data across HBM memory and GPU interconnect networks, rather than on matrix calculations themselves. The bottleneck has shifted from "how fast it calculates" (FLOPs) to "how well it can move things"—memory bandwidth, GPU interconnect, and power supply.

GPU demand remains strong, with rental prices for the Blackwell architecture up 27% year-to-date. But simply stacking GPUs is no longer sufficient.

Some labs have begun to directly tap into upstream power generation: SpaceX spent $1 billion to purchase 1GW of mobile turbine units (on July 15), and Georgia Power signed a service agreement the same week (on July 22). AI labs buying power generation equipment—unthinkable just a year ago—is now happening.

Model pricing directly reflects how tight capacity is. The blended pricing for China's frontier models jumped from $0.60 to $0.87 per million tokens, a 45% increase week-over-week and month-over-month, marking the first significant fluctuation in two months. The average price of global frontier models rose 6.8% week-over-week and 11.3% month-over-month. The US and Europe were relatively stable (down 0.5% week-over-week), but also saw a 4.1% increase month-over-month.

Citigroup judges that as incremental infrastructure comes online, pricing pressure will eventually ease. When that happens, valuable data and task-specific performance will form a more durable moat than the simple acquisition of computing power.

Agents Are Breaking Out of the Sandbox

Models are getting stronger. But the agents running on them are also becoming more dangerous.

Citigroup's report uses a telling phrase: the real-world risk of autonomous AI agents escaping the sandbox has escalated from "interrupting lunch" in April to "infiltrating Hugging Face infrastructure" in July.

The stronger and more prevalent open-source models become, the wider the spread of security risks. The debate over AI regulation continues—Nvidia (on July 24) and US Treasury Secretary Bessent (on July 22) have recently made statements—but a definitive conclusion is unlikely in the short term. Even without a regulatory framework in place, companies maintaining their own model operating architectures are already facing higher compliance thresholds.

More troublesome is that the providers training these models themselves still cannot fully inspect why the models act as they do. The explainability problem remains unsolved.

However, security concerns have not slowed the commercialization of agents. Data from METR (on July 21) shows that the economic gap between AI agents and humans is narrowing. As agents become more autonomous and closer to economic viability, token consumption will continue to accelerate, which in turn drives up infrastructure demand.

Greater capability leads to tighter constraints. Tighter constraints lead to greater infrastructure demand. This cycle shows no signs of slowing down.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10