In recent years, the artificial intelligence (AI) sector has been predominantly focused on a fundamental question: can the industry build sufficient computing power fast enough to keep pace with demand? The rise of large language models (LLMs) has triggered an unprecedented infrastructure race, with investors fully engaged. Nvidia sits at the heart of the AI ecosystem, while hyperscale cloud providers are expanding their data centers at a record pace. Furthermore, AI-native cloud services have emerged to serve organizations that cannot access computing power through traditional channels. Cranes continue to rise, and power purchase agreements are being signed. Industry attention has naturally followed these developments, with discussions centered on accelerators, network architecture, power consumption, cooling systems, and data center construction. Investors track GPU shipments; enterprises compete for limited available computing power; and cloud providers vie for increasingly scarce infrastructure. This focus is understandable, as AI demand has grown faster than the industry's ability to meet it. Moreover, this demand is expected to continue accelerating. A Tirias Research whitepaper forecasts that by 2030, annual LLM inference volume will grow from 990 trillion tokens in 2024 to over 100 quintillion tokens, with significant growth also expected in image and video generation.
Evolving Challenges
However, as AI deployments mature, the industry is encountering a different kind of bottleneck. As AI applications evolve from simple chat to intelligent agent workloads, the challenge is shifting from acquiring AI infrastructure to effectively utilizing it. The core of this challenge lies in the fact that running a model is largely a computational problem, whereas running an agent is a systems problem. Agents change how infrastructure is used and what it consumes. Tirias Research describes this evolution as a shift from the "first wave" to the "second wave." In the first wave, AI is primarily used as a conversational assistant responding to individual prompts. In the second wave, autonomous agents can reason, call upon tools, maintain context, and execute complex, multi-step tasks with minimal human intervention. These agents often run continuously rather than ending after a single response. Consequently, the research estimates that an average "second wave" user consumes 40 times more tokens than a traditional "first wave" chat user, with technical users running asynchronous agents consuming orders of magnitude more.
The New Bottleneck
As workloads become more persistent, autonomous, and interconnected, the challenge has moved beyond simply providing computing power to orchestrating complex systems around that power. Recently, Nvidia CEO Jensen Huang described the industry as reaching another "inflection point" driven by agentic AI. An inflection point is not merely growth; it is a point where the trajectory changes, in this case, a point of acceleration. While the underlying technology remains crucial, the focus has shifted. Agents must plan, access tools, interact with external systems, maintain memory, and coordinate operations across multiple services. Infrastructure remains essential but is no longer sufficient. Success increasingly depends on systems that connect models, tools, data, and actions into reliable workflows. As summarized by Kevin Krewell, a senior analyst at Tirias Research, "The question isn't whether AI infrastructure still matters, but what becomes important next?"
Infrastructure as the Starting Point
The first phase of modern AI was about proving what the technology could do. The second phase has been about building the capacity to do it at scale. This meant training clusters, inference capacity, storage systems, and high-performance networks. The emergence of AI-native cloud services is a direct response to this need. From this perspective, AI clouds are evolving beyond infrastructure optimized solely for AI workloads. They are becoming platforms that combine compute, models, data, tools, orchestration, and operational services to provide an environment for building and running AI systems. But infrastructure is increasingly looking like the beginning of a journey, not the destination. Modern production-grade AI deployments typically involve more than just a model and a GPU. Organizations combine retrieval systems, vector databases, observability platforms, evaluation frameworks, agent architectures, security controls, orchestration layers, and third-party integrations. The resulting hardware and software stack can span multiple vendors and multiple clouds. This challenge is even more pronounced in areas like "Physical AI." Building autonomous systems also requires simulation, synthetic data generation, model training, validation, deployment, telemetry collection, and continuous retraining. The scale of the systems built around the model can easily dwarf the model itself. Acquiring infrastructure remains necessary, but knowing how to assemble everything around it is becoming the harder problem.
A Familiar Evolution
This is not the first time the technology industry has encountered such a transition. The history of computing can be seen as a history of abstraction. Operating systems reduced the need to understand underlying hardware; virtualization eliminated the burden of managing physical servers; cloud computing removed the need to build data centers; and serverless computing removed infrastructure deployment altogether. Each step moved users further from the underlying technology and closer to the outcomes they sought. AI appears to be following a very similar path. Most organizations still build AI from the ground up. They start by selecting a model, configuring infrastructure, setting up frameworks, connecting APIs, and assembling workflows step by step. Considerable work is required before anything useful is built. However, customers increasingly seem to want something different. They want to start with the problem they are trying to solve and let the platform handle the rest.
The Rise of Workflow-Centric AI Clouds
This is where the current generation of AI cloud services becomes interesting. Most companies still call themselves infrastructure providers. For instance, Nebius explicitly positions itself as an AI cloud. But examining its customer case studies, product announcements, and conference themes reveals something more nuanced. The recurring theme is not infrastructure capacity, but improving the systems around the models. Case studies focus on orchestration, retrieval, observability, evaluation, and deployment, not just model performance. The emphasis is consistently on reducing the complexity of building reliable AI systems, not just providing additional infrastructure components. Nebius's product development remains deeply rooted in its technology stack, but the company is increasingly packaging these capabilities around customer workflows. This is evident in its support for open-source models, third-party tools, ecosystem partnerships, and higher-level platform services. This pattern suggests the company is focused not only on providing infrastructure but also on making that infrastructure easier to apply to real-world AI development and deployment challenges. Historically, technology vendors often built products and expected customers to adapt their workflows accordingly. The emerging AI cloud model is reversing this relationship, starting with customer workflows and evolving the platform to support them. In essence, customer workflows are becoming the product requirements.
When Workflows Become the Product
The emergence of workflow-oriented platforms illustrates this trend well. In the past, cloud platforms exposed infrastructure components. Customers selected virtual machines, databases, storage systems, networking services, queues, and observability tools, then assembled these pieces into applications. The platform provided the building blocks, and the customer provided the architecture. Now, AI platforms are increasingly operating at a higher level of abstraction. Instead of exposing individual services, they are beginning to expose workflows. Nebius's recent launch of "Agents Blueprint" provides a practical case study of this shift. Infrastructure-as-code (IaC) tools like Terraform made it possible to continuously define and deploy infrastructure, reducing operational complexity and improving repeatability. Nebius's "Agents Blueprint" extends this concept beyond infrastructure by packaging complete AI workflows—including models, retrieval systems, orchestration frameworks, observability tools, and supporting services—into reusable patterns. The goal is no longer just to deploy resources, but to accelerate the creation of working AI systems by reusing established system templates. Use cases like customer support agents, research assistants, search applications, or autonomous systems are increasingly seen as reusable implementation patterns rather than unique integration projects. Blueprints capture architectures, tools, and operational practices validated in earlier deployments, allowing organizations to start with a proven system rather than assembling from scratch. Each deployment builds on prior experience, accelerating delivery while reducing the risk of repeating mistakes. In the cloud computing era, infrastructure became a service. In the emerging AI era, the workflow itself is becoming the product. This view also helps explain why some AI cloud providers are expanding from infrastructure into higher-level platform capabilities. These platforms are increasingly focused on helping customers assemble, deploy, and operate AI applications while abstracting away much of the underlying infrastructure complexity, rather than merely configuring infrastructure resources.
A Broader Industry Trend
"This shift reflects a broader industry trend. If customer workflows become the primary unit of consumption, then understanding how customers build AI systems becomes strategically important," Krewell notes. "Product roadmaps are increasingly being derived from observing successful implementation patterns, rather than simply adding new infrastructure features." Platforms evolve by helping customers move more effectively from intent to execution. This is one of the most significant shifts currently happening in the AI industry.
From Self-Service Infrastructure to Self-Service Outcomes
Cloud computing transformed software development by dramatically reducing friction. Developers could enter credit card details, configure resources, and start building immediately. Infrastructure was available on-demand, eliminating procurement cycles, hardware purchases, and much of the associated operational overhead of traditional IT. AI cloud providers are extending this concept. The first generation offered self-service infrastructure; the next generation is offering self-service outcomes. Instead of requiring customers to manually configure models, inference endpoints, observability systems, security controls, and orchestration frameworks, platforms are increasingly taking on the responsibility of assembling these components. Recent efforts to embed agents directly into cloud platforms offer an early glimpse of this future. Instead of navigating hundreds of APIs, cloud services, and deployment decisions, platforms incorporate an intelligent layer that can understand customer goals and configure the environment accordingly. Nebius recently launched Nebius Echo, an AI assistant built directly into its cloud platform that allows customers to interact with their infrastructure using natural language. Echo is more than a chatbot; it begins to shift cloud operations from manual configuration to intent-driven execution, where the platform can interpret services, inspect resources, and perform infrastructure operations with appropriate user approval. The customer describes the desired outcome, and the platform manages the execution. Infrastructure remains essential, but it increasingly serves as the foundation upon which customer expertise delivers value, rather than being the end result itself.
The True Inflection Point
The common assumption is that AI infrastructure providers are racing to build the largest clusters. To some extent, this is true. Demand for computing power continues to grow, and GPUs will undoubtedly remain the industry's foundation for years to come. However, focusing solely on infrastructure risks missing the larger transformation underway. Even under the baseline scenario in the aforementioned Tirias Research whitepaper, actual infrastructure deployment begins to lag behind projected demand around 2028, and assuming no major geopolitical or supply chain disruptions, there would still be an unmet annual inference demand of approximately 72 quintillion tokens by 2030. This gap is not merely a call to build more data centers. Given current constraints in real estate, power, and cooling, that path may be impractical even if one wanted to pursue it. Instead, it highlights that AI is evolving into a systems problem. As workloads become more persistent, autonomous, and interconnected, value is shifting from raw computing capacity to platforms that can efficiently orchestrate models, memory, tools, data, and workflows. The first phase of AI clouds was about acquiring computing power. The emerging phase is about abstraction and focusing on lowering the expertise barrier required to use that power effectively. Agentic AI makes this transition urgent. Agents don't just consume compute; they consume orchestration, judgment, and workflows. A single model answering questions is relatively easy to deploy. An agent that can autonomously plan, manage memory, call external tools, and execute multi-step workflows exposes every gap in the technology stack. Pure infrastructure was never designed to fill these gaps. This is the true inflection point. It is not a departure from the AI cloud, but its maturation. As Tirias Research's Krewell states, "The most successful AI cloud platforms will not be the ones that expose the most technology, but the ones that allow customers to think about technology the least." The cloud computing era taught organizations how to consume infrastructure. The next phase of the AI cloud is teaching infrastructure providers how to consume customer intent. When this happens, the defining characteristic of an AI cloud will no longer be the hardware it runs on, but how effectively it can translate customer goals into working AI systems—and how little the customer has to think about in between.