Goldman Sachs's Silicon Valley Field Trip Yields Key Findings: AI Agents Enter the Execution Era, Competition Shifts to Workflows, and World Models Rise

Deep News
Yesterday

Artificial intelligence is entering a new phase, moving from merely "providing answers" to "taking action."

According to a recent report from the trading desk at Goldman Sachs, AI commercialization is shifting from a "per-seat subscription" model to charging based on consumption volume, transaction volume, and results achieved. Simultaneously, AI agents are evolving from auxiliary tools into workflow executors. The locus of industrial value is also migrating away from the models themselves toward proprietary data, business context, and domain-specific expertise.

This signals a fundamental shift in AI industry competition: the battleground is no longer "whose model is more powerful," but rather "who can truly master the workflow." While model capability remains important, the ability to integrate into enterprise production environments, understand business context, and reliably complete tasks is becoming the more critical competitive moat.

These conclusions stem from Goldman Sachs's recent on-the-ground research across the Silicon Valley AI ecosystem. On August 18-19, Goldman Sachs conducted its third consecutive annual tour, visiting AI startups, leading venture capital firms, and researchers from Stanford University, UC Berkeley, and UC San Francisco. The firm believes that as agents accelerate their deployment, the value distribution among frontier models, open-source models, world models, enterprise software, and proprietary data will undergo significant realignment.

Agent Deployment: The Real Enterprise Bottleneck is 'Controllability,' Not Capability

If the past AI paradigm focused on "helping humans complete tasks," agents are now attempting to "complete the entire task autonomously." However, in large-scale enterprise deployment, the biggest obstacle is no longer model capability, but rather how responsibility is allocated and whether the entire execution process can be effectively controlled.

The report, citing Stanford researchers, notes that most enterprises are still operating under human-supervision models. Especially in fields like law, risk management, insurance, and auditing, the questions of who bears responsibility when a model errs, how to trace the process, and whether corrections can be made promptly are just as critical as the model's inherent capability.

Consequently, the workflows most likely to achieve automation first typically share three characteristics: clear decision boundaries, verifiable results, and the ability to roll back errors. Invoice processing is a prime example. The AI extracts fields and performs checks, low-confidence cases are routed for human review, and the process concludes with a reversible entry into the ERP system.

This also implies that information service providers possessing trusted content, validated domain models, and mature regulatory relationships are more likely to be the first to penetrate enterprise production environments.

Model Competition: Frontier and Open-Source Models Head Towards Division of Labor

Regarding the "open-source versus closed-source" debate, Goldman Sachs's findings do not suggest a binary choice, but rather that different models will correspond to different tiers of workflows.

The frontier model camp argues that enterprise benchmarks often underestimate model capabilities. In real-world production environments, the business losses from degraded model accuracy can far outweigh the savings in inference costs. As a result, several AI-native companies, despite publicly claiming a multi-model strategy, remain heavily reliant on frontier models for their core production needs.

The opposing view contends that the vast majority of enterprise workflows do not require frontier-level intelligence. As open-source model performance continues to improve, customers are increasingly willing to trade a small performance deficit for significantly lower inference costs. One venture capital firm projects that within the next 12 to 18 months, approximately 90% of inference tokens will be processed by open-source models.

This suggests a clearer division of labor in the future AI model market: frontier models will handle high-value, high-reliability complex tasks, while open-source models will manage larger-scale, standardized tasks and absorb the majority of token consumption.

World Models: AI Compute Could Find a Second Growth Curve

Over the past 18 months, researchers have increasingly shifted their focus from large language models (LLMs) to "world models."

Unlike LLMs that primarily train on internet data, world models require an understanding of environments, causality, physical laws, and dynamic interactions within the real world. Their data sources are more heavily drawn from physical systems, specific industries, and operational scenarios. This suggests that the importance of proprietary data is likely to increase further.

Goldman Sachs believes that the problem space for fields like physics, industry, science, and robotics is far larger than that of text generation. Moreover, these workflows typically demand significantly higher compute investments. As AI transitions further from the digital realm into the physical world, the computational requirements for model training, simulation, and inference could well create a new growth curve.

Goldman Sachs projects that compute demand could grow roughly 24-fold over the next five years, potentially extending the period of supply-demand tightness. Cloud computing and compute infrastructure companies such as Microsoft, Oracle, and CoreWeave are expected to be direct beneficiaries.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10