AI Chipmaker Cerebras Reports Record Quarterly Cloud Revenue, Sees 287% Surge Driven by OpenAI, AWS Bedrock Launch Set for Early 2027

Stock News
Aug 13

Cerebras Systems (CBRS.US) posted record core revenue for the second quarter and raised its full-year guidance after the market close on Wednesday, with CEO Andrew Feldman stating that all key metrics exceeded expectations. The company is forecasting that core revenue will more than triple by 2027, with multiples of growth continuing in subsequent years. Management is positioning 2026 as a "foundation year" to prepare for the fulfillment of over $25 billion in remaining performance obligations (RPO).

Second-quarter core revenue reached $209.9 million, a 103% increase year-over-year. Core cloud and other services revenue surged 287% to $127.7 million, while core hardware revenue grew 17% to $82.1 million. Core gross margin was 40.6%, up approximately 940 basis points from the same period last year, but declined sequentially from 46.5% in the first quarter. This was primarily due to the company temporarily leasing back some systems from cloud customers to meet private cloud demand, a high-cost arrangement that reduced gross margins by about 500 basis points. The company expects the third quarter to be the low point for gross margins, with improvement in the fourth quarter as its own data centers come online, and it maintains a long-term target of over 60%. Core operating loss was $33.6 million, with an operating margin of -16%, improving from -42% in the prior year. For the third quarter, Cerebras forecasts core revenue of $214 million to $216 million, gross margin of 38% to 40%, and operating margin of -25% to -23%. Full-year core revenue guidance was raised to $880 million to $890 million, with gross margin guidance of 41% to 43% and operating margin guidance of -19% to -17%.

On capacity expansion, data center space remains an industry bottleneck. The company has secured over 600 megawatts of capacity (operational or contracted through the end of 2027), with a future expansion pipeline measured in gigawatts. Deployments span multiple US states and locations in France, Finland, and Canada. Manufacturing capacity is being expanded through partners Flex and Sanmina, with capacity quadrupling compared to the first half of 2025, expected to increase more than tenfold in 2026, and triple to quadruple again in 2027. Supply from TSMC (TSM.US) for 5-nanometer wafers is stable, and the company avoids reliance on high-bandwidth memory, CoWoS packaging, or 3-nanometer capacity, reducing supply chain risk.

Technologically, this quarter saw the addition of support for OpenAI's GPT-5.6 Sol model, with a tenfold speed improvement. The company is advancing disaggregated inference collaborations with AMD (AMD.US) and Amazon (AMZN.US) AWS, where GPUs handle the prefill phase and Cerebras systems manage the decoding, aiming to boost throughput by five times, with deployment expected in the fourth quarter. Cerebras will unveil its fourth-generation CS-4 system at the Supernova conference, with the CS-5 planned for the second half of 2027. Performance is expected to double annually for several years, with throughput increasing more than 20-fold over the next 18 months.

On customer expansion, Cerebras expects its products to be fully available through the AWS Bedrock platform in the first quarter of 2027, with the first hyperscale revenue generated by mid-year. However, the $25.4 billion in RPO as of June 30 does not include hyperscalers like AWS. In the second quarter, the company signed six deals worth over $30 million each, with customers including Figma (FIG.US), Cognition, Lovable, Block (XYZ.US), AlphaSense, GSK (GSK.US), and CrowdStrike (CRWD.US). OpenAI will remain a significant share next year, but AWS, coding, and security applications are expected to grow, and "new cloud" providers will become an important part of the business by 2027. The following is the full transcript of the Cerebras Systems earnings call.

Where to begin

Operator: Good afternoon, and welcome to the Cerebras Systems fiscal 2026 second quarter earnings conference call. Please note that today's call is being recorded. I will now turn the call over to Head of Investor Relations, Sean Dorsey. Please go ahead.

Sean Dorsey: Thank you, operator. Good afternoon, everyone, and welcome to the Cerebras Systems fiscal 2026 second quarter earnings conference call. Earlier today, we issued a press release and posted a supplemental earnings presentation on the Investor Relations section of our website. A replay of this webcast will also be available on our Investor Relations website following the conclusion of this meeting. Joining me on the call today are our Co-founder, CEO, and President, Andrew Feldman, and our Chief Financial Officer, Bob Komin. Before we begin, I would like to remind everyone that today's discussion will contain forward-looking statements made under the safe harbor provisions of the Private Securities Litigation Reform Act of 1995. These statements include, but are not limited to, statements regarding our future financial performance, business strategy, market opportunities, customer demand, product roadmap, technology leadership, supply chain, operating model, and our outlook for the third quarter and full fiscal year 2026. Forward-looking statements are based on our current expectations and assumptions and are subject to various risks and uncertainties that could cause actual results to differ materially from those expressed or implied in the statements. These risks are described in our filings with the SEC, including the final prospectus related to our initial public offering and our future periodic filings with the SEC. We undertake no obligation to update these forward-looking statements except as required by law. On today's call, we will also discuss certain non-GAAP financial measures. A reconciliation between GAAP and non-GAAP results is included in today's press release and supplemental materials, which are available on the Investor Relations page of our website. I will now turn the call over to Andrew.

Why just 10 ASX 200 shares?

Andrew Feldman: Thank you, Sean. Good afternoon, everyone, and thank you for joining us today. The second quarter was a strong one. We completed our public offering, but that did not distract us from executing against our plan. We achieved record core revenue and exceeded expectations across all metrics: core revenue, core gross margin, and core operating margin. Looking ahead, we see insatiable demand for fast inference. The market is realizing that speed is not just a benchmark metric. Speed changes user engagement, changes agentic performance, and changes AI productivity. Fast inference unlocks new applications and new markets. As we have shared with you before, 2026 is a foundation-building year for Cerebras. In the seven weeks since our last earnings call, we have made excellent progress on multiple fronts, preparing us for enormous growth in 2027, 2028, and 2029 to fulfill the $25 billion in remaining performance obligations we currently have on our books. As a result of this progress, we expect core revenue to more than triple in 2027 and continue to grow by multiples in subsequent years. We see progress in three areas: capacity, capability, and customers. We are expanding capacity by signing new data center contracts globally, expanding manufacturing capabilities, and working with suppliers to secure supply and support our hypergrowth. We are enhancing our capability by inventing new technologies that extend our advantages in performance, throughput, and energy efficiency. And we are expanding our customer base by accelerating AI productivity in existing markets like code generation and agentic workflows, and opening entirely new domains like security, where speed creates unprecedented opportunities. On capacity, data center space remains a bottleneck for the entire industry, and we are no exception. The faster we and our customers can bring new data centers online, the faster we can grow. So, over the past seven months, we have been fully focused on securing and building data centers. We have two advantages. First, because we serve inference, we do not require the gigawatt-scale infrastructure needed for training clusters. This gives us more flexibility in scaling capacity across multiple locations globally. Second, we have built a repeatable process for site selection, cluster deployment, and customer activation, which is the operational capability needed to convert gigawatts of power into usable production tokens at global scale. I am pleased to report that our efforts have been very successful. We now have data centers that are either operational or contracted in Alabama, Dallas, Denver, Minneapolis, Santa Clara, Stockton, and internationally in France, Finland, Manitoba, Montreal, Norway, Saskatchewan, and Toronto. In total, over the past seven months, we have secured over 600 megawatts of data center capacity that is either online or will be delivered by the end of 2027. While this is far from sufficient to meet our demand, our pipeline of data center projects for future expansion continues to grow and is now measured in the gigawatts. To put it in perspective, as we continue to build our first-party cloud, it will become one of the largest AI clouds outside of the hyperscalers. And while at the end of 2025 we were still on a steep learning curve, today I am pleased to report that we have become quite proficient at data center build-out and have a clear path to excellence. Another key dimension of capacity is manufacturing and supply chain. Here, we have successfully scaled our manufacturing capabilities and are building new factories with Flex and Sanmina, with the expectation of increasing manufacturing capacity more than tenfold in 2026 and continuing to expand this capacity in 2027, again, to prepare for the extraordinary growth we expect in the coming years. Our relationships with supply chain vendors have also translated into significant advantages. TSMC has been very supportive, and we have the wafer supply needed to drive growth. Our ability to secure wafer supply is also aided by the fact that we can achieve industry-leading performance on TSMC's 5-nanometer node, which has lower wafer costs and relatively less tight supply. Our decades-long relationships with our supply partners further bolster our confidence in achieving our future growth plans. These relationships are scarce and valuable, especially during periods of tight supply. Finally, it is important to remember that most of the key supply chain constraints facing the industry today do not apply to us. For example, we do not use HBM memory, CoWoS packaging, or require 3-nanometer wafer manufacturing capacity. On capability, in the second quarter, we enabled support for OpenAI's GPT-5.6 Sol, the largest and most capable frontier model. In fact, Cerebras delivers GPT-5.6 Sol at 10x the speed. With GPT-5.6 Sol, any residual doubts about our ability to support large frontier models have been dispelled. Being the delivery partner for GPT-5.6 Sol and serving it on our cloud speaks volumes about the maturity of our software stack. Reaching a level that can deliver hyperscale-quality and reliability requires millions of system hours of grueling production environment testing. We are proud that our inference cloud can meet the demands of the most demanding customers. Our work serving frontier models has opened up new and significant strategic advantages that previously only NVIDIA possessed. Closed-source frontier models contain a continuous stream of new insights and new technologies. Serving these models allows us to see the future and prepare for it. Our hardware and software stack roadmap now reflects the trends we are seeing and will provide us with compounding advantages in the coming years. Continuing on the theme of capability, let's talk about disaggregation. We now have disaggregated inference solutions with two leading chip companies: AMD (using its Helios systems) and AWS (using its Trainium chips). Disaggregation expands the market for both GPU vendors and Cerebras. Disaggregation allows GPUs to participate in a market that is currently closed to them: the fast inference market. Disaggregation allows Cerebras to expand our opportunity to more price-sensitive customers and improve the profitability of our data centers. Let's look at how it works. As with any computing market, as inference grows and matures, opportunities for specialization arise. Disaggregation is a form of specialization, particularly applicable to workloads with known traffic patterns. In these cases, disaggregation provides an advantage by splitting inference into two phases—prefill and decode—and using a different processor for each phase. Prefill processes input from the user or agent. This is a parallelizable workload. Therefore, prefill is well-suited for GPUs and their HBM-based memory architectures. Decode is responsible for generating output tokens. This is the technically more difficult problem and is where most of the computation in a disaggregated solution resides. It is sequential and extremely memory bandwidth-intensive, making it particularly well-suited for our wafer-scale engine. The prefill and decode processors need to be connected to form an end-to-end solution. And in this area, our standards-based I/O and open collaboration strategy make integration with Cerebras straightforward. A few weeks ago, we announced a partnership with AMD to build disaggregated inference solutions. These solutions combine their Helios racks with our CS systems. The combined solution increases throughput by 5x while maintaining Cerebras' speed. To understand how powerful this is, it's crucial to distinguish between 'speed' and 'throughput'. Speed is a metric for a single user. It is measured in 'tokens per second per user'. It measures how quickly your query gets a response or how long it takes an agent to complete a task. This is the X-axis. Throughput, on the other hand, is the total number of tokens a solution can generate per second. It is measured by summing all tokens from all concurrent users. This is typically shown on the Y-axis. Speed is critical to user experience. Throughput is critical to the economics of inference. GPU solutions can support high throughput, but only at low speeds. When configured to support even moderate speeds, GPU throughput drops dramatically. This is not just true for GPUs, but also for ASICs and all solutions using HBM. HBM memory architecture forces a trade-off between throughput and speed. SRAM-based architectures like Cerebras are the complete opposite. We support extremely fast token speeds, but have moderate throughput. So, GPUs want to become faster without sacrificing throughput. Cerebras wants to achieve higher throughput without sacrificing speed. This is the advantage of our disaggregated solution. It provides Cerebras' speed at 5x higher throughput. Increasing throughput by 5x while maintaining our industry-leading speed has a profound impact on the unit economics of token generation. This means up to 5x more high-speed, high-value tokens per Cerebras system. More tokens per system at lower cost means higher revenue and higher gross margins. More tokens per CS system also means more tokens per watt, making every data center more profitable. And perhaps most importantly in a data center resource-constrained environment, the disaggregated solution allows us to fulfill more of the demand in our RPO. Finally, we believe this disaggregation approach will deliver performance and economic benefits on any GPU. For operators with large GPU fleets already deployed, disaggregating with Cerebras provides them with an opportunity to create significant leverage on their existing investments by pairing some of their GPUs with Cerebras solutions, thereby significantly enhancing the value and utility of their data center assets. Continuing on the theme of capability, let's turn to our roadmap. Our engineering execution is progressing steadily. We expect to deliver new systems that double in speed every year for the next several years. Remember, we are doubling performance on top of a performance advantage that is already 15x faster than anyone else in the industry. Furthermore, while maintaining performance leadership over the next 18 months, we plan to deliver solutions that increase throughput by more than 20x. Next week, at our annual Supernova conference, we will unveil our fourth-generation system, the CS4. This will be a big event with many product launches, so I recommend attending. Finally, we are currently on track to launch our CS5 in the second half of 2027. Looking further ahead, our innovation engine is humming. We have significant partnerships with the US government on delivering stacked memory solutions and integrated wafer-scale optical solutions. In the coming years, you can expect innovations from us in chips and chip architecture, as well as all aspects of system design, including packaging, I/O, and power delivery. To summarize the capability section: we expect to continue delivering groundbreaking advances in products and technology to increase speed and throughput, reduce energy per token, and dramatically lower the per-token cost of our solutions. Now let's turn to customers. Fast tokens are in high demand and command a premium in the market, and fast tokens with frontier intelligence are only available through the OpenAI and Cerebras partnership. Our collaboration with AWS continues, and we expect our solution to be formally launched on the AWS Bedrock platform in the first quarter of 2027. This AWS partnership expands our market opportunity and gives us global reach through an industry leader trusted by nearly every enterprise worldwide. Discussions with other hyperscalers are also progressing well. We expect to generate initial revenue starting in mid-2027, with scale ramping in 2028 and beyond. Amid all this progress, I think it's important to remember that our $25 billion RPO does not currently reflect any backlog from AWS or any other hyperscaler. Our business outside of OpenAI and hyperscalers continues to grow healthily. For example, in the second quarter, we signed 6 deals worth over $30 million each. AI code generation continues to grow at a high rate. In our experience, no one is happy with slow tokens when programming. So, it's no surprise that our presence in the code generation space is expanding. We signed new agreements with public companies like Figma and leading startups like Cognition, and we made significant progress in Europe with a major deal with Lovable. Agentic workflows are growing rapidly, and as agent operations quickly evolve into multi-step, multi-agent solutions, the value of speed compounds. Several companies, including Block, AlphaSense, and GSK, signed new agreements with Cerebras in the second quarter, leveraging fast inference to power their custom AI agents. Fast AI is also opening up new markets, expanding Cerebras' total addressable market (TAM). Security is one such example. Our recent partnership with CrowdStrike is an application only possible if AI is fast enough. Fast AI enables AI-based security appliances to be embedded in enterprise traffic, using large language models to secure traffic so quickly that it goes unnoticed. Fast AI allows large language models to provide security that is invisible to the user. AI provides security, and speed creates the 'stealth' that makes security possible without latency or disruption. Given the rapid evolution of the threat landscape, we expect such security protection to become the norm. Enterprises will soon expect the vast majority of their traffic to be inspected in this way, creating a massive new opportunity exclusively enabled by fast AI. Frontier labs, hyperscalers, leading chipmakers, the fastest-growing startups, and large enterprises are now all customers and partners of Cerebras, benefiting from our extremely fast inference service. In summary, it was a strong quarter. We successfully completed our IPO. We exceeded expectations on all metrics: core revenue, core gross margin, and core operating margin. We made progress in three key areas: capacity, capability, and customers. These are the foundations for our substantial growth in 2027 and 2028, and for maintaining this remarkable growth rate in 2029 and beyond. I will now turn the call over to Bob. Bob?

Key drivers for growth

Robert Komin: Thank you, Andrew, and good afternoon, everyone. We made tremendous progress in the first half of fiscal 2026. As we have described, fiscal 2026 is a year of laying the foundation for multi-fold growth in the coming years. At the beginning of the year, we won one of the largest technology deals ever, creating over $25 billion in RPO. This required us to immediately begin a massive ramp-up in three key components of capacity. First, we needed to increase wafer supply. As Andrew described, thanks to our strong relationship and support from TSMC, we have achieved this, not only for the remainder of this year but also for next year. Second, we needed to expand manufacturing capacity. Our manufacturing capacity is now 4x what it was in the first half of 2025, and we will increase it by more than 10x in 2026. So we have made great progress there. Third, we needed to dramatically increase data center capacity. We have made significant progress, with over 600 megawatts of capacity already online or contracted, expected to be delivered by the end of 2027, in addition to a gigawatt-scale pipeline. Therefore, the foundation for core revenue to more than triple in 2027 and achieve additional multi-fold growth in future years has been laid. This growth also paves the way for significant margin expansion in 2027 and beyond. Now, turning to the financial results for the second quarter. We had another strong quarter, exceeding expectations on all guidance metrics. We achieved record core revenue. We also exceeded expectations on core gross margin and core operating margin. I will continue to use the core business framework introduced last quarter to describe our progress. The definition of core business metrics and the reconciliation of all metrics to GAAP are included in today's earnings release and on our website. Core revenue was $209.9 million, up 103% year-over-year. Our private cloud business is growing at an astonishing rate. Core cloud and other services revenue was $127.7 million, up 287% year-over-year. This nearly fourfold growth reflects the strong market demand for our Cerebras fast inference service. Core hardware revenue for the quarter was $82.1 million, up 17% year-over-year. We focus on total core revenue rather than the composition between the two, as the mix can fluctuate significantly quarter to quarter based on the timing of large new cloud capacity additions and hardware shipments. In the second quarter, the majority of the growth in total core revenue was attributable to the growth in our core cloud services, reflecting the acceleration of our OpenAI deployment, increased usage from other cloud customers, and some hardware customers also being impacted by the timing of new data center capacity coming online. Demand for fast inference remains strong, with several late-stage hardware transactions, potential order values in the hundreds of millions from new customers, and significant new cloud services deals for 2027. The existing fast inference market is growing, and new markets are being activated. We view disaggregation as a key enabler for new applications, as it significantly improves the unit economics of inference and data center ownership. Currently, this means up to 5x more tokens per CS system, resulting in significantly lower power and cost per token. And, with continued strong investment in R&D and the product roadmap, Cerebras will quadruple our existing industry-leading speed by the end of 2027 and increase throughput by more than 20x, dramatically improving our performance and the unit economics of inference. Now, turning to gross margin. Year-over-year, core gross margin improved significantly. Core gross margin for the second quarter was 40.6%, approximately 940 basis points higher than the second quarter of 2025. This reflects the market's recognition of the value of fast inference, our continuous improvements, and the benefits of scale. Breaking down total core gross margin, core cloud and other services gross margin was 41.8%, 1,600 basis points higher than the second quarter of 2025. Core hardware gross margin was 38.8%, 510 basis points higher than last year. As we discussed last quarter, to meet the enormous demand for our fast inference service, we are temporarily leasing back some of our own systems from cloud customers and providing them through the Cerebras cloud. Meeting this inference demand early enhances our ability to satisfy cloud customer demand and grow with them over time. We believe this will create additional long-term value for Cerebras and its shareholders. In the short term, the higher cost of leased-back capacity reduces gross margins. Consequently, sequentially, core gross margin was 40.6%, compared to 46.5% in the first quarter of 2026. If not for the increased cost from leasing back more systems to increase private cloud capacity, core gross margin would be approximately 500 basis points higher than it is now, closer to last quarter's level. Looking ahead, we expect the third quarter to be the low point for core gross margin, followed by a significant improvement in the fourth quarter of 2026 as more data centers with lower-cost Cerebras-owned systems come online. This will drive a recovery in core cloud gross margins. Core gross margin will also continue to improve in 2027 and trend towards our target of over 60%, for the following reasons: The market has recognized that fast tokens have higher value. This supports higher pricing, which will be reflected in hardware and cloud services deals recognized in future quarters. In the coming quarters, we will gradually phase out the higher-cost leased-back systems and replace them with lower-cost owned systems in our private cloud. Our product roadmap shows a 20x increase in throughput over the next 18 months. This reduces the cost of producing tokens per system and per watt of power. As we scale, our supply chain bill of materials costs will further optimize. Because we use a 5-nanometer node, our wafer costs are lower than those using 3-nanometer or 2-nanometer nodes. And we are purchasing wafers in much larger quantities. Finally, we do not rely on HBM, which puts pressure on HBM-dependent vendors to either raise prices or lose margin points. We do not have this risk, which we believe will enhance our value proposition and pricing flexibility. Now, on operating margin. Core operating loss was $33.6 million. Core operating margin was negative 16%, compared to negative 42% in the same period last year, an improvement of approximately 2,600 basis points year-over-year. The significant improvement in core operating margin while doubling revenue and increasing investments across the board demonstrates the strong operating leverage inherent in our business model. Today, we are investing in world-class talent, manufacturing and data center capabilities, and corporate infrastructure to support the significant scale expansion we expect to achieve in the coming years. As of June 30, 2026, remaining performance obligations were $25.4 billion. This backlog provides us with visibility, confidence in future revenue growth, and the ability to make necessary upfront investments. Our existing large strategic customers provide validation, contract visibility, and the necessary economic support for us to build new capacity at scale. At the same time, we are succeeding in expanding our addressable market and customer base. OpenAI provides us with, among other things, scale and frontier insights. AWS provides global enterprise reach. The recent partnership with AMD expands the market opportunity to include disaggregated inference, and fast inference is opening up new markets like security. We ended the second quarter with over $8.6 billion in cash, cash equivalents, restricted cash, and marketable securities. We also have access to an $850 million revolving credit facility, which remains undrawn. Our liquidity and balance sheet are strong and were further strengthened after the IPO in the second quarter. This is a significant advantage, providing us with the flexibility and capital to invest in the opportunities emerging in this dynamic and high-growth market environment. Additionally, we have an advantage with net capital expenditure per megawatt being significantly lower than the vast majority of AI cloud providers, for two key reasons: First, our capital expenditure is primarily for deploying our own hardware in our data centers, which has a much lower bill of materials cost and does not include the high-profit margins that many other companies must pay. Second, our largest customer compensates us as a data center pass-through cost, covering a significant portion of our remaining data center construction capital expenditure. Now, turning to our outlook. For the third quarter of fiscal 2026, we expect core revenue to be between $214 million and $216 million, core gross margin between 38% and 40%, and core operating margin between negative 25% and negative 23%. For the full fiscal year 2026, we are raising our core revenue guidance to between $880 million and $890 million. We are raising our core gross margin guidance to between 41% and 43%, and we are raising our core operating margin guidance to between negative 19% and negative 17%. In summary, the second quarter was a very strong quarter of continued execution and growth for Cerebras. We achieved record core revenue, cloud and services revenue nearly quadrupled, and gross margins and operating margins significantly exceeded our expectations. We raised our guidance for all metrics for the full year. We made great progress in building capacity, capability, and customers, and we ended the quarter with over $8.6 billion in cash, cash equivalents, and investments to continue executing our growth plan. We are well-positioned to achieve revenue growth of more than three times in 2027, and significant additional growth in the coming years, while significantly expanding gross margins and operating margins towards our targets. I will now turn the call back to Andrew for closing remarks.

Foundation for the future

Andrew Feldman: Thank you, Bob. Over a decade ago, we founded Cerebras with the strong belief that we could build a better processor for AI, and that to deliver such a processor, we would need to build a complete accelerator system and rack. Today, Cerebras is one of only four companies globally—Google, Amazon, NVIDIA, and Cerebras—that can independently design processors, systems, data centers, and provide AI-based cloud services to customers. Thank you for listening to our presentation. Operator, please open the line for questions.

Q&A Session

Operator: Our first question comes from Timothy Arcuri of UBS.

Timothy Arcuri: Andrew, I want to ask about customer concentration. You did mention that revenue will more than triple next year. Obviously, we know OpenAI is growing rapidly now. So that will be a big chunk of your incremental revenue. I guess AWS could contribute $1 billion or more next year. So, how do you think about customer concentration next year? For instance, could these two customers account for two-thirds of your revenue? Also, can you talk about discussions with other companies like Google, Microsoft?

Andrew Feldman: Sure. That's a good question. I think some historical context might help. In 2021, people complained we only had government customers. Then, when we won a billion-dollar sovereign cloud deal, people worried we only had one sovereign cloud customer. Then we won the largest lab, the frontier lab. Then people worried we didn't have a hyperscaler customer, and then we won AWS. So, I think in every case, we have been able to use the momentum from the previous step to expand our business. I think OpenAI is a huge customer, and they are a significant part of our business, as well as the business of all other companies in this field. I think they will still be a large part of our business next year. But you are absolutely right, AWS and other customers, whether it's fast-growing code generation companies or some application cases in security, will account for a larger share, and OpenAI's percentage of our revenue will decline over time. But I think they will still be a significant part of our revenue next year.

Timothy Arcuri: Okay. A quick follow-up. Bob, you mentioned capacity will increase 10x year-over-year this year. Can you give us an idea of the capacity growth expectation for next year? I know Andrew said revenue will more than triple, but can you talk about how much manufacturing capacity will grow year-over-year next year?

Robert Komin: Yes. By the end of this year, our manufacturing capacity will have increased more than tenfold from the beginning. We have signed additional facility contracts for growth in 2027, sufficient to increase capacity by another three to four times, and we still have time to sign more contracts. So, we are looking at significant growth in the coming years.

Operator: Our next question comes from Joshua Buchalter of TD Cowen.

Joshua Buchalter: Maybe following up on Tim's question. Could you elaborate on how the financial mechanics of the Amazon deal will work? Is the plan to offer the service in the AWS cloud next year, and then we will basically see how the actual demand is, making it hard to predict now?

Andrew Feldman: Sure. The service will be available. It will be deployed within Amazon's data centers. It will be offered through Amazon's API service, Bedrock. We are currently organizing the deployment. So, I think that's roughly the framework for thinking about it. We expect the service to go live in the first quarter.

Joshua Buchalter: Understood. A follow-up on the AMD partnership. Can you provide more details on how the joint solution will be brought to market, such as integration with Helios racks and the timeline for revenue generation? Also, regarding this partnership with AMD, they recently acquired an inference hardware company. Can you talk about how this fits into their broader product portfolio and how it relates to your solution?

Andrew Feldman: Sure. A few points. I think we will announce more details about the AMD partnership arrangement in due course. But I think the joint solution of placing the Helios rack in front of the Cerebras system, with the Helios rack handling prefill and Cerebras handling decode, is a very powerful one—a solution that can provide much faster speed than Helios alone and much higher throughput than Cerebras alone. This solution is very attractive, and we already have buyers. Your second question about AMD's recent acquisition. Look, I think the company they acquired is interesting and innovative, and we are more than happy to see innovative hardware. I think acquiring hardware startups takes a long time between acquisition and product delivery. We were impressed by the research of that company, and I think it can find many applications within AMD's product portfolio. However, I think its first application scenario will not be data center inference.

Operator: Our next question comes from Tom O'Malley of Barclays.

Kyle Bleustein: This is Kyle Bleustein on for Tom O'Malley. I wanted to follow up on the economics of the AMD deal. I want to confirm, is the approach that you purchase AMD Helios racks and install them in your own cloud, with all the revenue from customers leasing the disaggregated inference solution going to you? Or is there some form of revenue sharing agreement?

Andrew Feldman: Yes, the answer is the first part of your question.

Kyle Bleustein: Okay. My follow-up is, the AWS deal is deployed in their cloud. Do you see a path where Trainium and CS3 could be co-deployed in your cloud or other hyperscaler clouds? I'm just trying to understand how disaggregated inference might evolve in future deployments.

Andrew Feldman: I think we are very interested in this approach. As you see with Google, hyperscalers have the opportunity to deploy their own designed components outside their own data center footprint. This is a direction we are interested in, not just with AWS but also with other cloud providers. So, I think this is very possible in future collaborations with AWS.

Operator: Our next question comes from Quinn Bolton of Needham & Co.

Quinn Bolton: A quick confirmation on the AMD deal. Andrew, if you purchase and deploy Helios racks into the Cerebras cloud, but the service is provided to OpenAI (under your contract), does that represent an additional revenue opportunity? Or how should we think about the revenue potential in this scenario? I have a follow-up.

Andrew Feldman: I think anytime you can increase throughput while maintaining performance, you increase the revenue opportunity, right? Throughput is the number of customers you can support simultaneously. If you can do that without sacrificing speed, then each system can generate more revenue. However, I don't want to go into the specifics of our relationship with OpenAI. But one of the reasons we are so excited about this collaboration is that it preserves our speed while increasing throughput, making each system more profitable. Each system produces more tokens. This means lower cost per token and less power consumed per token. Therefore, we not only increase revenue but also improve margins; we improve not only margins and revenue but also make each data center investment more valuable, as data centers are limited by power capacity. If you can get more tokens from a given power capacity, that translates into more revenue. So it's a very powerful advantage. And, with AWS and AMD both collaborating with us on disaggregation, we have captured about half of the leading chipmakers. So it's a very, very compelling story.

Quinn Bolton: Then, my follow-up is, you mentioned several times that by the end of 2027, you will increase the throughput of the wafer-scale engine by 20x.

Andrew Feldman: Correct.

Quinn Bolton: Does this eliminate the need for some disaggregated computing or heterogeneous inference? Or will it just make the overall throughput of the heterogeneous solution faster, or reach a higher throughput level?

Andrew Feldman: I think we are exploring various ways to increase throughput, right? If your throughput increases by 20x while costs remain the same, you are in a very good position. So, our systems are improving throughput. We are looking for ways to increase the throughput of the disaggregated solution. We are researching various different inventions, technologies, and partnerships to continue our pattern of industry-leading performance and significantly increased throughput. That's a very good question. I mean, it's a central point in our roadmap thinking.

Operator: Our next question comes from Joe Moore of Morgan Stanley.

Joseph Moore: Following up on what you just talked about regarding disaggregated decoding, where are you in the commercialization process? We understand you can do fast inference at scale. You have also done it with disaggregation. Is it ready for deployment? What work still needs to be done over the next year or so to achieve the level of improvement you are talking about?

Andrew Feldman: Okay. I think our lab is now running disaggregated inference using GPUs. I think it will be deployed and available in the fourth quarter.

Joseph Moore: Okay. And you mentioned the ability to work with deployed GPUs. So, when you work closely with Amazon and AMD, will these collaborative solutions be more effective than what you could achieve with deployed NVIDIA GPUs?

Andrew Feldman: I think it's fair to say we haven't attempted that yet. Of course, the Helios rack will provide us with a larger and better solution than using 355-type GPUs. Trainium 3 will also provide a better solution than using Trainium 2. If we use other GPUs, the current top-generation will undoubtedly provide better performance than the previous generation. But I think it's also fair to say that in a disaggregated solution, the previous generation of GPUs will also perform much better than if the disaggregated solution were not used, right? In an environment where everyone is trying to extend the life of their hardware and continuously generate valuable tokens, this is an important choice. Is that clear?

Operator: Our next question comes from Vijay Rakesh of Mizuho Securities.

Vijay Rakesh: A couple of quick questions. Regarding the 600 megawatts of contracted capacity and power you mentioned, and the goal of increasing capacity 10x by the end of the year, will this help accelerate the growth momentum in 2027?

Andrew Feldman: Yes, absolutely. I think we are searching for data center capacity globally every day. We are doing this because there is enormous demand for fast inference. The faster we deploy, the faster our revenue grows. This applies not only to our cloud business but also to our customers' on-premise deployments, right? The faster they get data centers, the faster we can ship them hardware. So this is our top priority and something I spend a lot of time on. We now have a full team, and we are quite good at tracking data centers globally. And, Vijay, as you know, the work is far from over after signing a contract. We have people on-site every day. We are in communication with developers and construction companies at every stage, doing our best to ensure the project stays on schedule. Therefore, I think the faster we make progress in this area, the faster we can increase our revenue.

Vijay Rakesh: Understood. Then, regarding partnerships, I know you mentioned hyperscalers, but there is also an emerging group of 'new cloud providers'. They are getting funding, and Wall Street is developing many financing structures. How is this channel developing for you?

Andrew Feldman: Okay. I think early on, new cloud providers were very focused on NVIDIA. I think as the business becomes clearer to investors, some new cloud providers are diversifying and finding themselves less dependent on a single hardware vendor. Therefore, we have a great opportunity in this category. Some new cloud providers are multi-vendor, some are AMD-only, and some are founded by people who own power assets. I think in 2027, this will become an important part of our business.

Operator: Thank you. There are no further questions at this time. This concludes today's call. Thank you for your participation. You may now disconnect.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10