DFZQ: FPGA's Expanding Role in AI Servers Drives Increasing Value

Stock News
Jun 05

DFZQ has released a research report highlighting that compared to ASICs, FPGAs offer the advantage of flexible reconfiguration to adapt to algorithmic iterations. When compared to CPUs/GPUs, their low-latency advantage is significant. Within AI servers, FPGAs are utilized for control, interconnect, and computing functions, with their value per server gradually increasing. The path of FPGA adoption in AI servers is expected to evolve from control and interconnect towards computing applications. The main points from DFZQ are outlined below.

Internally, an FPGA primarily consists of three components: input/output elements, programmable logic elements, and programmable routing elements. The basic element of a logic block is the Look-Up Table (LUT), which is essentially static random-access memory (SRAM). Its logic capacity is determined by the number of input signals. Beyond the LUT, a logic block also includes components like flip-flops for implementing sequential circuits and multiplexers. The core of a logic block is the logic element.

Input/output blocks connect the I/O pins with the internal routing elements, enabling the connection between the chip and external circuits while also handling signal driving and matching for inputs and outputs.

Routing elements, which include switch blocks, connection blocks, and routing blocks, serve to interconnect different logic blocks to form the desired functionality. The switch blocks within the routing elements can be programmed and configured to establish arbitrary routing paths.

High Flexibility and Low Latency are Key FPGA Advantages

Compared to ASICs, FPGAs offer greater flexibility. For instance, when downstream algorithms are updated or protocols are upgraded, there is no need to redesign the hardware; simply updating the configuration file suffices. This flexibility is particularly valuable in fields like communications where protocols frequently upgrade and in AI training where algorithms evolve rapidly. The inference frameworks for large AI models undergo significant evolution monthly, each posing new demands on the computational patterns of the underlying hardware. FPGAs can equip new operators through logic reconfiguration without changing the physical hardware, extending the hardware lifecycle from a single model generation to cross-generation reuse. While a primary application of this flexibility in traditional fields has been communication base stations, its adoption in AI is expected to accelerate.

Compared to CPUs/GPUs, FPGAs exhibit lower latency. This is primarily due to three reasons. First, direct hardware translation of algorithms: FPGAs map algorithms directly into hardware circuits via programmable logic units, bypassing the instruction parsing and scheduling steps required by CPUs/GPUs. Second, dataflow-driven execution: computing units process data directly in the order it arrives, eliminating the need for frequent memory access or waiting for global synchronization. Third, parallel computation: FPGAs contain a large number of logic units that can operate simultaneously, supporting both task-level and pipeline-level parallelism.

FPGAs Enable Control, Interconnect, and Computing in AI Servers, Boosting Per-Server Value

For control functions, CPLDs/FPGAs are responsible for system management and power sequencing. As cabinet power consumption increases and the power-up sequencing and overcurrent protection requirements for multi-chip systems (including GPUs, HBM, switch chips, DPUs, etc.) become more stringent, CPLDs/FPGAs, which utilize non-volatile storage media and are operational immediately upon power-up, can complete power rail protection responses within seconds.

For interconnect, taking NVIDIA's Groq3LPX as an example, it needs to handle multiple physical layer interfaces and protocols simultaneously. FPGAs can deploy multiple protocol stacks within a single chip and support dynamic online reconfiguration. The continued coexistence of multiple protocols in the Scale-Up domain, such as NVLink, UALink, SUE, and UB, further strengthens the position of FPGA-based interconnect solutions.

For computing, Altera's Agilex series, which integrates AI Tensor Blocks, supports direct invocation by mainstream inference frameworks. Microsoft's Project Brainwave has already validated that FPGAs can achieve throughput improvements in data center inference scenarios. FPGAs hold advantages in three dimensions: performance-per-watt (efficiency), latency determinism, and functional integration. As the demand for large model inference grows, the path for FPGA penetration in AI servers is expected to iterate from control and interconnect towards computing.

The report concludes with risk warnings, including potential shortfalls in downstream demand, slower-than-expected technology R&D, and intensifying industry competition.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10