BTC $77,184.58 -0.16%
ETH $2,521.75 -0.36%
BNB $726.39 +0.08%
XRP $1.37 +0.29%
SOL $101.45 -1.01%
TRX $0.3398 +0.43%
DOGE $0.0848 +0.60%
ADA $0.2078 +1.07%
BCH $225.35 -1.60%
LINK $11.51 -0.55%
HYPE $79.87 -0.85%
AAVE $125.93 +0.79%
SUI $0.7226 -0.75%
XLM $0.1800 +0.81%
ZEC $1,123.84 -4.95%
AAPL $333.13 +0.16%
AMZN $256.77 -0.07%
GOOGL $340.95 +0.69%
MSFT $495.00 -0.06%
META $648.72 +0.25%
NVDA $218.40 -0.02%
TSLA $367.12 +0.52%
SNDK $1,622.67 -0.50%
INTC $102.02 -0.86%
SPCX $150.13 -0.54%
MU $966.03 -0.87%
AMD $515.91 +0.18%
BTC $77,184.58 -0.16%
ETH $2,521.75 -0.36%
BNB $726.39 +0.08%
XRP $1.37 +0.29%
SOL $101.45 -1.01%
TRX $0.3398 +0.43%
DOGE $0.0848 +0.60%
ADA $0.2078 +1.07%
BCH $225.35 -1.60%
LINK $11.51 -0.55%
HYPE $79.87 -0.85%
AAVE $125.93 +0.79%
SUI $0.7226 -0.75%
XLM $0.1800 +0.81%
ZEC $1,123.84 -4.95%
AAPL $333.13 +0.16%
AMZN $256.77 -0.07%
GOOGL $340.95 +0.69%
MSFT $495.00 -0.06%
META $648.72 +0.25%
NVDA $218.40 -0.02%
TSLA $367.12 +0.52%
SNDK $1,622.67 -0.50%
INTC $102.02 -0.86%
SPCX $150.13 -0.54%
MU $966.03 -0.87%
AMD $515.91 +0.18%
hot_img

Cerebras releases the fourth generation AI inference system CS-4: performance doubled, power consumption doubled, more flexible deployment

2026-08-19 10:38:29

Cerebras released its fourth-generation AI inference system CS-4 this week, based on the same 5nm WSE-3 wafer, achieving double the performance by doubling the clock frequency and power consumption. A single CS-4 cabinet accommodates 3 wafers (CS-3 has 2), featuring a modular "backpack" design that simplifies manufacturing and deployment, with a TDP of approximately 125 to 135kW. The CS-4 can provide an inference speed of nearly 4000 tokens/second/user, about twice that of the CS-3, and supports decomposed inference with heterogeneous systems such as AMD and AWS Trainium.

Cerebras claims that the CS-4 offers about 2000 times the on-chip memory bandwidth of NVIDIA's Rubin (43PB/s), but the 44GB SRAM capacity remains unchanged, and long-context inference still requires multi-wafer stacking. For example, with the DeepSeek V4 Pro (1.6T parameters), approximately 20 systems are needed for a 1M context window, and about 40 systems are required for 256 concurrent users, corresponding to a CAPEX exceeding 20 million USD. Cerebras is collaborating with clients such as OpenAI and plans to achieve approximately double performance improvements each year, aiming for a 20-fold throughput increase by 2027. The "backpack" cabinet design of the CS-4 will continue into the next-generation "Nexus" platform.

app_icon
ChainCatcher Building the Web3 world with innovations.