BTC $77,726.70 +1.04%
ETH $2,520.69 +0.61%
BNB $723.98 +0.44%
XRP $1.39 +2.67%
SOL $101.39 +1.22%
TRX $0.3394 -0.05%
DOGE $0.0840 +0.73%
ADA $0.2102 +2.90%
BCH $222.10 -0.69%
LINK $11.38 +0.24%
HYPE $79.70 +1.61%
AAVE $126.21 +1.14%
SUI $0.7245 +1.60%
XLM $0.1841 +3.56%
ZEC $1,135.64 +1.35%
AAPL $330.73 -0.55%
AMZN $253.98 -0.53%
GOOGL $336.39 -0.71%
MSFT $493.32 -0.01%
META $639.78 -0.90%
NVDA $213.22 -1.30%
TSLA $358.85 -1.73%
SNDK $1,556.19 -1.71%
INTC $97.15 -2.72%
SPCX $148.24 -0.99%
MU $929.73 -1.45%
AMD $490.81 -3.03%
BTC $77,726.70 +1.04%
ETH $2,520.69 +0.61%
BNB $723.98 +0.44%
XRP $1.39 +2.67%
SOL $101.39 +1.22%
TRX $0.3394 -0.05%
DOGE $0.0840 +0.73%
ADA $0.2102 +2.90%
BCH $222.10 -0.69%
LINK $11.38 +0.24%
HYPE $79.70 +1.61%
AAVE $126.21 +1.14%
SUI $0.7245 +1.60%
XLM $0.1841 +3.56%
ZEC $1,135.64 +1.35%
AAPL $330.73 -0.55%
AMZN $253.98 -0.53%
GOOGL $336.39 -0.71%
MSFT $493.32 -0.01%
META $639.78 -0.90%
NVDA $213.22 -1.30%
TSLA $358.85 -1.73%
SNDK $1,556.19 -1.71%
INTC $97.15 -2.72%
SPCX $148.24 -0.99%
MU $929.73 -1.45%
AMD $490.81 -3.03%

cere

All
Article
Flash

hot_img Cerebras releases the fourth generation AI inference system CS-4: performance doubled, power consumption doubled, more flexible deployment

Cerebras released its fourth-generation AI inference system CS-4 this week, based on the same 5nm WSE-3 wafer, achieving double the performance by doubling the clock frequency and power consumption. A single CS-4 cabinet accommodates 3 wafers (CS-3 has 2), featuring a modular "backpack" design that simplifies manufacturing and deployment, with a TDP of approximately 125 to 135kW. The CS-4 can provide an inference speed of nearly 4000 tokens/second/user, about twice that of the CS-3, and supports decomposed inference with heterogeneous systems such as AMD and AWS Trainium.Cerebras claims that the CS-4 offers about 2000 times the on-chip memory bandwidth of NVIDIA's Rubin (43PB/s), but the 44GB SRAM capacity remains unchanged, and long-context inference still requires multi-wafer stacking. For example, with the DeepSeek V4 Pro (1.6T parameters), approximately 20 systems are needed for a 1M context window, and about 40 systems are required for 256 concurrent users, corresponding to a CAPEX exceeding 20 million USD. Cerebras is collaborating with clients such as OpenAI and plans to achieve approximately double performance improvements each year, aiming for a 20-fold throughput increase by 2027. The "backpack" cabinet design of the CS-4 will continue into the next-generation "Nexus" platform.
app_icon
ChainCatcher Building the Web3 world with innovations.