BTC $77,609.78 +1.12%
ETH $2,510.46 +0.73%
BNB $722.23 +0.68%
XRP $1.38 +2.84%
SOL $101.31 +1.65%
TRX $0.3400 +0.10%
DOGE $0.0838 +0.68%
ADA $0.2092 +2.56%
BCH $220.76 -1.00%
LINK $11.33 +0.41%
HYPE $79.48 +2.45%
AAVE $125.84 +1.30%
SUI $0.7228 +1.68%
XLM $0.1836 +3.40%
ZEC $1,130.04 +3.40%
AAPL $330.82 -0.36%
AMZN $253.09 -0.69%
GOOGL $335.49 -0.76%
MSFT $491.03 -0.33%
META $639.10 -0.73%
NVDA $211.54 -1.82%
TSLA $358.18 -1.77%
SNDK $1,543.64 -2.23%
INTC $96.26 -3.04%
SPCX $147.31 -1.43%
MU $923.93 -1.52%
AMD $486.50 -3.66%
BTC $77,609.78 +1.12%
ETH $2,510.46 +0.73%
BNB $722.23 +0.68%
XRP $1.38 +2.84%
SOL $101.31 +1.65%
TRX $0.3400 +0.10%
DOGE $0.0838 +0.68%
ADA $0.2092 +2.56%
BCH $220.76 -1.00%
LINK $11.33 +0.41%
HYPE $79.48 +2.45%
AAVE $125.84 +1.30%
SUI $0.7228 +1.68%
XLM $0.1836 +3.40%
ZEC $1,130.04 +3.40%
AAPL $330.82 -0.36%
AMZN $253.09 -0.69%
GOOGL $335.49 -0.76%
MSFT $491.03 -0.33%
META $639.10 -0.73%
NVDA $211.54 -1.82%
TSLA $358.18 -1.77%
SNDK $1,543.64 -2.23%
INTC $96.26 -3.04%
SPCX $147.31 -1.43%
MU $923.93 -1.52%
AMD $486.50 -3.66%

rdw

All
Article
Flash

first_img Google claims that the cost of AI server memory has exceeded 75%, promoting a dual-track strategy for software and hardware

The SEMICON Taiwan 2026 Memory Summit took place on the 1st, where Nikhil Cherian, Senior Director of Supply Chain Infrastructure at Google Cloud under Alphabet, pointed out that with the popularity of multimodal and mixed expert architectures, AI computation has shifted from being power-limited to memory-limited, with high-performance memory accounting for over 75% of the cost of AI server hardware bill of materials. In the face of capacity, bandwidth, and power consumption bottlenecks, Google is breaking through the AI memory bottleneck through a dual-track strategy of hardware offloading for inference and training, and lossless quantization software algorithms.Google adopts an offloading strategy in hardware architecture, launching TPU 8i for low-latency inference and TPU 8t specialized for large-scale training. The TPU 8i is equipped with 288 GB of high-bandwidth memory, with SRAM capacity on the chip increased threefold to 384 MiB, placing dynamic conversation states and key-value caches on the chip itself to achieve zero chip-off latency. The TPU 8t forms a super-large computing cluster with 9600 chips, achieving a shared pool of HBM at a scale of 2 PB, eliminating chip-off data transfer bottlenecks, along with TPU Direct Storage technology.Google has developed the training-free TurboQuant lossless quantization algorithm, compressing the key-value cache of large models from 32 bits to 3 bits, reducing memory usage by six times without loss of accuracy, resulting in an eightfold acceleration in attention computation, and integrating old-generation DRAM technology to extend the lifecycle of components.

first_img DingTalk launches AI office app QwenNote, hardware QwenNote A2 exposed

According to "DuJia," DingTalk is advancing a brand new AI office application QwenNote (Qianwen Listening Note). This application is positioned as an AI personal assistant, integrating real-time voice transcription, summarization, and translation through a combination of software and hardware, and deeply integrating with AI Agent, embedding Agent capabilities into voice input, promoting a shift from simple recording to automated execution. The application supports real-time transcription and bilingual recognition in Chinese and English, as well as language switching, and can generate structured meeting minutes, outlines, and to-do lists, with built-in AI Q&A and quick commands based on listening materials.QwenNote offers a voice memo function, requiring a QR code scan to connect to the recording device. By long-pressing the button on the back of the device, users can record inspirations, and once the recording is complete, it will automatically archive, generate a title and brief summary, and timestamp it. The product also features a stealth protection mode, which, when activated, will physically delete the original audio and only retain the transcribed text to accommodate confidential scenarios. The associated hardware QwenNote A2 has already been showcased, which is part of the Qianwen Listening Note hardware ecosystem, which also includes DingTalk A1, DingTalk A1 Pro, Cleer H1, etc., and users can complete the binding by scanning a QR code.Reports indicate that DingTalk hopes to complement the offline voice collection entry through the integration of software and hardware, forming a closed loop of on-site audio collection, real-time bilingual transcription, AI meeting minutes Q&A, and DingTalk organizational collaboration. Listening materials can be synchronized to DingTalk AI Listening Note and support personal private isolation. On the software side, it continues to embed large models into documents, meetings, and IM scenarios, while on the hardware side, it expands the Qianwen Listening Note product line. Relevant hardware has not yet been widely publicly searched for more formal release information.
app_icon
ChainCatcher Building the Web3 world with innovations.