NVIDIA 将 Groq 技术整合进机架级产品

Rohan Paul · @rohanpaul_ai · X·2026-08-25 02:54·1天前
AI 导读

NVIDIA 在获得 Groq 非独家授权 8 个月后,首次将 Groq 技术整合进机架级产品,Groq 机架将于今年上线。NVIDIA 将智能体 AI 工作负载拆分:Rubin GPU 负责重计算,Groq 3 LPX 处理延迟敏感的 token 生成,Vera CPU 运行代码与数据处理。相比 Cerebras,Groq 3 LPX 声称带来 4 倍响应速度提升,对长任务影响尤为显著。

Rohan Paul@rohanpaul_ai
49AI 编辑部评分,满分 100

NVIDIA 将 Groq 技术整合进机架级产品

2026-08-25 02:54· 1天前
AI 导读

NVIDIA 在获得 Groq 非独家授权 8 个月后,首次将 Groq 技术整合进机架级产品,Groq 机架将于今年上线。NVIDIA 将智能体 AI 工作负载拆分:Rubin GPU 负责重计算,Groq 3 LPX 处理延迟敏感的 token 生成,Vera CPU 运行代码与数据处理。相比 Cerebras,Groq 3 LPX 声称带来 4 倍响应速度提升,对长任务影响尤为显著。

8 months after NVIDIA’s non-exclusive Groq licensing arrangement, Groq technology finally is appearing inside a rack-scale NVIDIA product.

Groq racks will be online this year.

So NVIDIA is splitting agentic AI work across specialized processors instead of treating the GPU as the whole machine.

Rubin GPUs handle heavy model computation, Groq 3 LPX targets latency-sensitive token generation, and Vera CPUs will run code, tools, data processing and simulation around the model.

Agents make token-generation latency far more consequential than it is in ordinary chat because later steps often wait for earlier ones to finish.

Nvidia's claimed 4x responsiveness (Nvidia's Groq 3 LPX vs. Cerebras' inference platform) improvement can therefore compound across a long task rather than just make individual responses appear faster. One big reason inference hardware will be fragmenting by workload.