AB
AiBoss
News

Nvidia says Groq 3 LPX is in full production and will be used to scale Vera Rubin inference.

NVIDIA Groq 3 LPX has entered full production and is positioned as a low-latency token generation extension for Vera Rubin NVL72. Nebius has become the first AI cloud service provider to adopt this platform.

NVIDIA announced that the NVIDIA Groq 3 LPX is now in full production.The platform works in conjunction with Vera Rubin NVL72, with the Rubin GPU handling large-scale contexts and LPX handling latency-sensitive token generation tasks.

Nvidia claims that rack-mount deployments can include up to 256 LP30 accelerators, and Nebius is the first AI cloud service provider to adopt this platform. The performance comparisons in the announcement are based on specific models and test conditions.

refer to:NVIDIA Announcement