The Decoder:AI News(RSS)
67AI 编辑部评分,满分 100

Nvidia 发布 Nemotron 3.5 Lightning:以速度优先于极致智能

2026-08-11 23:07· 1天前· Matthias Bastian
AI 导读

Nvidia 发布 Nemotron 3.5 Lightning,这是 Nemotron 3.5 系列首款模型,总参数 316 亿、激活参数仅 36 亿。该模型在 Artificial Analysis 智能指数上得 24 分,与 OpenAI gpt-oss-120b 持平,但推理速度近每秒 670 tokens,为对比中最高。

Image description

Nvidia's new Nemotron 3.5 Lightning is a compact open-weights model that matches OpenAI's gpt-oss-120b on intelligence benchmarks with a quarter of the parameters while delivering the fastest inference speeds in its class.

Nvidia has released Nemotron 3.5 Lightning, the first model in its new Nemotron 3.5 lineup. The model directly succeeds the Nemotron 3 Nano 30B A3B and keeps its hybrid Mamba-Transformer architecture, with 31.6 billion total parameters and only 3.6 billion active at any given time.

According to the independent benchmarking platform Artificial Analysis, the model scores 24 on the Intelligence Index, a nine-point jump from its predecessor (15). That puts Lightning on par with OpenAI's gpt-oss-120b (24) and just behind Nvidia's own Nemotron 3 Super (26), which is about four times larger. The smartest small models in the same size class, like Qwen3.6 35B A3B (32) and Meta's new Muse Glimmer (35), still hold a clear lead.

Nemotron 3.5 Lightning scores 24 on the Artificial Analysis Intelligence Index, tying with gpt-oss-120b. | Image: Artificial Analysis

Nvidia is targeting a different spot on the efficiency frontier with Lightning. In pre-release tests using the final NVFP4 weights, the model hits nearly 670 tokens per second, the highest measured throughput among all compared models and almost twice as fast as Google's Gemini 3.5 Flash-Lite (386 tokens/s). A task from the Intelligence Index takes about 0.5 minutes to complete, while Qwen3.6 35B A3B needs around 3.5 minutes, and Gemma 4 31B takes roughly 5.8 minutes.

At 669 tokens per second, Lightning is the fastest model in the comparison. Despite generating a similar number of tokens per task as its predecessor Nemotron 3 Nano, it delivers much better results. | Image: Artificial Analysis

Proprietary models still dominate the overall efficiency frontier. Gemini 3.5 Flash-Lite scores 37 on the Intelligence Index with a similar time per task, and GPT-5.6 Luna (max) reaches 52 points in under two minutes.

Agentic benchmarks show the biggest gains

The biggest improvements show up in agentic benchmarks, according to Artificial Analysis. On GDPval-AA v2, Lightning reaches an Elo rating of 824, a 334-point gain over Nemotron 3 Nano. That beats both gpt-oss-120b (800) and the larger Nemotron 3 Super (698). On Terminal-Bench v2.1, the score jumps from 7 to 24.3 percent, nearly matching gpt-oss-120b at 26.2 percent.

On agentic benchmarks, Lightning surpasses both gpt-oss-120b and the larger Nemotron 3 Super with an Elo rating of 824. | Image: Artificial Analysis

Nvidia ships the model under the permissive OpenMDW-1.1 license, positioning it as a high-throughput workhorse for agent-based pipelines. Artificial Analysis reports that Nvidia worked with partners like CodeRabbit and Harvey on post-training to boost performance in specific domains.

Availability

Nvidia provides the model in both BF16 and NVFP4 weights. The NVFP4 variant also scores 24 on the Intelligence Index with minimal quality loss compared to the higher-precision version, according to Artificial Analysis. The reasoning model handles text only and supports a context window of one million tokens. Weights are available now, and serverless inference is offered by DeepInfraFireworksFriendliAICoreWeave, GMI Cloud, Nebius, and Crusoe, among others.

Nvidia's push for efficiency over size isn't new. In a widely discussed paper last year, its researchers argued that models under 10 billion parameters can handle most agent workloads as well as 70- to 175-billion-parameter models at one-tenth to one-thirtieth the cost. Nemotron 3.5 Lightning has 31.6 billion parameters but activates only 3.6 billion per step, putting it in the same lightweight class. At nearly 670 tokens per second, it also beats gpt-oss-120b and the larger Nemotron 3 Super on agentic benchmarks, making it the clearest product-level proof of that thesis yet.

Nvidia

来源:The Decoder:AI News(RSS) · the-decoder.com

事件后续 · 5查看事件全部 →

Nvidia 发布 Nemotron 3.5 Lightning:以速度优先于极致智能

The Decoder:AI News(RSS)·2026-08-11 23:07·1天前·Matthias Bastian
AI 导读

Nvidia 发布 Nemotron 3.5 Lightning,这是 Nemotron 3.5 系列首款模型,总参数 316 亿、激活参数仅 36 亿。该模型在 Artificial Analysis 智能指数上得 24 分,与 OpenAI gpt-oss-120b 持平,但推理速度近每秒 670 tokens,为对比中最高。

原文 · 保持原样,未翻译
Image description

Nvidia's new Nemotron 3.5 Lightning is a compact open-weights model that matches OpenAI's gpt-oss-120b on intelligence benchmarks with a quarter of the parameters while delivering the fastest inference speeds in its class.

Nvidia has released Nemotron 3.5 Lightning, the first model in its new Nemotron 3.5 lineup. The model directly succeeds the Nemotron 3 Nano 30B A3B and keeps its hybrid Mamba-Transformer architecture, with 31.6 billion total parameters and only 3.6 billion active at any given time.

According to the independent benchmarking platform Artificial Analysis, the model scores 24 on the Intelligence Index, a nine-point jump from its predecessor (15). That puts Lightning on par with OpenAI's gpt-oss-120b (24) and just behind Nvidia's own Nemotron 3 Super (26), which is about four times larger. The smartest small models in the same size class, like Qwen3.6 35B A3B (32) and Meta's new Muse Glimmer (35), still hold a clear lead.

Nemotron 3.5 Lightning scores 24 on the Artificial Analysis Intelligence Index, tying with gpt-oss-120b. | Image: Artificial Analysis

Nvidia is targeting a different spot on the efficiency frontier with Lightning. In pre-release tests using the final NVFP4 weights, the model hits nearly 670 tokens per second, the highest measured throughput among all compared models and almost twice as fast as Google's Gemini 3.5 Flash-Lite (386 tokens/s). A task from the Intelligence Index takes about 0.5 minutes to complete, while Qwen3.6 35B A3B needs around 3.5 minutes, and Gemma 4 31B takes roughly 5.8 minutes.

At 669 tokens per second, Lightning is the fastest model in the comparison. Despite generating a similar number of tokens per task as its predecessor Nemotron 3 Nano, it delivers much better results. | Image: Artificial Analysis

Proprietary models still dominate the overall efficiency frontier. Gemini 3.5 Flash-Lite scores 37 on the Intelligence Index with a similar time per task, and GPT-5.6 Luna (max) reaches 52 points in under two minutes.

Agentic benchmarks show the biggest gains

The biggest improvements show up in agentic benchmarks, according to Artificial Analysis. On GDPval-AA v2, Lightning reaches an Elo rating of 824, a 334-point gain over Nemotron 3 Nano. That beats both gpt-oss-120b (800) and the larger Nemotron 3 Super (698). On Terminal-Bench v2.1, the score jumps from 7 to 24.3 percent, nearly matching gpt-oss-120b at 26.2 percent.

On agentic benchmarks, Lightning surpasses both gpt-oss-120b and the larger Nemotron 3 Super with an Elo rating of 824. | Image: Artificial Analysis

Nvidia ships the model under the permissive OpenMDW-1.1 license, positioning it as a high-throughput workhorse for agent-based pipelines. Artificial Analysis reports that Nvidia worked with partners like CodeRabbit and Harvey on post-training to boost performance in specific domains.

Availability

Nvidia provides the model in both BF16 and NVFP4 weights. The NVFP4 variant also scores 24 on the Intelligence Index with minimal quality loss compared to the higher-precision version, according to Artificial Analysis. The reasoning model handles text only and supports a context window of one million tokens. Weights are available now, and serverless inference is offered by DeepInfraFireworksFriendliAICoreWeave, GMI Cloud, Nebius, and Crusoe, among others.

Nvidia's push for efficiency over size isn't new. In a widely discussed paper last year, its researchers argued that models under 10 billion parameters can handle most agent workloads as well as 70- to 175-billion-parameter models at one-tenth to one-thirtieth the cost. Nemotron 3.5 Lightning has 31.6 billion parameters but activates only 3.6 billion per step, putting it in the same lightweight class. At nearly 670 tokens per second, it also beats gpt-oss-120b and the larger Nemotron 3 Super on agentic benchmarks, making it the clearest product-level proof of that thesis yet.

Nvidia

来源:The Decoder:AI News(RSS)· the-decoder.com

事件后续 · 5查看事件全部 →