# NVIDIA 发布 Nemotron 3.5 Lightning 长时运行智能体模型

- 来源：Rohan Paul (@rohanpaul_ai)
- 发布时间：2026-08-11 23:22
- AIHOT 分数：46
- AIHOT 链接：https://aihot.virxact.com/items/cmsou3ijw06qmrohdffqxlepd
- 原文链接：https://x.com/rohanpaul_ai/status/2087198053036167532

## AI 摘要

NVIDIA 发布 Nemotron 3.5 Lightning，专为长时运行智能体设计，30B 总参数/3B 激活参数，支持 1M token 上下文，可商用。支持单卡部署（DGX Spark 或 H100），并加入多 token 预测与 NVFP4 量化，输出速度最高提升 4 倍。PinchBench 上准确率 86%，完成速度比 Qwen3.6 35B 快 30%。

## 正文

NVIDIA released Nemotron 3.5 Lightning for long-running agent execution.

• 30B total / 3B active parameters, with up to 1M tokens of context.
Ready for commercial use.

• Local deployment: NVIDIA ships BF16 and much smaller NVFP4 checkpoints, lists 1× DGX Spark or 1× H100 for single-GPU deployment, and also lists RTX 5090 among supported hardware.

• NVIDIA then adds multi-token prediction, DSpark and DFlash speculative decoding, plus NVFP4 quantization to push generation speed further.

• NVIDIA claims up to 4× output speed; on PinchBench, 86% accuracy and 30% faster completion than Qwen3.6 35B at similar accuracy.

### 引用推文

> Jensen Huang：Lightning strikes for continuous and long-run agents! Nemotron 3.5 Lightning is smart, fast, efficient and open.
