跳到正文
原文
TensorRT-LLM Releases· tongyuantongyu·· 24 天前精选AI 评分62

NVIDIA TensorRT-LLM 发布 v1.3.0rc26,新增 Rubin SM107 与 NVFP4 KV cache 支持

v1.3.0rc26

AI 导读

NVIDIA TensorRT-LLM 发布 v1.3.0rc26,新增 Rubin SM107 的 GEMM、MoE 与 CuTe DSL 支持,并加入 DeepSeek-V4 Hopper、Qwen3.8-Flash-Next、Kimi K3 on B200 等模型支持。

推荐理由

该版本为 Rubin SM107 与 Blackwell 平台补齐 GEMM、MoE 与 CuTe DSL 支持,并新增 NVFP4 KV cache 与流水线化 KV 传输,适合关注新硬件适配与推理吞吐的部署方评估。

来源:TensorRT-LLM Releases · github.com