TensorRT-LLM Releases· tongyuantongyu·· 24 天前精选AI 评分62
NVIDIA TensorRT-LLM 发布 v1.3.0rc26,新增 Rubin SM107 与 NVFP4 KV cache 支持
v1.3.0rc26
AI 导读
NVIDIA TensorRT-LLM 发布 v1.3.0rc26,新增 Rubin SM107 的 GEMM、MoE 与 CuTe DSL 支持,并加入 DeepSeek-V4 Hopper、Qwen3.8-Flash-Next、Kimi K3 on B200 等模型支持。
推荐理由
该版本为 Rubin SM107 与 Blackwell 平台补齐 GEMM、MoE 与 CuTe DSL 支持,并新增 NVFP4 KV cache 与流水线化 KV 传输,适合关注新硬件适配与推理吞吐的部署方评估。
来源:TensorRT-LLM Releases · github.com