跳到正文
原文
TensorRT-LLM Releases· tongyuantongyu·· 10 天前精选AI 评分62

TensorRT-LLM 发布 v1.3.0rc28,默认启用 KV-cache manager V2

v1.3.0rc28

AI 导读

NVIDIA TensorRT-LLM 发布 v1.3.0rc28,为 Llama、Llama4 及 Nemotron 多模态模型默认启用 KV-cache manager V2,并移除 TensorRT serve、evaluation、benchmark 等废弃路径。

推荐理由

该版本为 Llama、Llama4 与 Nemotron 多模态模型默认启用 KV-cache manager V2,并新增 FP8 KV-cache 与 NCCL-EP 0.2 低延迟专家并行支持。

来源:TensorRT-LLM Releases · github.com