overview
What is TensorRT-LLM?
TensorRT-LLM is an NVIDIA toolkit designed for optimizing Large Language Model (LLM) inference, combining the power of TensorRT kernels with Triton integration. It's the go-to solution for enterprises looking to streamline AI workflows while ensuring high efficiency and performance.
- Supports various LLM architectures including decoder-only and encoder-decoder models.
- Designed for deployment on the latest NVIDIA GPUs for maximum performance.
- Perfect for AI developers, researchers, and production teams.
