overview
What is TensorRT-LLM?
TensorRT-LLM is a state-of-the-art solution designed to enhance the performance of Transformer models through optimized kernels and advanced quantization techniques. Whether you’re deploying AI applications in cloud or edge environments, our tool ensures that you can deliver low-latency inference effortlessly.
- Supports various Transformer architectures.
- Designed for both training and serving phases.
- Easily integrates with existing NVIDIA ecosystems.
