TensorRT-LLM is an open-source library used by companies to optimize large language model inference performance on NVIDIA GPUs through kernel fusion and quantization.