tensorrt-llm

0

Описание

TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and build TensorRT engines that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM also contains components to create Python and C++ runtimes that execute those TensorRT engines.

Языки

  • C++99,5%
  • Python0,4%
  • Cuda0,1%

Ashwinkumar J S

2 года назад
2 года назад
3 года назад
3 года назад
3 года назад
README.md