vLLM

ModelOpen Source

High-throughput and memory-efficient inference and serving engine for LLMs. PagedAttention + continuous batching.

Visit website View on GitHub
Built with
PythonCUDAPyTorch
Share this part