vLLM
vllm.ai · ML Frameworks
High-throughput open-source inference and serving engine built around PagedAttention and continuous batching, with an OpenAI-compatible server that scales from one GPU to clusters.
About the ML Frameworks category
The numerical foundations: tensor libraries, autodiff engines, compilers and inference runtimes that every model above them depends on.
vLLM is one of 14 ml frameworks tools indexed on Axiomi. Facts on this page come from the tool's official site and public APIs; prices and features change, so confirm details on the official website before you commit.