← Back to Repositories Hub Language License vllm-project/vllm A high-throughput and memory-efficient LLM serving engine. Maintainer: vllm-project Last Reviewed: 2026-06-22 Next Review: 2026-12-31 Editorial Owner: AI Delivery Hub Editorial Description & Purpose A high-throughput and memory-efficient LLM serving engine. Architecture Summary Built as a modular framework in Python, leveraging asynchronous processing and event loops to orchestrate LLM requests. Core Use Cases Deploying high-concurrency model endpoints, PagedAttention optimizations. Alternatives Evaluated Alternative open-source orchestration packages 💻 View on GitHub 📖 View Documentation