AI models
vLLM
High-throughput model serving for GPUs.
open source self-hostable vLLM project · ?The production engine for serving open models on your own GPUs.
Website ↗ Source ↗ Plan “AI models” with it
How you can run it
| Way | Where the data lives | Price basis | Verdict |
|---|---|---|---|
| Host it myself | your own server or hardware | free | green — You run it yourself: no data leaves your own infrastructure. |
Facts
| Company / project | vLLM project (jurisdiction unknown) |
|---|---|
| Licence | Apache-2.0 |
| Self-hosting footprint | 8192 MB RAM · 4 vCPU · 50 GB — needs a GPU server |
| Commonly replaces | OpenAI API |
| Provenance | curated, verified 2026-09-04 |
Machine-readable: /api/tools/vllm.json. Something wrong? Tell us.