AI for serving LLMs
High-throughput inference servers.
| # | Tool | Category | Pricing | Visit |
|---|---|---|---|---|
| 1 | vLLM High-throughput LLM inference server with continuous batching and PagedAttention | LLM Frameworks | Free | Visit |
| 2 | SGLang Fast LLM serving framework — efficient serving for LLMs and vision-language models | LLM Frameworks | Free | Visit |
| 3 | LoRAX Multi-LoRA inference server — serve hundreds of fine-tuned adapters on a single GPU | LLM Frameworks | Free | Visit |
| 4 | Modal Serverless Python and GPU cloud — deploy ML models, batch jobs, and APIs with decorator syntax | LLM Frameworks | Freemium | Visit |
| 5 | RunPod GPU cloud for AI — rent A100/H100 GPUs for training, inference, and fine-tuning | LLM Frameworks | Freemium | Visit |
| 6 | Baseten ML model deployment platform — deploy any model as a production API in minutes | LLM Frameworks | Freemium | Visit |