Skip to main content

AI for serving LLMs

High-throughput inference servers.

#ToolCategoryPricingVisit
1vLLM

High-throughput LLM inference server with continuous batching and PagedAttention

LLM FrameworksFreeVisit
2SGLang

Fast LLM serving framework — efficient serving for LLMs and vision-language models

LLM FrameworksFreeVisit
3LoRAX

Multi-LoRA inference server — serve hundreds of fine-tuned adapters on a single GPU

LLM FrameworksFreeVisit
4Modal

Serverless Python and GPU cloud — deploy ML models, batch jobs, and APIs with decorator syntax

LLM FrameworksFreemiumVisit
5RunPod

GPU cloud for AI — rent A100/H100 GPUs for training, inference, and fine-tuning

LLM FrameworksFreemiumVisit
6Baseten

ML model deployment platform — deploy any model as a production API in minutes

LLM FrameworksFreemiumVisit