Grouping multiple requests into one forward pass to keep GPU utilisation high. Continuous batching lets requests join and leave mid-generation.
inference
Batching (Inference)
Related terms
Tools that implement Batching (Inference)
Together AI
FeaturedFine-tuning · Llama / Mistral / Qwen / DeepSeek and others
8.6
Fine-tune & serve open-weight models (Llama, Mistral, DeepSeek).
Paid· Pay-per-token; fine-tuning per-tokenopen modelsfine-tuning
vLLM
Fine-tuning · Multi-model (open-weight LLMs: Llama, Qwen, DeepSeek, Mistral, Gemma, Phi, etc.)
8.3
Open-source high-throughput inference engine for serving LLMs with PagedAttention and continuous batching.
Free· Free and open-source (Apache 2.0); self-hosted infrastructure costs applyllm-servingself-hosted-inference
Fireworks AI
Fine-tuning · Multi-model (DeepSeek, Qwen, GLM, Kimi, Gemma, Minimax, others)
7.9
Production inference and fine-tuning platform for open-source LLMs, tuned for speed and enterprise economics.
Freemium· Free signup credits; pay-per-token from ~$0.14/M in; enterprise reserved capacity on requestllm-fine-tuningserverless-inference