Skip to main content
📖 The AI Tool Bible

AI serverless GPUs

Editorial picks for "serverless gpu".

23 tools

All Fine-tuning →
TA

Together AI

Featured
Fine-tuning · Llama / Mistral / Qwen / DeepSeek and others
8.6

Fine-tune & serve open-weight models (Llama, Mistral, DeepSeek).

Paid· Pay-per-token; fine-tuning per-tokenopen modelsfine-tuning
MO

Modal

Fine-tuning · Infrastructure (any model you can host)
8.7

Serverless GPUs and infra for training & serving ML.

Freemium· $30/mo free credits; pay-as-you-go GPU ratesserverless GPUfine-tuning
RE

Replicate

Fine-tuning · Thousands of community + first-party models
8.5

One-API platform for running and fine-tuning open-source models.

Paid· Pay-per-second of GPU timemodel hostingfine-tuning
CL

ClearML

Agents · Model-agnostic
8.3

End-to-end MLOps and GenAI platform with open-source experiment tracking and enterprise GPU orchestration.

Freemium· Community: $0 · Pro: $15 Per User/Month + Usage · Scale: Custom Quote · Enterprise: Request a Quoteexperiment-trackinggpu-orchestration
RU

RunPod

Fine-tuning · Bring-your-own (any open-weight or custom model)
8.3

On-demand GPU cloud and serverless inference platform built specifically for AI workloads.

Paid· Pod: $7.39/hr · Pod: $4.39/hr · Pod: $5.89/hr · Pod: $1.99/hr · Pod: $3.19/hrllm-fine-tuninggpu-rental
VL

vLLM

Fine-tuning · Multi-model (open-weight LLMs: Llama, Qwen, DeepSeek, Mistral, Gemma, Phi, etc.)
8.3

Open-source high-throughput inference engine for serving LLMs with PagedAttention and continuous batching.

Free· Free and open-source (Apache 2.0); self-hosted infrastructure costs applyllm-servingself-hosted-inference
BE

BentoML

Agents · Multi-model
8.2

Open-source framework and managed platform for serving and scaling AI models in production.

Freemium· OSS free (Apache 2.0); managed Bento cloud has free tier + usage-based pricingmodel-servingllm-inference
CO

CoreWeave

Fine-tuning · DeepSeek
8.2

AI-native GPU cloud built for large-scale training, fine-tuning, and inference on NVIDIA hardware.

Enterprise· NVIDIA GB300 NVL72: Contact sales · NVIDIA GB200 NVL72: $42.00 · NVIDIA HGX B300: Contact sales · NVIDIA HGX B200: $68.80 · NVIDIA RTX PRO 6000 Blackwell Server Edition: $20.00model-trainingfine-tuning
DA

Daytona

Agents
8.2

Secure, isolated sandboxes for running AI-generated code with sub-90ms cold starts.

Freemium· Pay-per-second from $0.000014/sec; $200 free creditagent-sandboxescode-interpreters
SG

SGLang

Fine-tuning · Multi-model (DeepSeek, Qwen, Llama, Mistral, GLM, GPT-OSS)
8.2

Open-source high-throughput inference engine for LLMs and multimodal models with OpenAI-compatible serving.

Free· Free, open-source (Apache 2.0); self-hosted infra cost onlyllm-servingmultimodal-inference
LA

Lambda

Fine-tuning · NVIDIA VR200 NVL72, NVIDIA GB300 NVL72, NVIDIA HGX B200, NVIDIA HGX B300, NVIDIA H100
8.1

On-demand NVIDIA GPU cloud built specifically for training, fine-tuning, and serving large AI models.

Paid· Basic: $10 · Pro: $20 · Enterprise: Contact salesllm-trainingfine-tuning
FA

Fal.ai

Image Generation · Multi-model (Flux, Stable Diffusion, video/audio models)
8.0

Serverless GPU inference platform optimized for fast diffusion and generative media APIs.

Paid· Usage-based; serverless from ~$1.89/GPU-hour, per-output pricing on model APIstext-to-imagetext-to-video
AN

Anyscale

Fine-tuning · Infrastructure (any model)
7.9

Ray-powered platform for training, serving, and scaling LLMs.

Paid· Enterprise / contact salesdistributed trainingRay
FA

Fireworks AI

Fine-tuning · Multi-model (DeepSeek, Qwen, GLM, Kimi, Gemma, Minimax, others)
7.9

Production inference and fine-tuning platform for open-source LLMs, tuned for speed and enterprise economics.

Freemium· Free signup credits; pay-per-token from ~$0.14/M in; enterprise reserved capacity on requestllm-fine-tuningserverless-inference
FE

FedML

Fine-tuning · Bring-your-own (PyTorch, Hugging Face)
7.3

Distributed training, fine-tuning, and serving platform with federated learning roots.

Freemium· Open-source library free; managed GPU usage pay-as-you-gofine-tuningdistributed-training
BE

Beam

Coding · openai/gpt-oss-20b
7.2

Serverless GPU infrastructure for AI workloads with sub-second cold starts and bring-your-own-cloud support.

Freemium· $30 free credit refreshed monthly; usage-based beyond thatgpu-inferenceagent-sandboxes
PG

Paperspace Gradient

Fine-tuning · Bring-your-own (PyTorch, TensorFlow, Hugging Face)
7.2

End-to-end MLOps platform with GPU notebooks, training jobs, and model deployment, now folded into DigitalOcean.

Freemium· Free: $0 · Pro: $8 · Growth: $39 · T0: $0 · T1: $12model-trainingfine-tuning
NR

Neu.ro

Agents · Multi-model
7.1

Infrastructure-agnostic MLOps platform for the full ML/DL lifecycle across public, hybrid, and on-prem clouds.

Enterprise· Contact sales; no public pricingmlopsmodel-training
FO

Forefront

Fine-tuning · Multi-model (Mistral-7B, Mixtral, Phi-2)
7.0

Fine-tune and serve open-source LLMs on your own data without managing GPUs.

Paid· Basic: $20 · Pro: $50 · Enterprise: Contact salesfine-tuningopen-source-llms
VE

Velda

Fine-tuning
6.7

Serverless GPU orchestration that runs AI training and batch jobs without Docker or Kubernetes.

Freemium· Free monthly credits on Velda Cloud; Enterprise contact salesdistributed-trainingbatch-inference
E2

E2B

Agents · LLM-agnostic (works with GPT-4o, Claude, Mistral, Llama, and self-hosted models)

Secure cloud sandboxes for running AI-generated code

Freemium· Hobby: Free · Pro: $150 · Ultimate: Contact us for custom solution with special pricing · Enterprise: Contact us for custom solution with special pricingCode interpreter for chat appsAutonomous coding agents
SC

Scalattice

Agents · Open-weight models including Qwen 2.5 Coder, Llama 3.3 70B, DeepSeek R1, Gemma

Pay less for AI, earn from your GPU.

Paid· Per-token pay-as-you-go on prepaid credits. Example rates per 1M tokens: Qwen-2.5-Coder-7B $0.029 in / $0.077 out; Llama-3.3-70B $0.098 in / $0.306 out. No platform fee for GPU providers to connect.Cheap open-model inference behind an OpenAI-compatible SDKAgent loops on Llama 3.3 70B or DeepSeek R1
X4

X402vps

Agents

Pay-per-hour Docker containers for autonomous AI agents, billed in USDC via the x402 protocol

Paid· Usage-based hourly: ATOM $0.01/hr (0.1 vCPU, 256MB RAM, 2GB disk, no internet); CELL $0.02/hr (0.25 vCPU, 512MB, 5GB, 100MB/hr bandwidth); NODE $0.04/hr (0.5 vCPU, 1GB, 10GB, 500MB/hr); CORE $0.08/hr (1 vCPU, 2GB, 20GB, 2GB/hr). Bandwidth overage $0.01/GB. Exec calls are free. Paid in USDC on Base.Ephemeral web scraping sandboxesLLM code-interpreter backends