Skip to main content
📖 The AI Tool Bible

AI serverless GPUs

Editorial picks for "serverless gpu".

43 tools

All Fine-tuning
Modal preview image
Modal logo

Modal

Fine-tuning · Infrastructure (any model you can host)
8.7

Serverless GPUs and infra for training & serving ML.

Freemium· $30/mo free credits; pay-as-you-go GPU ratesserverless GPUfine-tuning
CoreWeave preview image
CoreWeave logo

CoreWeave

Fine-tuning · DeepSeek
8.2

AI-native GPU cloud built for large-scale training, fine-tuning, and inference on NVIDIA hardware.

Enterprise· NVIDIA GB300 NVL72: Contact sales · NVIDIA GB200 NVL72: $42.00 · NVIDIA HGX B300: Contact sales · NVIDIA HGX B200: $68.80 · NVIDIA RTX PRO 6000 Blackwell Server Edition: $20.00model-trainingfine-tuning
Forefront preview image
Forefront logo

Forefront

Fine-tuning · Multi-model (Mistral-7B, Mixtral, Phi-2)
7.0

Fine-tune and serve open-source LLMs on your own data without managing GPUs.

Paid· Basic: $20 · Pro: $50 · Enterprise: Contact salesfine-tuningopen-source-llms
Together AI preview image
Together AI logo

Together AI

Featured
Fine-tuning · Llama / Mistral / Qwen / DeepSeek and others
8.6

Fine-tune & serve open-weight models (Llama, Mistral, DeepSeek).

Paid· Pay-per-token; fine-tuning per-tokenopen modelsfine-tuning
Stable Diffusion preview image
Stable Diffusion logo

Stable Diffusion

Image Generation · SD 3.5 / SDXL
8.8

Open-source image generation — run anywhere, fine-tune anything.

Free· Free open weights; optional Stability APIlocalfine-tuning
Llama 3 preview image
Llama 3 logo

Llama 3

Writing · Llama 3 / 3.1 (8B, 70B, 405B)
8.3

Meta's open-weights LLM family that put serious frontier-adjacent models in everyone's hands.

Free· Weights free under Meta Llama Community License; inference cost via self-hosting or 3rd-party providerschatlong-context reasoning
LM Studio preview image
LM Studio logo

LM Studio

Agents · Multi-model (gpt-oss, Qwen3, Gemma, DeepSeek-R1, Llama, others)
8.3

Desktop app for discovering, downloading, and running open-weight LLMs locally with an OpenAI-compatible server.

Freemium· Free: $0local-llm-inferenceprivate-chat
RunPod preview image
RunPod logo

RunPod

Fine-tuning · Bring-your-own (any open-weight or custom model)
8.3

On-demand GPU cloud and serverless inference platform built specifically for AI workloads.

Paid· Pod: $7.39/hr · Pod: $4.39/hr · Pod: $5.89/hr · Pod: $1.99/hr · Pod: $3.19/hrllm-fine-tuninggpu-rental
vLLM preview image
vLLM logo

vLLM

Fine-tuning · Multi-model (open-weight LLMs: Llama, Qwen, DeepSeek, Mistral, Gemma, Phi, etc.)
8.3

Open-source high-throughput inference engine for serving LLMs with PagedAttention and continuous batching.

Free· Free and open-source (Apache 2.0); self-hosted infrastructure costs applyllm-servingself-hosted-inference
Daytona preview image
Daytona logo

Daytona

Agents
8.2

Secure, isolated sandboxes for running AI-generated code with sub-90ms cold starts.

Freemium· Pay-per-second from $0.000014/sec; $200 free creditagent-sandboxescode-interpreters
Jan preview image
Jan logo

Jan

Writing · Multi-model (local open-weights + OpenAI/Claude/Gemini via API)
8.2

Open-source desktop ChatGPT alternative that runs local LLMs and routes to cloud providers from one app.

Free· Free and open source; bring-your-own keys for cloud modelslocal-llm-chatprivate-ai-assistant
SGLang preview image
SGLang logo

SGLang

Fine-tuning · Multi-model (DeepSeek, Qwen, Llama, Mistral, GLM, GPT-OSS)
8.2

Open-source high-throughput inference engine for LLMs and multimodal models with OpenAI-compatible serving.

Free· Free, open-source (Apache 2.0); self-hosted infra cost onlyllm-servingmultimodal-inference
Lambda preview image
Lambda logo

Lambda

Fine-tuning · NVIDIA VR200 NVL72, NVIDIA GB300 NVL72, NVIDIA HGX B200, NVIDIA HGX B300, NVIDIA H100
8.1

On-demand NVIDIA GPU cloud built specifically for training, fine-tuning, and serving large AI models.

Paid· Basic: $10 · Pro: $20 · Enterprise: Contact salesllm-trainingfine-tuning
Ray Tune preview image
Ray Tune logo

Ray Tune

Fine-tuning
8.1

Open-source Python library for distributed hyperparameter tuning at any scale.

Free· Open-source (Apache 2.0); managed via Anyscale offers a $100 starting credithyperparameter-tuningdistributed-training
Together AI Fine-tuning preview image
Together AI Fine-tuning logo

Together AI Fine-tuning

Fine-tuning · Multi-model (any Hugging Face open-source model)
8.1

Managed fine-tuning platform for open-source LLMs and vision models with LoRA, full fine-tuning, and RL support.

Paid· Usage-based; cost estimator in-product, no public price listllm-fine-tuningvision-fine-tuning
Fal.ai preview image
Fal.ai logo

Fal.ai

Image Generation · Multi-model (Flux, Stable Diffusion, video/audio models)
8.0

Serverless GPU inference platform optimized for fast diffusion and generative media APIs.

Paid· Usage-based; serverless from ~$1.89/GPU-hour, per-output pricing on model APIstext-to-imagetext-to-video
Genmo preview image
Genmo logo

Genmo

Video · Mochi 1
8.0

Open-source text-to-video model (Mochi 1) with a hosted playground for turning prompts into short clips.

Freemium· Free: $0/month · Lite: Loading... · Standard: Loading...text-to-videogenerative-video
Fireworks AI preview image
Fireworks AI logo

Fireworks AI

Fine-tuning · Multi-model (DeepSeek, Qwen, GLM, Kimi, Gemma, Minimax, others)
7.9

Production inference and fine-tuning platform for open-source LLMs, tuned for speed and enterprise economics.

Freemium· Free signup credits; pay-per-token from ~$0.14/M in; enterprise reserved capacity on requestllm-fine-tuningserverless-inference
Langchain-Chatchat preview image
Langchain-Chatchat logo

Langchain-Chatchat

RAG · Multi-model (GLM-4, Qwen2, Llama 3, etc. via Xinference/Ollama/LocalAI/FastChat)
7.4

Self-hostable RAG and agent framework that wires LangChain to any local open-source LLM and a knowledge base.

Free· Apache-2.0 open source; self-hosted, infra costs onlyprivate-knowledge-baseoffline-rag
CogVideoX preview image
CogVideoX logo

CogVideoX

Video · CogVideoX / CogVideoX1.5 (diffusion transformer)
7.3

Open-source text-to-video and image-to-video diffusion transformer from Zhipu AI, runnable on consumer GPUs.

Free· Open-source weights; commercial API via bigmodel.cntext-to-videoimage-to-video
DeepSeek preview image
DeepSeek logo

DeepSeek

Coding · DeepSeek R1, V3, V2, Coder V2, VL
7.3

Chinese AI lab shipping open-weight reasoning models that punch well above their API price.

Freemium· deepseek-v4-flash: 0.02元 (缓存命中) / 1元 (缓存未命中) / 2元 (输出) · deepseek-v4-pro: 0.025元 (缓存命中) / 3元 (缓存未命中) / 6元 (输出)reasoningcode-generation
Dia preview image
Dia logo

Dia

Audio · Dia-1.6B
7.3

Open-weights 1.6B text-to-dialogue model that generates ultra-realistic multi-speaker conversations in one pass.

Free· Free, open weights (Apache 2.0); hosted larger version waitlisteddialogue-generationvoice-cloning
FedML preview image
FedML logo

FedML

Fine-tuning · Bring-your-own (PyTorch, Hugging Face)
7.3

Distributed training, fine-tuning, and serving platform with federated learning roots.

Freemium· Open-source library free; managed GPU usage pay-as-you-gofine-tuningdistributed-training
Beam preview image
Beam logo

Beam

Coding · openai/gpt-oss-20b
7.2

Serverless GPU infrastructure for AI workloads with sub-second cold starts and bring-your-own-cloud support.

Freemium· $30 free credit refreshed monthly; usage-based beyond thatgpu-inferenceagent-sandboxes
LLaMA Factory preview image
LLaMA Factory logo

LLaMA Factory

Fine-tuning · Multi-model (LLaMA, Mistral, Qwen, Gemma, Phi, LLaVA, ChatGLM, Yi)
7.2

Open-source, no-code WebUI for fine-tuning 100+ open LLMs with LoRA, QLoRA, DPO, and PPO.

Free· Free, open-source (Apache-2.0); self-hostedlora-fine-tuningqlora
Paperspace Gradient preview image
Paperspace Gradient logo

Paperspace Gradient

Fine-tuning · Bring-your-own (PyTorch, TensorFlow, Hugging Face)
7.2

End-to-end MLOps platform with GPU notebooks, training jobs, and model deployment, now folded into DigitalOcean.

Freemium· Free: $0 · Pro: $8 · Growth: $39 · T0: $0 · T1: $12model-trainingfine-tuning
BGE (BAAI General Embedding) preview image
BGE (BAAI General Embedding) logo

BGE (BAAI General Embedding)

RAG · BGE / bge-m3 / bge-reranker
7.1

Open-source embedding and reranker models from BAAI that anchor a huge share of production RAG stacks.

Free· Free, open-source (MIT-style license); self-hosted inference cost onlysemantic-searchrag-retrieval
Groq preview image
Groq logo

Groq

Coding · Multi-model (Llama, Mixtral, Gemma, Qwen, Whisper)
7.1

Custom-silicon LPU inference platform serving open models at GPU-trouncing latency via an OpenAI-compatible API.

Freemium· Free API key with rate limits; per-token paid tiers; enterprise contractslow-latency inferencevoice agents
Iguazio preview image
Iguazio logo

Iguazio

Agents · Multi-model
7.1

Enterprise MLOps and GenAI platform for taking models from notebook to production at scale.

Enterprise· Contact sales; free trial availablemlopsllm-fine-tuning
PostgresML preview image
PostgresML logo

PostgresML

RAG · Multi-model (Llama, Mistral, open-source embeddings)
7.1

PostgreSQL extension that runs embeddings, vector search, and LLM inference inside your database.

Freemium· Serverless: From $7.50 per query hour · Dedicated: From $0.60 per instance hour · Enterprise: Custom pricingvector-searchrag
AI Horde (Stable Horde) preview image
AI Horde (Stable Horde) logo

AI Horde (Stable Horde)

Image Generation · Stable Diffusion (SD 1.5, SDXL, community checkpoints) + open LLMs
7.0

Crowdsourced, volunteer-run cluster for free Stable Diffusion image and LLM text generation.

Free· Free; optional Kudos earned by donating GPU time or purchased via Patreontext-to-imageimg2img
ONNX preview image
ONNX logo

ONNX

Fine-tuning
7.0

Open standard for representing and exchanging machine learning models across frameworks and runtimes.

Free· Free and open source (Apache-2.0); Linux Foundation AI projectmodel-interchangeedge-deployment
Apache SINGA preview image
Apache SINGA logo

Apache SINGA

Fine-tuning
6.9

Apache-licensed distributed deep learning library focused on scalable training across GPUs and nodes.

Free· Free, Apache 2.0 licenseddistributed trainingdeep learning research
WhisperAPI preview image
WhisperAPI logo

WhisperAPI

Audio · OpenAI Whisper
6.9

Hosted OpenAI Whisper transcription with a pay-as-you-go API and drop-in web dashboard.

Paid· 20 API Credits: $5 · 100 API Credits: $20 · 200 API Credits: $30 · Custom: $100.00audio-transcriptionvideo-subtitles
Velda preview image
Velda logo

Velda

Fine-tuning
6.7

Serverless GPU orchestration that runs AI training and batch jobs without Docker or Kubernetes.

Freemium· Free monthly credits on Velda Cloud; Enterprise contact salesdistributed-trainingbatch-inference
Cloudflare MCP Server logo

Cloudflare MCP Server

MCP Servers

Official suite of remote MCP servers that let Claude, Cursor, and other agents read and control your Cloudflare account.

Freemium· Free: $0 USD · Team: $4 USD per user/month · Enterprise: $21 USD per user/monthDebugging Workers with live logs from chatDeploying and configuring Workers bindings
Colossal-AI preview image
Colossal-AI logo

Colossal-AI

Fine-tuning · Framework-agnostic; used with LLaMA, GPT, Stable Diffusion, ViT, and other PyTorch-based open-weight models

Making large AI models cheaper, faster, and more accessible through distributed training

Free· Open-source (Apache 2.0). Enterprise support, consulting, and managed training services available from HPC-AI Technology on request.LLM pretraining across multi-node GPU clustersFull-parameter and LoRA fine-tuning of open-weight LLMs
HiDream preview image
HiDream logo

HiDream

Image Generation · HiDream-I1 (17B diffusion transformer) — Full, Dev, Fast variants

Open-source 17B image model with frontier quality, hosted on Vivago and free to self-host under MIT.

Freemium· Free tier on vivago.ai; Basic $9.90/mo, Plus $29.90/mo, Pro $69.90/mo (monthly). Annual saves up to 50% (Basic $7.90/mo, Plus $19.90/mo, Pro $59.90/mo). Third-party API pricing (e.g. Fal): ~$1 per 20 Full runs, ~$1 per 33 Dev runs, ~$1 per 100 Fast runs.text-to-image generationconcept art and illustration
LLM GPU Checker (KO) preview image
LLM GPU Checker (KO) logo

LLM GPU Checker (KO)

Evaluation · Catalog covers open models on Hugging Face (Llama, Qwen, Mistral, Gemma, etc.)

Match LLMs to GPUs and plan multi-model AI stacks by VRAM, bandwidth and precision.

Free· Free (open web tool hosted on GitHub Pages).GPU sizing for self-hosted LLMsMulti-GPU RAG stack planning
LTX Video preview image
LTX Video logo

LTX Video

Video · LTX-Video (LTX-2), in-house DiT — 13B full, 13B distilled, 2B distilled, FP8 quantized variants

Open-source DiT video model with synchronized audio, 4K output, and multi-keyframe control

Freemium· Model weights free under OpenRail-M (self-host). Hosted access via LTX Studio (free tier + paid plans), Fal.ai and Replicate pay-per-generation (typically fractions of a cent per second of video). Enterprise licensing available from Lightricks.Text-to-video generationImage-to-video animation
Magic.dev preview image
Magic.dev logo

Magic.dev

Coding · In-house frontier code models (including a long-term-memory 'LTM' model family with reported 100M-token context)

Frontier code models with ultra-long context aimed at automating software engineering

Enterprise· No public pricing. Access is via research partnerships and enterprise engagements; no self-serve tier or public API published as of writing.Whole-repo refactorsLong-horizon feature implementation
QuantProbe preview image
QuantProbe logo

QuantProbe

Evaluation

Physics-based calculator that predicts LLM decode speed, memory fit, and quantization quality on any hardware.

Free· Free and open source; install with `pip install quantprobe` or use the hosted web calculator.GPU and workstation sizing for local LLM inferenceQuantization tier selection (Q4/Q5/Q8) under a latency budget
Tabby preview image
Tabby logo

Tabby

Coding · DeepSeek-Coder, Qwen2.5-Coder, StarCoder, CodeLlama, or any OpenAI-compatible endpoint (Mistral, etc.)

Open-source, self-hosted AI coding assistant

Freemium· Community: $0 user/month · Team: $19 user/month · Enterprise: Contact usSelf-hosted code completionIn-IDE AI chat