AI serverless GPUs
Editorial picks for "serverless gpu".
43 tools

Modal
Serverless GPUs and infra for training & serving ML.

CoreWeave
AI-native GPU cloud built for large-scale training, fine-tuning, and inference on NVIDIA hardware.

Forefront
Fine-tune and serve open-source LLMs on your own data without managing GPUs.

Together AI
FeaturedFine-tune & serve open-weight models (Llama, Mistral, DeepSeek).

Stable Diffusion
Open-source image generation — run anywhere, fine-tune anything.

Llama 3
Meta's open-weights LLM family that put serious frontier-adjacent models in everyone's hands.

LM Studio
Desktop app for discovering, downloading, and running open-weight LLMs locally with an OpenAI-compatible server.

RunPod
On-demand GPU cloud and serverless inference platform built specifically for AI workloads.

vLLM
Open-source high-throughput inference engine for serving LLMs with PagedAttention and continuous batching.

Daytona
Secure, isolated sandboxes for running AI-generated code with sub-90ms cold starts.

Jan
Open-source desktop ChatGPT alternative that runs local LLMs and routes to cloud providers from one app.

SGLang
Open-source high-throughput inference engine for LLMs and multimodal models with OpenAI-compatible serving.

Lambda
On-demand NVIDIA GPU cloud built specifically for training, fine-tuning, and serving large AI models.

Ray Tune
Open-source Python library for distributed hyperparameter tuning at any scale.

Together AI Fine-tuning
Managed fine-tuning platform for open-source LLMs and vision models with LoRA, full fine-tuning, and RL support.

Fal.ai
Serverless GPU inference platform optimized for fast diffusion and generative media APIs.

Genmo
Open-source text-to-video model (Mochi 1) with a hosted playground for turning prompts into short clips.

Fireworks AI
Production inference and fine-tuning platform for open-source LLMs, tuned for speed and enterprise economics.

Langchain-Chatchat
Self-hostable RAG and agent framework that wires LangChain to any local open-source LLM and a knowledge base.

CogVideoX
Open-source text-to-video and image-to-video diffusion transformer from Zhipu AI, runnable on consumer GPUs.

DeepSeek
Chinese AI lab shipping open-weight reasoning models that punch well above their API price.

Dia
Open-weights 1.6B text-to-dialogue model that generates ultra-realistic multi-speaker conversations in one pass.

FedML
Distributed training, fine-tuning, and serving platform with federated learning roots.

Beam
Serverless GPU infrastructure for AI workloads with sub-second cold starts and bring-your-own-cloud support.

LLaMA Factory
Open-source, no-code WebUI for fine-tuning 100+ open LLMs with LoRA, QLoRA, DPO, and PPO.

Paperspace Gradient
End-to-end MLOps platform with GPU notebooks, training jobs, and model deployment, now folded into DigitalOcean.

BGE (BAAI General Embedding)
Open-source embedding and reranker models from BAAI that anchor a huge share of production RAG stacks.

Groq
Custom-silicon LPU inference platform serving open models at GPU-trouncing latency via an OpenAI-compatible API.

Iguazio
Enterprise MLOps and GenAI platform for taking models from notebook to production at scale.

PostgresML
PostgreSQL extension that runs embeddings, vector search, and LLM inference inside your database.

AI Horde (Stable Horde)
Crowdsourced, volunteer-run cluster for free Stable Diffusion image and LLM text generation.

ONNX
Open standard for representing and exchanging machine learning models across frameworks and runtimes.

Apache SINGA
Apache-licensed distributed deep learning library focused on scalable training across GPUs and nodes.

WhisperAPI
Hosted OpenAI Whisper transcription with a pay-as-you-go API and drop-in web dashboard.

Velda
Serverless GPU orchestration that runs AI training and batch jobs without Docker or Kubernetes.
Cloudflare MCP Server
Official suite of remote MCP servers that let Claude, Cursor, and other agents read and control your Cloudflare account.

Colossal-AI
Making large AI models cheaper, faster, and more accessible through distributed training

HiDream
Open-source 17B image model with frontier quality, hosted on Vivago and free to self-host under MIT.

LLM GPU Checker (KO)
Match LLMs to GPUs and plan multi-model AI stacks by VRAM, bandwidth and precision.

LTX Video
Open-source DiT video model with synchronized audio, 4K output, and multi-keyframe control

Magic.dev
Frontier code models with ultra-long context aimed at automating software engineering

QuantProbe
Physics-based calculator that predicts LLM decode speed, memory fit, and quantization quality on any hardware.

Tabby
Open-source, self-hosted AI coding assistant