Skip to main content
📖 The AI Tool Bible

AI LLM deployment

Editorial picks for "deploy llm".

48 tools

All Fine-tuning →
TA

Together AI

Featured
Fine-tuning · Llama / Mistral / Qwen / DeepSeek and others
8.6

Fine-tune & serve open-weight models (Llama, Mistral, DeepSeek).

Paid· Pay-per-token; fine-tuning per-tokenopen modelsfine-tuning
MO

Modal

Fine-tuning · Infrastructure (any model you can host)
8.7

Serverless GPUs and infra for training & serving ML.

Freemium· $30/mo free credits; pay-as-you-go GPU ratesserverless GPUfine-tuning
AB

AWS Bedrock

Agents · Multi-model: Anthropic Claude, Meta Llama, Mistral, Cohere, AI21, Amazon Nova/Titan, DeepSeek, Stability, OpenAI GPT
8.6

Build and scale generative AI applications with foundation models

Paid· Standard: Contact sales · Flex: Contact sales · Priority: Contact sales · Reserved: Contact salesEnterprise RAG chatbot over private documentsMulti-step tool-using agents via AgentCore
GV

Google Vertex AI

Agents · Gemini 2.5 (Pro/Flash/Nano), Imagen, Veo, Chirp, plus Model Garden (Llama, Mistral, Claude via partner)
8.6

Google Cloud's unified platform for building, deploying, and scaling enterprise AI agents and models.

Paid· Image Data - Training (Classification): $3.465 / 1 hour · Image Data - Training (Object Detection): $3.465 / 1 hour · Image Data - Deployment and Online Prediction: $1.375 / 1 hour · Image Data - Batch Prediction: $2.222 / 1 hour · Tabular Data - Training (Classification/Regression): $21.252 / 1 hourEnterprise RAG chatbotMulti-agent customer service
IW

IBM watsonx

Agents · IBM Granite (3.x, Code, Time Series), Meta Llama 3.x, Mistral, plus other curated open models
8.6

Enterprise AI platform for building, deploying, and governing models and agents

Enterprise· watsonx.ai has a free tier on IBM Cloud with limited tokens; paid usage is metered per 1M tokens by model family (Granite, Llama, Mistral, etc.). watsonx.governance and watsonx.data are quoted per environment. Enterprise deals via IBM sales; on-prem/Cloud Pak for Data is separately licensed.Enterprise RAG chatbot over private documentsCustomer service agents with guardrails
RE

Replicate

Fine-tuning · Thousands of community + first-party models
8.5

One-API platform for running and fine-tuning open-source models.

Paid· Pay-per-second of GPU timemodel hostingfine-tuning
CL

ClearML

Agents · Model-agnostic
8.3

End-to-end MLOps and GenAI platform with open-source experiment tracking and enterprise GPU orchestration.

Freemium· Community: $0 · Pro: $15 Per User/Month + Usage · Scale: Custom Quote · Enterprise: Request a Quoteexperiment-trackinggpu-orchestration
DA

Dataiku

Agents · Multi-model (LLM Mesh: OpenAI, Anthropic, Bedrock, Vertex, OSS)
8.3

Enterprise AI platform unifying data, ML, LLMs, and agents under one governed workflow.

Enterprise· Basic: $10 · Pro: $20 · Enterprise: Contact salesenterprise-aiagent-orchestration
HA

H2O.ai

Agents · Multi-model
8.3

Enterprise AI platform combining AutoML, generative AI, and vertical agents for regulated industries.

Enterprise· Contact sales; live demo on requestenterprise-agentsautoml
LL

Llama

Fine-tuning · Llama 4 (Maverick, Scout), Llama 3.3/3.2/3.1
8.3

Meta's open-weight LLM family covering 1B mobile models up to 405B frontier and natively multimodal 10M-context Llama 4 variants.

Freemium· Basic: $15 · Pro: $30 · Enterprise: $100self-hosted-llmfine-tuning
LS

LM Studio

Agents · Multi-model (gpt-oss, Qwen3, Gemma, DeepSeek-R1, Llama, others)
8.3

Desktop app for discovering, downloading, and running open-weight LLMs locally with an OpenAI-compatible server.

Freemium· Free: $0local-llm-inferenceprivate-chat
RU

RunPod

Fine-tuning · Bring-your-own (any open-weight or custom model)
8.3

On-demand GPU cloud and serverless inference platform built specifically for AI workloads.

Paid· Pod: $7.39/hr · Pod: $4.39/hr · Pod: $5.89/hr · Pod: $1.99/hr · Pod: $3.19/hrllm-fine-tuninggpu-rental
VL

vLLM

Fine-tuning · Multi-model (open-weight LLMs: Llama, Qwen, DeepSeek, Mistral, Gemma, Phi, etc.)
8.3

Open-source high-throughput inference engine for serving LLMs with PagedAttention and continuous batching.

Free· Free and open-source (Apache 2.0); self-hosted infrastructure costs applyllm-servingself-hosted-inference
BE

BentoML

Agents · Multi-model
8.2

Open-source framework and managed platform for serving and scaling AI models in production.

Freemium· OSS free (Apache 2.0); managed Bento cloud has free tier + usage-based pricingmodel-servingllm-inference
CO

CoreWeave

Fine-tuning · DeepSeek
8.2

AI-native GPU cloud built for large-scale training, fine-tuning, and inference on NVIDIA hardware.

Enterprise· NVIDIA GB300 NVL72: Contact sales · NVIDIA GB200 NVL72: $42.00 · NVIDIA HGX B300: Contact sales · NVIDIA HGX B200: $68.80 · NVIDIA RTX PRO 6000 Blackwell Server Edition: $20.00model-trainingfine-tuning
DA

Daytona

Agents
8.2

Secure, isolated sandboxes for running AI-generated code with sub-90ms cold starts.

Freemium· Pay-per-second from $0.000014/sec; $200 free creditagent-sandboxescode-interpreters
OP

OpenPipe

Fine-tuning · Llama, Mistral, Qwen and other open-weight base models
8.2

Fine-tuning and reinforcement learning platform for turning expensive prompts into cheap, fast, task-specific models.

Freemium· Free tier available; usage-based pricing for training and hosted inference; enterprise plans on requestllm-cost-reductionfine-tuning
SG

SGLang

Fine-tuning · Multi-model (DeepSeek, Qwen, Llama, Mistral, GLM, GPT-OSS)
8.2

Open-source high-throughput inference engine for LLMs and multimodal models with OpenAI-compatible serving.

Free· Free, open-source (Apache 2.0); self-hosted infra cost onlyllm-servingmultimodal-inference
UN

Unsloth

Fine-tuning · Llama, Mistral, Gemma, Qwen, GLM (multi-model)
8.2

Open-source LLM fine-tuning toolkit with custom kernels that train 2-30x faster and use up to 90% less VRAM.

Freemium· Free open-source; Pro and Enterprise contact saleslora-finetuningqlora
LA

Lambda

Fine-tuning · NVIDIA VR200 NVL72, NVIDIA GB300 NVL72, NVIDIA HGX B200, NVIDIA HGX B300, NVIDIA H100
8.1

On-demand NVIDIA GPU cloud built specifically for training, fine-tuning, and serving large AI models.

Paid· Basic: $10 · Pro: $20 · Enterprise: Contact salesllm-trainingfine-tuning
LO

LocalAI

Writing · Multi-model (llama.cpp, diffusers, whisper, etc.)
8.1

Self-hosted OpenAI-compatible API for running LLMs, image, and audio models on your own hardware.

Free· Free and open source (MIT)local-llm-inferenceopenai-api-replacement
TA

Together AI Fine-tuning

Fine-tuning · Multi-model (any Hugging Face open-source model)
8.1

Managed fine-tuning platform for open-source LLMs and vision models with LoRA, full fine-tuning, and RL support.

Paid· Usage-based; cost estimator in-product, no public price listllm-fine-tuningvision-fine-tuning
MI

MiniMax

Agents · MiniMax M3, Hailuo 2.3, Speech 2.8, Music 2.6
8.0

Chinese frontier-model lab shipping multimodal foundation models with a 1M-context coding/agent stack.

Freemium· Free tier; Token plan from ~$20/mo (~12.5B tokens); enterprise pricing on requestcoding-agentlong-context
AN

Anyscale

Fine-tuning · Infrastructure (any model)
7.9

Ray-powered platform for training, serving, and scaling LLMs.

Paid· Enterprise / contact salesdistributed trainingRay
FA

Fireworks AI

Fine-tuning · Multi-model (DeepSeek, Qwen, GLM, Kimi, Gemma, Minimax, others)
7.9

Production inference and fine-tuning platform for open-source LLMs, tuned for speed and enterprise economics.

Freemium· Free signup credits; pay-per-token from ~$0.14/M in; enterprise reserved capacity on requestllm-fine-tuningserverless-inference
DA

DataRobot

Agents · Multi-model
7.8

Enterprise platform for building, deploying, and governing AI agents alongside classic predictive ML.

Enterprise· Contact sales; free trial availableenterprise-agentspredictive-ml
MO

Modzy

Agents · Multi-model
7.7

Enterprise ModelOps platform for deploying and running AI/ML models across cloud, on-prem, and edge.

Enterprise· Basic: $10 · Pro: $20 · Enterprise: Contact salesmodel-deploymentedge-ai
OM

oMLX

Coding · Multi-model (Qwen, Llama, Mistral, Gemma, DeepSeek, MiniMax, GLM)
7.5

Native macOS LLM inference server built on MLX, with paged SSD KV caching for Apple Silicon agents.

Free· Free, Apache 2.0 open sourcelocal-llm-inferencecoding-agents
BI

Bifrost

Agents · Multi-model (OpenAI, Anthropic, Bedrock, Vertex, 1000+ via providers)
7.3

Open-source AI gateway that unifies 1000+ models behind one OpenAI-compatible endpoint with failover, budgets, and MCP routing.

Freemium· Free Forever: Free · Enterprise: Contact salesllm-gatewaymulti-provider-routing
FE

FedML

Fine-tuning · Bring-your-own (PyTorch, Hugging Face)
7.3

Distributed training, fine-tuning, and serving platform with federated learning roots.

Freemium· Open-source library free; managed GPU usage pay-as-you-gofine-tuningdistributed-training
KA

Kong AI Gateway

Agents · Multi-model
7.3

Enterprise API gateway extended to route, govern, and observe LLM and agent traffic across providers.

Freemium· Free trial: $0 · Plus: Charged per Gateway per month · Enterprise: Custom pricing · Essentials: $0 · Pro: $12llm-gatewaymulti-llm-routing
KU

Kubeflow

Agents · Multi-framework (PyTorch, JAX, XGBoost, TensorFlow)
7.3

Open-source toolkit for running the full ML lifecycle on Kubernetes.

Free· Free and open source; commercial distributions and managed offerings priced separately by vendorsml-pipelinesdistributed-training
PA

Pachyderm

Fine-tuning
7.3

Kubernetes-native data versioning and pipeline engine for reproducible ML at petabyte scale.

Freemium· Basic: $10 · Pro: $30 · Enterprise: Contact salesdata-versioningml-pipelines
SE

Seldon

Agents · Multi-model (bring your own)
7.3

Kubernetes-native MLOps platform for deploying and orchestrating ML and generative AI models in production.

Freemium· Basic: $10 · Pro: $20 · Enterprise: Contact salesmodel-servinginference-pipelines
VA

Valohai

Agents · Multi-model
7.3

MLOps platform for versioned pipelines, distributed training, and LLM evaluation across any cloud.

Enterprise· Free trial; contact sales for pricingmlopsllm-evaluation
BE

Beam

Coding · openai/gpt-oss-20b
7.2

Serverless GPU infrastructure for AI workloads with sub-second cold starts and bring-your-own-cloud support.

Freemium· $30 free credit refreshed monthly; usage-based beyond thatgpu-inferenceagent-sandboxes
DD

Domino Data Lab

Agents · Multi-model
7.2

Enterprise AI platform for building, deploying, and governing models and agents at scale.

Enterprise· Domino Cloud: Contact sales · Premium: Contact sales · Enterprise: Contact salesenterprise mlopsagentic ai
GE

Geniusrise

Agents · Multi-model
7.2

Open-source framework for building, deploying, and scaling AI microservices across text, vision, and audio.

Free· Free, open source; self-hostedinference-servingfine-tuning
OL

Ollama

Coding · Multi-model (Llama, Qwen, Gemma, DeepSeek, Mistral, Phi, etc.)
7.2

The de facto runtime for running open-weights LLMs locally, now with a paid cloud tier for bigger models.

Freemium· Free local; Pro $20/mo; Max $100/molocal-llmself-hosted-inference
PG

Paperspace Gradient

Fine-tuning · Bring-your-own (PyTorch, TensorFlow, Hugging Face)
7.2

End-to-end MLOps platform with GPU notebooks, training jobs, and model deployment, now folded into DigitalOcean.

Freemium· Free: $0 · Pro: $8 · Growth: $39 · T0: $0 · T1: $12model-trainingfine-tuning
WA

Wallaroo.AI

Agents · Multi-model
7.2

Production AI inference platform for deploying and monitoring models across cloud, on-prem, and edge.

Enterprise· Starter: $500 · Wallaroo Community Edition: Free · Ampere Community Edition: Freemodel-deploymentmlops
GR

Groq

Coding · Multi-model (Llama, Mixtral, Gemma, Qwen, Whisper)
7.1

Custom-silicon LPU inference platform serving open models at GPU-trouncing latency via an OpenAI-compatible API.

Freemium· Free API key with rate limits; per-token paid tiers; enterprise contractslow-latency inferencevoice agents
IG

Iguazio

Agents · Multi-model
7.1

Enterprise MLOps and GenAI platform for taking models from notebook to production at scale.

Enterprise· Contact sales; free trial availablemlopsllm-fine-tuning
NR

Neu.ro

Agents · Multi-model
7.1

Infrastructure-agnostic MLOps platform for the full ML/DL lifecycle across public, hybrid, and on-prem clouds.

Enterprise· Contact sales; no public pricingmlopsmodel-training
SG

Scale GenAI Platform

Fine-tuning · Multi-model (OpenAI, Google, Meta, Mistral)
7.1

Enterprise agent platform from Scale AI that connects your data, orchestrates multi-agent workflows, and learns from human feedback inside your own VPC.

Enterprise· Contact sales; enterprise contracts onlyenterprise-agentsrag-over-internal-data
AS

Amazon SageMaker

Agents · Multi-model
7.0

AWS's end-to-end platform for building, training, and deploying machine learning models and AI agents at enterprise scale.

Paid· Pay-as-you-go; free tier available for new AWS accountsmodel-trainingmodel-deployment
FO

Forefront

Fine-tuning · Multi-model (Mistral-7B, Mixtral, Phi-2)
7.0

Fine-tune and serve open-source LLMs on your own data without managing GPUs.

Paid· Basic: $20 · Pro: $50 · Enterprise: Contact salesfine-tuningopen-source-llms
JS

Jina Serve

Agents · StableLM
7.0

Open-source Python framework for serving multimodal AI models as scalable gRPC/HTTP microservices.

Freemium· Basic: $10 · Pro: $20 · Enterprise: Contact salesmodel-servingmultimodal-pipelines