Skip to main content
📖 The AI Tool Bible

AI private LLM hosting

Editorial picks for "private llm hosting".

48 tools

All Fine-tuning →
TA

Together AI

Featured
Fine-tuning · Llama / Mistral / Qwen / DeepSeek and others
8.6

Fine-tune & serve open-weight models (Llama, Mistral, DeepSeek).

Paid· Pay-per-token; fine-tuning per-tokenopen modelsfine-tuning
MO

Modal

Fine-tuning · Infrastructure (any model you can host)
8.7

Serverless GPUs and infra for training & serving ML.

Freemium· $30/mo free credits; pay-as-you-go GPU ratesserverless GPUfine-tuning
AB

AWS Bedrock

Agents · Multi-model: Anthropic Claude, Meta Llama, Mistral, Cohere, AI21, Amazon Nova/Titan, DeepSeek, Stability, OpenAI GPT
8.6

Build and scale generative AI applications with foundation models

Paid· Standard: Contact sales · Flex: Contact sales · Priority: Contact sales · Reserved: Contact salesEnterprise RAG chatbot over private documentsMulti-step tool-using agents via AgentCore
GV

Google Vertex AI

Agents · Gemini 2.5 (Pro/Flash/Nano), Imagen, Veo, Chirp, plus Model Garden (Llama, Mistral, Claude via partner)
8.6

Google Cloud's unified platform for building, deploying, and scaling enterprise AI agents and models.

Paid· Image Data - Training (Classification): $3.465 / 1 hour · Image Data - Training (Object Detection): $3.465 / 1 hour · Image Data - Deployment and Online Prediction: $1.375 / 1 hour · Image Data - Batch Prediction: $2.222 / 1 hour · Tabular Data - Training (Classification/Regression): $21.252 / 1 hourEnterprise RAG chatbotMulti-agent customer service
IW

IBM watsonx

Agents · IBM Granite (3.x, Code, Time Series), Meta Llama 3.x, Mistral, plus other curated open models
8.6

Enterprise AI platform for building, deploying, and governing models and agents

Enterprise· watsonx.ai has a free tier on IBM Cloud with limited tokens; paid usage is metered per 1M tokens by model family (Granite, Llama, Mistral, etc.). watsonx.governance and watsonx.data are quoted per environment. Enterprise deals via IBM sales; on-prem/Cloud Pak for Data is separately licensed.Enterprise RAG chatbot over private documentsCustomer service agents with guardrails
HA

H2O.ai

Agents · Multi-model
8.3

Enterprise AI platform combining AutoML, generative AI, and vertical agents for regulated industries.

Enterprise· Contact sales; live demo on requestenterprise-agentsautoml
LL

Llama

Fine-tuning · Llama 4 (Maverick, Scout), Llama 3.3/3.2/3.1
8.3

Meta's open-weight LLM family covering 1B mobile models up to 405B frontier and natively multimodal 10M-context Llama 4 variants.

Freemium· Basic: $15 · Pro: $30 · Enterprise: $100self-hosted-llmfine-tuning
LS

LM Studio

Agents · Multi-model (gpt-oss, Qwen3, Gemma, DeepSeek-R1, Llama, others)
8.3

Desktop app for discovering, downloading, and running open-weight LLMs locally with an OpenAI-compatible server.

Freemium· Free: $0local-llm-inferenceprivate-chat
RU

RunPod

Fine-tuning · Bring-your-own (any open-weight or custom model)
8.3

On-demand GPU cloud and serverless inference platform built specifically for AI workloads.

Paid· Pod: $7.39/hr · Pod: $4.39/hr · Pod: $5.89/hr · Pod: $1.99/hr · Pod: $3.19/hrllm-fine-tuninggpu-rental
VL

vLLM

Fine-tuning · Multi-model (open-weight LLMs: Llama, Qwen, DeepSeek, Mistral, Gemma, Phi, etc.)
8.3

Open-source high-throughput inference engine for serving LLMs with PagedAttention and continuous batching.

Free· Free and open-source (Apache 2.0); self-hosted infrastructure costs applyllm-servingself-hosted-inference
BE

BentoML

Agents · Multi-model
8.2

Open-source framework and managed platform for serving and scaling AI models in production.

Freemium· OSS free (Apache 2.0); managed Bento cloud has free tier + usage-based pricingmodel-servingllm-inference
CO

CoreWeave

Fine-tuning · DeepSeek
8.2

AI-native GPU cloud built for large-scale training, fine-tuning, and inference on NVIDIA hardware.

Enterprise· NVIDIA GB300 NVL72: Contact sales · NVIDIA GB200 NVL72: $42.00 · NVIDIA HGX B300: Contact sales · NVIDIA HGX B200: $68.80 · NVIDIA RTX PRO 6000 Blackwell Server Edition: $20.00model-trainingfine-tuning
DA

Daytona

Agents
8.2

Secure, isolated sandboxes for running AI-generated code with sub-90ms cold starts.

Freemium· Pay-per-second from $0.000014/sec; $200 free creditagent-sandboxescode-interpreters
SG

SGLang

Fine-tuning · Multi-model (DeepSeek, Qwen, Llama, Mistral, GLM, GPT-OSS)
8.2

Open-source high-throughput inference engine for LLMs and multimodal models with OpenAI-compatible serving.

Free· Free, open-source (Apache 2.0); self-hosted infra cost onlyllm-servingmultimodal-inference
LA

Lambda

Fine-tuning · NVIDIA VR200 NVL72, NVIDIA GB300 NVL72, NVIDIA HGX B200, NVIDIA HGX B300, NVIDIA H100
8.1

On-demand NVIDIA GPU cloud built specifically for training, fine-tuning, and serving large AI models.

Paid· Basic: $10 · Pro: $20 · Enterprise: Contact salesllm-trainingfine-tuning
LO

LocalAI

Writing · Multi-model (llama.cpp, diffusers, whisper, etc.)
8.1

Self-hosted OpenAI-compatible API for running LLMs, image, and audio models on your own hardware.

Free· Free and open source (MIT)local-llm-inferenceopenai-api-replacement
TA

Together AI Fine-tuning

Fine-tuning · Multi-model (any Hugging Face open-source model)
8.1

Managed fine-tuning platform for open-source LLMs and vision models with LoRA, full fine-tuning, and RL support.

Paid· Usage-based; cost estimator in-product, no public price listllm-fine-tuningvision-fine-tuning
AN

Anyscale

Fine-tuning · Infrastructure (any model)
7.9

Ray-powered platform for training, serving, and scaling LLMs.

Paid· Enterprise / contact salesdistributed trainingRay
FA

Fireworks AI

Fine-tuning · Multi-model (DeepSeek, Qwen, GLM, Kimi, Gemma, Minimax, others)
7.9

Production inference and fine-tuning platform for open-source LLMs, tuned for speed and enterprise economics.

Freemium· Free signup credits; pay-per-token from ~$0.14/M in; enterprise reserved capacity on requestllm-fine-tuningserverless-inference
LA

Lamini

Fine-tuning · Lamini (built on open base models)
7.7

Memory-tuning platform for grounding LLMs in your facts.

Paid· Enterprise / contact salesenterprise FTfactual recall
MO

Modzy

Agents · Multi-model
7.7

Enterprise ModelOps platform for deploying and running AI/ML models across cloud, on-prem, and edge.

Enterprise· Basic: $10 · Pro: $20 · Enterprise: Contact salesmodel-deploymentedge-ai
OM

oMLX

Coding · Multi-model (Qwen, Llama, Mistral, Gemma, DeepSeek, MiniMax, GLM)
7.5

Native macOS LLM inference server built on MLX, with paged SSD KV caching for Apple Silicon agents.

Free· Free, Apache 2.0 open sourcelocal-llm-inferencecoding-agents
FE

FedML

Fine-tuning · Bring-your-own (PyTorch, Hugging Face)
7.3

Distributed training, fine-tuning, and serving platform with federated learning roots.

Freemium· Open-source library free; managed GPU usage pay-as-you-gofine-tuningdistributed-training
SE

Seldon

Agents · Multi-model (bring your own)
7.3

Kubernetes-native MLOps platform for deploying and orchestrating ML and generative AI models in production.

Freemium· Basic: $10 · Pro: $20 · Enterprise: Contact salesmodel-servinginference-pipelines
VA

Valohai

Agents · Multi-model
7.3

MLOps platform for versioned pipelines, distributed training, and LLM evaluation across any cloud.

Enterprise· Free trial; contact sales for pricingmlopsllm-evaluation
BE

Beam

Coding · openai/gpt-oss-20b
7.2

Serverless GPU infrastructure for AI workloads with sub-second cold starts and bring-your-own-cloud support.

Freemium· $30 free credit refreshed monthly; usage-based beyond thatgpu-inferenceagent-sandboxes
DD

Domino Data Lab

Agents · Multi-model
7.2

Enterprise AI platform for building, deploying, and governing models and agents at scale.

Enterprise· Domino Cloud: Contact sales · Premium: Contact sales · Enterprise: Contact salesenterprise mlopsagentic ai
GE

Geniusrise

Agents · Multi-model
7.2

Open-source framework for building, deploying, and scaling AI microservices across text, vision, and audio.

Free· Free, open source; self-hostedinference-servingfine-tuning
OL

Ollama

Coding · Multi-model (Llama, Qwen, Gemma, DeepSeek, Mistral, Phi, etc.)
7.2

The de facto runtime for running open-weights LLMs locally, now with a paid cloud tier for bigger models.

Freemium· Free local; Pro $20/mo; Max $100/molocal-llmself-hosted-inference
PG

Paperspace Gradient

Fine-tuning · Bring-your-own (PyTorch, TensorFlow, Hugging Face)
7.2

End-to-end MLOps platform with GPU notebooks, training jobs, and model deployment, now folded into DigitalOcean.

Freemium· Free: $0 · Pro: $8 · Growth: $39 · T0: $0 · T1: $12model-trainingfine-tuning
WA

Wallaroo.AI

Agents · Multi-model
7.2

Production AI inference platform for deploying and monitoring models across cloud, on-prem, and edge.

Enterprise· Starter: $500 · Wallaroo Community Edition: Free · Ampere Community Edition: Freemodel-deploymentmlops
IG

Iguazio

Agents · Multi-model
7.1

Enterprise MLOps and GenAI platform for taking models from notebook to production at scale.

Enterprise· Contact sales; free trial availablemlopsllm-fine-tuning
NR

Neu.ro

Agents · Multi-model
7.1

Infrastructure-agnostic MLOps platform for the full ML/DL lifecycle across public, hybrid, and on-prem clouds.

Enterprise· Contact sales; no public pricingmlopsmodel-training
SG

Scale GenAI Platform

Fine-tuning · Multi-model (OpenAI, Google, Meta, Mistral)
7.1

Enterprise agent platform from Scale AI that connects your data, orchestrates multi-agent workflows, and learns from human feedback inside your own VPC.

Enterprise· Contact sales; enterprise contracts onlyenterprise-agentsrag-over-internal-data
AS

Amazon SageMaker

Agents · Multi-model
7.0

AWS's end-to-end platform for building, training, and deploying machine learning models and AI agents at enterprise scale.

Paid· Pay-as-you-go; free tier available for new AWS accountsmodel-trainingmodel-deployment
FO

Forefront

Fine-tuning · Multi-model (Mistral-7B, Mixtral, Phi-2)
7.0

Fine-tune and serve open-source LLMs on your own data without managing GPUs.

Paid· Basic: $20 · Pro: $50 · Enterprise: Contact salesfine-tuningopen-source-llms
JS

Jina Serve

Agents · StableLM
7.0

Open-source Python framework for serving multimodal AI models as scalable gRPC/HTTP microservices.

Freemium· Basic: $10 · Pro: $20 · Enterprise: Contact salesmodel-servingmultimodal-pipelines
OP

OpenSandbox

Agents · Model-agnostic
7.0

Open-source sandbox infrastructure for running AI-generated code, agents, and browsers in isolated Docker or Kubernetes environments.

Free· Open source (Apache 2.0); managed pricing not disclosedcode-executionagent-sandboxing
PR

PrivateGPT

RAG · Multi-model (BYO local LLM)
7.0

Production-ready, air-gapped RAG framework for querying your documents with local LLMs.

Freemium· OSS free; Zylon enterprise contract (contact sales)private-ragchat-with-documents
CH

Chassis

Agents
6.9

Open-source tool that auto-packages ML models into production-ready Docker containers with a prediction API.

Free· Free, open source (Apache-style community project)model-packagingedge-deployment
KA

Katonic AI

Agents · Multi-model (2,600+ via AI Gateway)
6.9

Sovereign enterprise platform for building, deploying, and governing AI agents on your own infrastructure.

Enterprise· Contact sales for quote; no public pricingenterprise-agentson-prem-llm
TR

TrueFoundry

Agents · Multi-model
6.9

Enterprise control plane for deploying, governing, and scaling agentic AI on your own infrastructure.

Enterprise· Developer: $0 · Pro*: $499 · Pro Plus: $2999 · Enterprise: Customagent-deploymentllm-serving
PG

Prediction Guard

Agents · Multi-model
6.6

Self-hosted AI control plane that lets regulated enterprises govern models, agents, and MCP servers behind their firewall.

Enterprise· Contact salesai-governanceself-hosted-llm
AU

AutotuneLLM

Fine-tuning

An open-source optimization layer that sits between your app and Ollama to squeeze more performance out of local LLMs.

Free· Free and open source (MIT licensed).Local LLM inference on Apple SiliconReducing KV cache RAM for Ollama models
LG

LLM Gateway

Agents · Routes to GPT-4o, Claude 3.5 Sonnet, Gemini 1.5, Llama 3.1, Mistral, and 200+ others

One API, 200+ models, transparent pricing, and no vendor lock-in.

Freemium· Free: $0 · Enterprise: CustomMulti-provider LLM routingAutomatic failover between model vendors
LG

LLM GPU Checker (KO)

Evaluation · Catalog covers open models on Hugging Face (Llama, Qwen, Mistral, Gemma, etc.)

Match LLMs to GPUs and plan multi-model AI stacks by VRAM, bandwidth and precision.

Free· Free (open web tool hosted on GitHub Pages).GPU sizing for self-hosted LLMsMulti-GPU RAG stack planning
OW

Open WebUI

Agents · Backend-agnostic: any Ollama, llama.cpp, vLLM, or OpenAI-compatible API (OpenAI GPT, Anthropic Claude, Llama 3.x, Qwen, Mistral, Gemma, etc.)

Self-hosted, extensible AI chat platform that runs on your infrastructure

Freemium· Free (self-hosted, MIT-style community license via pip/Docker) / Enterprise: custom pricing for SSO, RBAC, audit logs, air-gapped deployment, data-residency guaranteesSelf-hosted ChatGPT alternative for a teamPrivate RAG chatbot over internal documents
QU

QuantProbe

Evaluation

Physics-based calculator that predicts LLM decode speed, memory fit, and quantization quality on any hardware.

Free· Free and open source; install with `pip install quantprobe` or use the hosted web calculator.GPU and workstation sizing for local LLM inferenceQuantization tier selection (Q4/Q5/Q8) under a latency budget