AI private LLM hosting
Editorial picks for "private llm hosting".
48 tools
Together AI
FeaturedFine-tune & serve open-weight models (Llama, Mistral, DeepSeek).
Modal
Serverless GPUs and infra for training & serving ML.
AWS Bedrock
Build and scale generative AI applications with foundation models
Google Vertex AI
Google Cloud's unified platform for building, deploying, and scaling enterprise AI agents and models.
IBM watsonx
Enterprise AI platform for building, deploying, and governing models and agents
H2O.ai
Enterprise AI platform combining AutoML, generative AI, and vertical agents for regulated industries.
Llama
Meta's open-weight LLM family covering 1B mobile models up to 405B frontier and natively multimodal 10M-context Llama 4 variants.
LM Studio
Desktop app for discovering, downloading, and running open-weight LLMs locally with an OpenAI-compatible server.
RunPod
On-demand GPU cloud and serverless inference platform built specifically for AI workloads.
vLLM
Open-source high-throughput inference engine for serving LLMs with PagedAttention and continuous batching.
BentoML
Open-source framework and managed platform for serving and scaling AI models in production.
CoreWeave
AI-native GPU cloud built for large-scale training, fine-tuning, and inference on NVIDIA hardware.
Daytona
Secure, isolated sandboxes for running AI-generated code with sub-90ms cold starts.
SGLang
Open-source high-throughput inference engine for LLMs and multimodal models with OpenAI-compatible serving.
Lambda
On-demand NVIDIA GPU cloud built specifically for training, fine-tuning, and serving large AI models.
LocalAI
Self-hosted OpenAI-compatible API for running LLMs, image, and audio models on your own hardware.
Together AI Fine-tuning
Managed fine-tuning platform for open-source LLMs and vision models with LoRA, full fine-tuning, and RL support.
Anyscale
Ray-powered platform for training, serving, and scaling LLMs.
Fireworks AI
Production inference and fine-tuning platform for open-source LLMs, tuned for speed and enterprise economics.
Lamini
Memory-tuning platform for grounding LLMs in your facts.
Modzy
Enterprise ModelOps platform for deploying and running AI/ML models across cloud, on-prem, and edge.
oMLX
Native macOS LLM inference server built on MLX, with paged SSD KV caching for Apple Silicon agents.
FedML
Distributed training, fine-tuning, and serving platform with federated learning roots.
Seldon
Kubernetes-native MLOps platform for deploying and orchestrating ML and generative AI models in production.
Valohai
MLOps platform for versioned pipelines, distributed training, and LLM evaluation across any cloud.
Beam
Serverless GPU infrastructure for AI workloads with sub-second cold starts and bring-your-own-cloud support.
Domino Data Lab
Enterprise AI platform for building, deploying, and governing models and agents at scale.
Geniusrise
Open-source framework for building, deploying, and scaling AI microservices across text, vision, and audio.
Ollama
The de facto runtime for running open-weights LLMs locally, now with a paid cloud tier for bigger models.
Paperspace Gradient
End-to-end MLOps platform with GPU notebooks, training jobs, and model deployment, now folded into DigitalOcean.
Wallaroo.AI
Production AI inference platform for deploying and monitoring models across cloud, on-prem, and edge.
Iguazio
Enterprise MLOps and GenAI platform for taking models from notebook to production at scale.
Neu.ro
Infrastructure-agnostic MLOps platform for the full ML/DL lifecycle across public, hybrid, and on-prem clouds.
Scale GenAI Platform
Enterprise agent platform from Scale AI that connects your data, orchestrates multi-agent workflows, and learns from human feedback inside your own VPC.
Amazon SageMaker
AWS's end-to-end platform for building, training, and deploying machine learning models and AI agents at enterprise scale.
Forefront
Fine-tune and serve open-source LLMs on your own data without managing GPUs.
Jina Serve
Open-source Python framework for serving multimodal AI models as scalable gRPC/HTTP microservices.
OpenSandbox
Open-source sandbox infrastructure for running AI-generated code, agents, and browsers in isolated Docker or Kubernetes environments.
PrivateGPT
Production-ready, air-gapped RAG framework for querying your documents with local LLMs.
Chassis
Open-source tool that auto-packages ML models into production-ready Docker containers with a prediction API.
Katonic AI
Sovereign enterprise platform for building, deploying, and governing AI agents on your own infrastructure.
TrueFoundry
Enterprise control plane for deploying, governing, and scaling agentic AI on your own infrastructure.
Prediction Guard
Self-hosted AI control plane that lets regulated enterprises govern models, agents, and MCP servers behind their firewall.
AutotuneLLM
An open-source optimization layer that sits between your app and Ollama to squeeze more performance out of local LLMs.
LLM Gateway
One API, 200+ models, transparent pricing, and no vendor lock-in.
LLM GPU Checker (KO)
Match LLMs to GPUs and plan multi-model AI stacks by VRAM, bandwidth and precision.
Open WebUI
Self-hosted, extensible AI chat platform that runs on your infrastructure
QuantProbe
Physics-based calculator that predicts LLM decode speed, memory fit, and quantization quality on any hardware.