Skip to main content
📖 The AI Tool Bible

AI tools tagged On Premise

48 tools matching this tag.tech

All tags →
SD

Stable Diffusion

Image Generation · SD 3.5 / SDXL
8.8

Open-source image generation — run anywhere, fine-tune anything.

Free· Free open weights; optional Stability APIlocalfine-tuning
EV

Elasticsearch Vector Search

RAG · BYO embeddings (OpenAI, Cohere, Hugging Face, Mistral, Bedrock, Vertex, Azure) plus Elastic's built-in ELSER sparse model and E5 dense model
8.7

Hybrid vector + keyword search in the enterprise-grade Elasticsearch engine

Freemium· Resource based pricing: Pay as you go (monthly) or prepaid · Usage based pricing: Pay as you go (monthly) or prepaid · License based pricing: ?RAG chatbot over enterprise docsHybrid semantic + keyword product search
IW

IBM watsonx

Agents · IBM Granite (3.x, Code, Time Series), Meta Llama 3.x, Mistral, plus other curated open models
8.6

Enterprise AI platform for building, deploying, and governing models and agents

Enterprise· watsonx.ai has a free tier on IBM Cloud with limited tokens; paid usage is metered per 1M tokens by model family (Granite, Llama, Mistral, etc.). watsonx.governance and watsonx.data are quoted per environment. Enterprise deals via IBM sales; on-prem/Cloud Pak for Data is separately licensed.Enterprise RAG chatbot over private documentsCustomer service agents with guardrails
PA

Palantir AIP

Agents · Multi-model (GPT, Claude, Llama, customer-hosted)
8.4

Enterprise AI platform that grounds LLMs in your operational data and runs agents against real business systems.

Enterprise· Contact sales; typically bundled with Foundryenterprise-agentsoperational-ai
LL

Llama

Fine-tuning · Llama 4 (Maverick, Scout), Llama 3.3/3.2/3.1
8.3

Meta's open-weight LLM family covering 1B mobile models up to 405B frontier and natively multimodal 10M-context Llama 4 variants.

Freemium· Basic: $15 · Pro: $30 · Enterprise: $100self-hosted-llmfine-tuning
VL

vLLM

Fine-tuning · Multi-model (open-weight LLMs: Llama, Qwen, DeepSeek, Mistral, Gemma, Phi, etc.)
8.3

Open-source high-throughput inference engine for serving LLMs with PagedAttention and continuous batching.

Free· Free and open-source (Apache 2.0); self-hosted infrastructure costs applyllm-servingself-hosted-inference
VE

Vespa

RAG · Hosted search engine (not an LLM)
8.2

Yahoo's open-source search engine with vector + sparse retrieval.

Freemium· Free open-source; Vespa Cloud paidlarge-scale searchranking
LA

Lambda

Fine-tuning · NVIDIA VR200 NVL72, NVIDIA GB300 NVL72, NVIDIA HGX B200, NVIDIA HGX B300, NVIDIA H100
8.1

On-demand NVIDIA GPU cloud built specifically for training, fine-tuning, and serving large AI models.

Paid· Basic: $10 · Pro: $20 · Enterprise: Contact salesllm-trainingfine-tuning
LO

LocalAI

Writing · Multi-model (llama.cpp, diffusers, whisper, etc.)
8.1

Self-hosted OpenAI-compatible API for running LLMs, image, and audio models on your own hardware.

Free· Free and open source (MIT)local-llm-inferenceopenai-api-replacement
CO

Continue

Coding · BYO (any OpenAI-compatible API + Ollama for local)
7.9

Open-source, self-hostable VS Code/JetBrains AI assistant.

Free· Free / open-source; you pay model costsself-hostedopen source
DA

DataRobot

Agents · Multi-model
7.8

Enterprise platform for building, deploying, and governing AI agents alongside classic predictive ML.

Enterprise· Contact sales; free trial availableenterprise-agentspredictive-ml
LA

Lamini

Fine-tuning · Lamini (built on open base models)
7.7

Memory-tuning platform for grounding LLMs in your facts.

Paid· Enterprise / contact salesenterprise FTfactual recall
MO

Modzy

Agents · Multi-model
7.7

Enterprise ModelOps platform for deploying and running AI/ML models across cloud, on-prem, and edge.

Enterprise· Basic: $10 · Pro: $20 · Enterprise: Contact salesmodel-deploymentedge-ai
TA

Tabnine

Coding · Proprietary (Tabnine-trained, on-prem capable)
7.6

Privacy-focused AI autocomplete with on-prem options.

Paid· Tabnine Code Assistant: $39 · Tabnine Agentic Platform: $59enterpriseprivacy
KA

Kong AI Gateway

Agents · Multi-model
7.3

Enterprise API gateway extended to route, govern, and observe LLM and agent traffic across providers.

Freemium· Free trial: $0 · Plus: Charged per Gateway per month · Enterprise: Custom pricing · Essentials: $0 · Pro: $12llm-gatewaymulti-llm-routing
KU

Kubeflow

Agents · Multi-framework (PyTorch, JAX, XGBoost, TensorFlow)
7.3

Open-source toolkit for running the full ML lifecycle on Kubernetes.

Free· Free and open source; commercial distributions and managed offerings priced separately by vendorsml-pipelinesdistributed-training
PA

Pachyderm

Fine-tuning
7.3

Kubernetes-native data versioning and pipeline engine for reproducible ML at petabyte scale.

Freemium· Basic: $10 · Pro: $30 · Enterprise: Contact salesdata-versioningml-pipelines
SE

Seldon

Agents · Multi-model (bring your own)
7.3

Kubernetes-native MLOps platform for deploying and orchestrating ML and generative AI models in production.

Freemium· Basic: $10 · Pro: $20 · Enterprise: Contact salesmodel-servinginference-pipelines
DD

Domino Data Lab

Agents · Multi-model
7.2

Enterprise AI platform for building, deploying, and governing models and agents at scale.

Enterprise· Domino Cloud: Contact sales · Premium: Contact sales · Enterprise: Contact salesenterprise mlopsagentic ai
SV

SAS Viya

Agents · Multi-model
7.2

Enterprise-grade data and AI analytics platform with built-in governance, MCP server, and a Copilot for regulated industries.

Enterprise· Contact sales; 14-day free trialenterprise-analyticsai-governance
WA

Wallaroo.AI

Agents · Multi-model
7.2

Production AI inference platform for deploying and monitoring models across cloud, on-prem, and edge.

Enterprise· Starter: $500 · Wallaroo Community Edition: Free · Ampere Community Edition: Freemodel-deploymentmlops
FA

Fiddler AI

Evaluation · Fiddler Centor (proprietary evaluators)
7.1

Enterprise AI observability and guardrails platform for monitoring agents, LLMs, and ML models in production.

Enterprise· Free: Free · Developer: $0.002 per trace · Enterprise: Contact salesllm-observabilityagent-monitoring
IG

Iguazio

Agents · Multi-model
7.1

Enterprise MLOps and GenAI platform for taking models from notebook to production at scale.

Enterprise· Contact sales; free trial availablemlopsllm-fine-tuning
II

iii

Agents · Bring-your-own
7.1

Self-hosted distributed runtime for multi-language workers and AI agents under one protocol.

Free· Open-source; commercial pricing not publishedagent-orchestrationdistributed-workers
NR

Neu.ro

Agents · Multi-model
7.1

Infrastructure-agnostic MLOps platform for the full ML/DL lifecycle across public, hybrid, and on-prem clouds.

Enterprise· Contact sales; no public pricingmlopsmodel-training
PO

PostgresML

RAG · Multi-model (Llama, Mistral, open-source embeddings)
7.1

PostgreSQL extension that runs embeddings, vector search, and LLM inference inside your database.

Freemium· Serverless: From $7.50 per query hour · Dedicated: From $0.60 per instance hour · Enterprise: Custom pricingvector-searchrag
SG

Scale GenAI Platform

Fine-tuning · Multi-model (OpenAI, Google, Meta, Mistral)
7.1

Enterprise agent platform from Scale AI that connects your data, orchestrates multi-agent workflows, and learns from human feedback inside your own VPC.

Enterprise· Contact sales; enterprise contracts onlyenterprise-agentsrag-over-internal-data
TI

Tiledesk

Agents · Multi-model
7.1

Open-source conversational AI platform for building multi-channel support agents and automation workflows.

Freemium· 14-day free trial; freemium, paid, and enterprise tierscustomer-supportrag-knowledge-base
BI

Bisheng

Agents · Multi-model
7.0

Open-source enterprise AgentOps platform for building, orchestrating, and governing LLM agents at scale.

Freemium· Open-source community edition free; commercial enterprise edition via DataElement salesagent-orchestrationrag
KO

Kotaemon

RAG · Multi-model (OpenAI, LlamaCPP, any OpenAI-compatible endpoint)
7.0

Open-source RAG UI for chatting with your own documents, locally or self-hosted.

Free· Free, open-source (MIT-style); self-hosted infrastructure costs onlydocument-qaprivate-rag
MA

MaxKB

RAG · Multi-model
7.0

Open-source enterprise RAG and agent platform with built-in workflow engine and multi-LLM support.

Freemium· Community edition free (GPLv3); paid enterprise editionenterprise-knowledge-basecustomer-support-bots
OP

OpenSandbox

Agents · Model-agnostic
7.0

Open-source sandbox infrastructure for running AI-generated code, agents, and browsers in isolated Docker or Kubernetes environments.

Free· Open source (Apache 2.0); managed pricing not disclosedcode-executionagent-sandboxing
PR

PrivateGPT

RAG · Multi-model (BYO local LLM)
7.0

Production-ready, air-gapped RAG framework for querying your documents with local LLMs.

Freemium· OSS free; Zylon enterprise contract (contact sales)private-ragchat-with-documents
SY

SystemPrompt

Agents · Multi-model
7.0

Self-hosted AI governance gateway that audits, gates, and logs every LLM call before it leaves your network.

Freemium· Free self-hosted tier; commercial licensing on requestai-governancellm-gateway
CO

Cohere

RAG · Command, Embed, Rerank, Transcribe (proprietary)
6.9

Enterprise-grade LLM platform built for private, secure, and customizable deployment.

Enterprise· Embed 4 Small: $2,500 · Embed 4 Medium: $3,250 · Rerank 3.5 Medium: $3,250 · Rerank 4 Fast Medium: $3,250 · Rerank 4 Pro Medium: $3,250enterprise-ragsemantic-search
DE

DeepSearcher

RAG · Multi-model (DeepSeek, OpenAI o1/o3-mini, Claude, Llama, others)
6.9

Open-source agentic RAG framework for private enterprise data, built by the Zilliz/Milvus team.

Free· Free, Apache 2.0; bring your own LLM and vector DB costsenterprise-ragagentic-search
FR

Frigate

Video · Custom object detection (Coral/YOLO/Frigate+ models)
6.9

Open-source NVR that runs AI object detection locally on your security camera feeds.

Freemium· Free and open-source; Frigate+ custom-model subscription is paidsecurity-camera-nvrobject-detection
KA

Katonic AI

Agents · Multi-model (2,600+ via AI Gateway)
6.9

Sovereign enterprise platform for building, deploying, and governing AI agents on your own infrastructure.

Enterprise· Contact sales for quote; no public pricingenterprise-agentson-prem-llm
TR

TrueFoundry

Agents · Multi-model
6.9

Enterprise control plane for deploying, governing, and scaling agentic AI on your own infrastructure.

Enterprise· Developer: $0 · Pro*: $499 · Pro Plus: $2999 · Enterprise: Customagent-deploymentllm-serving
EA

EKHOS AI

Audio · Proprietary local models
6.8

Offline Windows transcription app with speaker diarization, GPU acceleration, and 98-language support.

Freemium· Premium: $9transcriptionspeaker-diarization
GP

GPTLocalhost

Writing · Bring-your-own (Ollama, LM Studio, llama.cpp, Foundry Local, etc.)
6.8

Run local LLMs directly inside Microsoft Word without sending text to the cloud.

Freemium· Free: Free · Monthly: ? · Lifetime: ?private draftingoffline writing assistant
PG

Prediction Guard

Agents · Multi-model
6.6

Self-hosted AI control plane that lets regulated enterprises govern models, agents, and MCP servers behind their firewall.

Enterprise· Contact salesai-governanceself-hosted-llm
CL

ClickHouse

RAG

The open-source columnar database powering real-time analytics — and, increasingly, LLM observability and RAG backends.

Freemium· Open-source self-managed: free. ClickHouse Cloud: from $50/month (usage-based on compute + storage, AWS/GCP/Azure). Enterprise tier available with dedicated support and BYOC options.LLM trace and cost analyticsRAG retrieval with hybrid vector + metadata filters
CO

CowAgent

Agents · Model-agnostic: Claude, GPT, Gemini, DeepSeek, Qwen, GLM, Kimi, MiniMax, Doubao

Open-source, self-hosted AI agent framework that plans, uses tools, and executes multi-step tasks.

Free· Free and open source under MIT License; users bring their own model API keys (e.g. OpenAI, Anthropic, DeepSeek) and pay those providers directly.Autonomous research assistantSelf-hosted coding agent
GO

Goose

Agents · Provider-agnostic — Claude 3.5/4, GPT-4o/4.1, Gemini 1.5/2, Llama 3.x via Ollama, plus Azure OpenAI, Bedrock, OpenRouter and others

Open-source, local-first AI agent for code, workflows, and everything in between.

Free· Free and open source under Apache 2.0. You bring your own model API keys (or run local models via Ollama); provider costs are pass-through.Multi-file code refactors and PR generationLocal repo Q&A and code review
JE

JeecgBoot

Coding · Model-agnostic: ChatGPT, DeepSeek (default), Qwen, Ollama-hosted local models

Open-source enterprise low-code platform with a pluggable AI code generator, RAG knowledge base and MCP agent tooling.

Freemium· Community Edition: free under Apache 2.0 license. Commercial/Enterprise editions with additional modules, priority support, and license flexibility are quoted on request via jeecg.com sales (QQ / hotline).AI-assisted CRUD app generationInternal ERP and CRM builds
ME

Meilisearch

RAG · Retrieval engine (Rust); pluggable embedders including OpenAI, Cohere, Hugging Face, Ollama and custom REST models

Open-source, lightning-fast search engine with built-in hybrid and vector search for RAG

Freemium· Cloud: $20/month · Usage-based: $30/month · Resource-based: $23/month · Instance: $18/monthE-commerce product searchDocumentation and knowledge base search
ME

MemPalace

Agents · Embedding backends: embedding-gemma-300m (multilingual) or all-MiniLM-L6-v2 (English). No LLM required for retrieval.

Open-source, local-first AI memory system for LLM agents

Free· Free and open-source under the MIT license. No paid tiers; self-hosted with zero API costs once installed.Persistent memory for Claude Code and Cursor sessionsPer-agent memory isolation in multi-agent systems