AI LLM deployment
Editorial picks for "deploy llm".
48 tools

Elasticsearch Vector Search
Hybrid vector + keyword search in the enterprise-grade Elasticsearch engine

Replicate
One-API platform for running and fine-tuning open-source models.

ChatGLM (Zhipu Qingyan)
Zhipu AI's bilingual ChatGLM assistant with agents, code, image and video generation

Palantir AIP
Enterprise AI platform that grounds LLMs in your operational data and runs agents against real business systems.

PyCaret
Low-code Python AutoML library that wraps scikit-learn, XGBoost, LightGBM and friends behind a few-line API.

Yi (01.AI)
Foundation models from 01.AI — open-weight Yi family plus frontier Yi-Lightning and Yi-Large

ClearML
End-to-end MLOps and GenAI platform with open-source experiment tracking and enterprise GPU orchestration.

Llama
Meta's open-weight LLM family covering 1B mobile models up to 405B frontier and natively multimodal 10M-context Llama 4 variants.

vLLM
Open-source high-throughput inference engine for serving LLMs with PagedAttention and continuous batching.

Arize AI
Enterprise observability and evaluation platform for LLM agents and generative AI applications.

Giskard
Continuous AI red teaming platform that stress-tests LLM agents for vulnerabilities before they hit production.

LibreChat
Open-source, self-hostable ChatGPT-style frontend that brings every major LLM provider under one roof.

Writer
Enterprise generative AI platform built around in-house Palmyra LLMs for regulated, brand-consistent content.

HoneyHive
OpenTelemetry-native observability and evaluation platform for LLM agents in production.

Lambda
On-demand NVIDIA GPU cloud built specifically for training, fine-tuning, and serving large AI models.

RAGFlow
Open-source RAG engine with deep document parsing, hybrid search, and visual agent orchestration.

PromptHub
Git-style prompt management, testing, and deployment platform for teams running multiple LLMs in production.

Rivet
Open-source visual IDE for building and debugging LLM agent graphs.

Deepgram
Production-grade speech-to-text, text-to-speech, and voice-agent APIs for real-time and batch audio.

Fireworks AI
Production inference and fine-tuning platform for open-source LLMs, tuned for speed and enterprise economics.

Patronus
Automated LLM evaluation for hallucinations, safety, and quality.

Tabnine
Privacy-focused AI autocomplete with on-prem options.

OpenMetadata
Open-source metadata platform that gives AI agents a semantic context graph over your data stack.

Agentset
Production-ready RAG infrastructure with agentic search, citations, and model-agnostic plumbing.

Epsilla
Agent-as-a-Service platform with managed RAG and a no-code builder for vertical enterprise AI.

OpenBB
Open-source financial workspace where analysts and AI agents share the same governed data.

Pathway
Live data framework for production RAG and streaming ETL pipelines in Python.

Seldon
Kubernetes-native MLOps platform for deploying and orchestrating ML and generative AI models in production.

Valohai
MLOps platform for versioned pipelines, distributed training, and LLM evaluation across any cloud.

Beam
Serverless GPU infrastructure for AI workloads with sub-second cold starts and bring-your-own-cloud support.

Domino Data Lab
Enterprise AI platform for building, deploying, and governing models and agents at scale.

Kiln AI
Open-source workbench for building, evaluating, and fine-tuning AI agents across 190+ models.

LobsterAI
Youdao's 24/7 personal agent that runs locally, plugs into multiple LLMs, and automates office workflows via dialogue and skills.

Nexent
Open-source, zero-code platform for spinning up production-grade AI agents from a single natural-language prompt.

OneKE
Open-source multi-agent framework for schema-guided knowledge extraction from documents.

Plano
Envoy-based data plane for AI agents that handles routing, guardrails, and observability outside your app code.

Portkey AI Gateway
Open-source AI gateway that routes a single API call across 1,600+ LLMs with caching, fallbacks, and observability.

SAS Viya
Enterprise-grade data and AI analytics platform with built-in governance, MCP server, and a Copilot for regulated industries.

Wallaroo.AI
Production AI inference platform for deploying and monitoring models across cloud, on-prem, and edge.

WeKnora
Tencent's open-source RAG framework that turns raw documents into a queryable knowledge base, ReAct agent, and self-maintaining wiki.

Iguazio
Enterprise MLOps and GenAI platform for taking models from notebook to production at scale.

Maxim AI
End-to-end evaluation, simulation, and observability platform for shipping production-grade AI agents.

Portkey
Production LLM gateway with observability, guardrails, and prompt management for teams shipping AI in anger.

Prompt Foundry
Prompt management and side-by-side LLM evaluation for OpenAI and Anthropic models.

Tableau
Salesforce-owned BI platform that bolted generative AI onto enterprise dashboards via Tableau Pulse and Tableau Agent.

AICamp
Team workspace that puts GPT, Claude, and Gemini behind one admin console with shared prompts, agents, and usage controls.

Bisheng
Open-source enterprise AgentOps platform for building, orchestrating, and governing LLM agents at scale.

Kotaemon
Open-source RAG UI for chatting with your own documents, locally or self-hosted.