Skip to main content
📖 The AI Tool Bible
OneKE preview image
OneKE logo

OneKE

Open-source multi-agent framework for schema-guided knowledge extraction from documents.

Free· Free, MIT-licensed; you pay for LLM API calls or self-hosted computeRAGMulti-model (OneKE-13B, LLaMA3, Qwen2.5, GPT, DeepSeek-R1)7.2 / 10
Visit website →

In short

OneKE extracts structured facts like entities and relations from various file formats using a multi-agent LLM pipeline. It is best for engineers who need a flexible, open-source tool to build knowledge graphs without managing complex infrastructure.

Best for

Pick OneKE if you're building a domain knowledge graph and want a flexible, open-source extraction pipeline that runs against either API or local LLMs.

Skip if

Skip it if you need a managed SaaS extraction API, English-first docs, or a turnkey solution without DevOps work.

OneKE is a dockerized, MIT-licensed knowledge extraction system from the ZJU-NLP lab and Ant Group's OpenSPG project. It uses a multi-agent LLM pipeline (schema agent, extraction agent, reflection agent) to pull structured facts out of plain text, HTML, PDF, Word, JSON, and TXT files, covering NER, relation extraction, event extraction, triple extraction, and open-ended information extraction. Output can be assembled directly into a visualizable knowledge graph.

The project's differentiator is flexibility on both ends: you can plug in OpenAI or DeepSeek-R1 via API, or run it fully locally against LLaMA3, Qwen2.5, ChatGLM4, MiniCPM3, or the bundled OneKE-13B model with optional vLLM acceleration. Schemas can be default, predefined, or self-deduced by the agent, and case-retrieval plus reflection loops let you trade speed for accuracy. It's aimed at researchers and engineers building domain-specific KGs who don't want to wire up extraction infrastructure from scratch.

Deployment is via Docker or Conda, with a Streamlit web UI for interactive runs and a HuggingFace Spaces demo. As an open-source academic-led project, the polish lags commercial extraction APIs, and the Yuque-hosted user guide is mostly Chinese, but the breadth of supported tasks, models, and file types is rare at this license tier.

Editor's take

OneKE is one of the more serious open-source attempts at productionizing LLM-based information extraction, and the multi-agent schema/reflection design is genuinely useful. The catch is that it's an academic-flavored release; expect to read Chinese docs and do real integration work. Worth it if you'd otherwise glue together LangChain agents yourself.

— The AI Tool Bible editorial team

Pros

  • Covers NER, RE, EE, and triple extraction in one framework
  • Works with API models or fully local LLMs via vLLM
  • Ingests PDF, Word, HTML, JSON, and plain text out of the box
  • Multi-agent schema + reflection loop improves extraction quality
  • MIT license with Docker and Streamlit UI included

Cons

  • ⚠️ Documentation is primarily Chinese and scattered across Yuque/GitHub
  • ⚠️ Self-hosting and tuning agents is non-trivial for non-researchers
  • ⚠️ No managed cloud offering; you bring the infrastructure
  • ⚠️ Quality depends heavily on the underlying LLM you wire in

Use cases

knowledge-graph-constructionnamed-entity-recognitionrelation-extractionevent-extractiondocument-parsing

Frequently asked

What types of files can OneKE process?
OneKE supports plain text, HTML, PDF, Word, JSON, and TXT files for knowledge extraction.
Can OneKE run locally without external API costs?
Yes, it can run fully locally against models like LLaMA3, Qwen2.5, or the bundled OneKE-13B, with optional vLLM acceleration.
What specific extraction tasks does OneKE support?
It covers named entity recognition, relation extraction, event extraction, triple extraction, and open-ended information extraction.
Is OneKE suitable for teams needing managed SaaS support?
No, it is an open-source project requiring self-hosting via Docker or Conda, and it lacks a managed cloud offering.

Explore related

Compare with similar tools

All in RAG
Pinecone preview image
Pinecone logo

Pinecone

Featured
RAG · Hosted vector DB (not an LLM)
8.8

Managed vector database for production-scale similarity search.

Freemium· Starter: Free · Builder: $20/month flat · Standard: $50/month min. usage · Enterprise: $500/month min. usagemanaged vector DBproduction RAG
LlamaIndex preview image
LlamaIndex logo

LlamaIndex

Featured
RAG · BYO (Claude / GPT / open)
8.7

Data framework for connecting LLMs to your data.

Freemium· Free open-source; LlamaCloud paidRAGdata ingestion
Elasticsearch Vector Search preview image
Elasticsearch Vector Search logo

Elasticsearch Vector Search

RAG · BYO embeddings (OpenAI, Cohere, Hugging Face, Mistral, Bedrock, Vertex, Azure) plus Elastic's built-in ELSER sparse model and E5 dense model
8.7

Hybrid vector + keyword search in the enterprise-grade Elasticsearch engine

Freemium· Resource based pricing: Pay as you go (monthly) or prepaid · Usage based pricing: Pay as you go (monthly) or prepaid · License based pricing: ?RAG chatbot over enterprise docsHybrid semantic + keyword product search
Snowflake Cortex preview image
Snowflake Cortex logo

Snowflake Cortex

RAG · Anthropic Claude, Meta Llama, Mistral Large 2, Snowflake Arctic
8.7

Generative AI and RAG built into the Snowflake data cloud

Enterprise· Standard: Contact sales · Enterprise: Contact sales · Business Critical: Contact sales · Virtual Private Snowflake: Contact salesEnterprise RAG chatbot over governed dataNatural-language SQL for business analysts
DataStax Astra DB preview image
DataStax Astra DB logo

DataStax Astra DB

RAG · Bring-your-own embeddings; integrates with OpenAI, Cohere, Hugging Face, Mistral, NVIDIA NIM, and Vertex AI via server-side vectorize
8.6

Serverless vector and document database for production RAG and AI agents

Freemium· Small On-Demand: Contact sales · Medium (Balanced): Contact sales · Medium (Storage Optimized): Contact sales · Large (Balanced): Contact sales · Large (Storage Optimized): Contact salesRAG chatbot over enterprise documentsAgent long-term memory store
MongoDB Atlas Vector Search preview image
MongoDB Atlas Vector Search logo

MongoDB Atlas Vector Search

RAG · Bring-your-own embeddings (OpenAI, Cohere, open models); native Voyage AI embeddings and rerankers
8.6

Vector search built into the operational database you're already using.

Freemium· Free: $0 · Flex: Up to $30 · Dedicated: Starts at $56.94RAG over enterprise documentsProduct and content recommendation engines