Hyper-Extract
A CLI and MCP server that extracts structured knowledge from documents: graphs, hypergraphs, and fact search with sources
Medium risk
We rate an entry medium when the tool runs code, makes network calls or reads project files. Check what exactly it does before installing.
Why this level
- Sends document content to the LLM and embedding provider to build the graph
- Stores provider keys locally in a configuration file
Install
In your terminal, with SkillFoxx CLI
npx skillfoxx add cli/hyper-extractDetects the agents on your machine, checks the risk and pins the version.
Other ways to install
Install the tool
pipx install hyperextractWithout pipx, pip install hyperextract works too.
Install the CLI: uv tool install hyperextract (or pipx install hyperextract). Configure a provider: he config init -p openai -k YOUR_KEY. For the MCP server, install pip install 'hyperextract[mcp]' and run he-mcp.
Other ways from the author
uv tool install hyperextractAlternative: pipx install hyperextract.
This is third-party code. Review the repository files before installing.
What it does
Hyper-Extract parses unstructured text and builds a knowledge graph from it using one of the built-in templates, such as a biography, timeline, or report. Every fact in the graph is tied back to its source in the original document, so a claim can be traced to where it came from. The resulting knowledge base can be queried with search or RAG-style questions, and exported to Obsidian, GraphML, CSV, JSON-LD, or Cypher. A separate MCP server gives agents such as Claude Desktop read and export access to already built knowledge bases without changing their content.
Who it is for. For analysts and technical writers who need to turn a set of documents into a verifiable knowledge base.
Good fit when
- You need to build a knowledge graph from articles, reports, or documentation
- You need to quickly find facts tied back to their source inside a large document
- You need to export extracted knowledge to Obsidian or a graph database
Not a fit when
- You just need a plain text summary without graph structure
- You have no access to any LLM or embedding provider
Example request
Build a knowledge graph from this report using the biography_graph template and find all the achievements mentioned in itLimitations
Extraction and search need an LLM and embedding provider key (OpenAI, DeepSeek, Anthropic, Google Gemini, Alibaba Bailian, or a local vLLM). Some cloud providers are unavailable from Russia without a VPN, while pairing DeepSeek with any embedder works as a cheaper option.
How to disable. Remove the package: uv tool uninstall hyperextract (or pipx uninstall hyperextract) and delete the ~/.he/config.toml config file.
Security check
- Sends document content to the LLM and embedding provider to build the graph
- Stores provider keys locally in a configuration file
README in short
The README describes Hyper-Extract as a CLI that turns documents into structured knowledge with one command. The tool builds graphs and hypergraphs from templates, supports several LLM and embedding providers, and exports to multiple formats including Obsidian. A separate, read-only MCP server is documented that connects knowledge bases to Claude Desktop and other MCP clients. It is Apache-2.0 licensed and distributed via PyPI.
FAQ
How is this different from plain RAG?
Hyper-Extract first builds an explicit fact graph tied to sources, and only then can you search or ask questions over it, instead of just indexing text chunks.
Can the MCP server modify my knowledge base?
No, it is read and export only for already built knowledge bases; you cannot create or delete them through it.
Related
166 skills for scientific work: bioinformatics, cheminformatics, clinical data, geospatial analysis and 100+ databases
Google's open-source MCP server for databases: ready tools for Postgres, MySQL, BigQuery, Spanner and more, plus custom tools in tools.yaml
Official Hugging Face skills: Hub operations via the hf CLI, datasets, model training, Spaces, evals and deployment
A visualization language for agents and an MCP server: neat charts from a simple semantic spec