Hindsight
A long-term memory system for AI agents: the agent stores experience and facts, then retrieves them through a built-in MCP server
High risk
We rate an entry high when the tool writes to external systems, handles money, production databases or secrets, or runs arbitrary commands. The CLI installs it only with your consent.
Why this level
- Requires LLM provider keys and handles secrets
- Sends memory content to an external LLM provider for processing
- Runs a persistent server and database, and the CLI installer reads git history and past sessions
Install
Manual install
docker run -it --pull always --name hindsight --restart unless-stopped -p 8888:8888 -p 9999:9999 -e HINDSIGHT_API_LLM_API_KEY=$OPENAI_API_KEY -v hindsight-data:/home/hindsight/.pg0 ghcr.io/vectorize-io/hindsight:latestRecommended way to run the server in Docker. API on port 8888, web UI on 9999, data in the hindsight-data volume.
This is third-party code. Review the repository files before installing.
What it does
Hindsight gives an agent memory that does not just keep chat history but accumulates knowledge and experience over time. The server runs in Docker, via pip or in Kubernetes and stores data in PostgreSQL, and a developer works with three operations: retain stores information, recall searches memory, reflect produces an answer that accounts for what was accumulated. Each server exposes a built-in MCP endpoint per memory bank, so any MCP client gets these operations as tools. Beyond MCP there are clients for Python, Node.js and Go, an embedded server mode with no separate process, a two-line LLM wrapper, and a memory installer for CLI coding agents that builds a bank from git history and past sessions. Memory runs on top of more than twenty LLM providers, including local ones.
Who it is for. For developers building AI agents and LLM applications who want to give them long-term memory.
Good fit when
- The agent needs to remember facts and past experience across sessions, not only the current dialog
- You want memory exposed to the agent as the MCP tools retain, recall and reflect
- You need to give a CLI coding agent project memory from git history and past sessions
Not a fit when
- A short context within a single dialog is enough, with no external store
- You cannot run a server and database or connect an LLM provider
Example request
Remember that we deploy this service via docker compose, and remind me next time I ask about deploymentLimitations
This is self-hosted infrastructure: it needs a running server, PostgreSQL and an LLM provider key, and memory content is sent to the chosen provider for processing. Local providers (ollama, lmstudio, llamacpp) and any OpenAI-compatible endpoints are supported, so the system can stay fully local. A paid Hindsight Cloud option exists as an alternative to running your own server. On Intel Macs (x86_64) install the hindsight-all-slim build instead of hindsight-all.
How to disable. Stop and remove the server container or process and its data volume, uninstall the clients (pip or npm), and for CLI agents remove the added memory integration from their configuration.
MCP
- Transport
- http
- Authentication
- API key
| Environment variables | |
|---|---|
| HINDSIGHT_API_LLM_API_KEY required, secret | Key for the chosen LLM provider that powers memory. Local providers may not need a key. |
| HINDSIGHT_API_LLM_PROVIDER | Selects the LLM provider: openai, anthropic, gemini, ollama, lmstudio and others. Defaults to openai. |
| HINDSIGHT_DB_PASSWORD secret | Database password when running with an external PostgreSQL via docker compose. |
Security check
- Requires LLM provider keys and handles secrets
- Sends memory content to an external LLM provider for processing
- Runs a persistent server and database, and the CLI installer reads git history and past sessions
README in short
The README presents Hindsight as an agent memory system that makes agents learn, not just remember history. It claims strong results on the LongMemEval benchmark and production use. The quick start shows running the server in Docker, via pip and in Kubernetes, storage in PostgreSQL and connecting clients in Python, Node.js, Go and over a CLI. It describes a built-in MCP endpoint per bank, a two-line LLM wrapper, an embedded serverless mode and a memory installer for CLI coding agents. Memory works with more than twenty LLM providers, including local ones, MIT licensed.
FAQ
How does the agent connect to memory?
Each server exposes a built-in MCP endpoint per memory bank at an address like http://localhost:8888/mcp/<bank_id>/. Point an MCP client at it and the retain, recall and reflect operations become tools. There are also clients for Python, Node.js and Go.
Can it work without external cloud LLMs?
Yes. Besides cloud providers it supports local ones (ollama, lmstudio, llamacpp) and any OpenAI-compatible endpoint, so memory can stay entirely on your own infrastructure.
Related
Open-source personal AI assistant on your own machine: answers in Telegram, Slack, Discord and WhatsApp, extended with skills and plugins
A self-improving agent from Nous Research with a TUI, messaging gateway, cron jobs and skills it writes itself
An open source coding agent for the terminal and desktop with build and plan modes
Open prompt library with a Claude Code plugin, MCP server and CLI: search, fetch and improve prompts and skills from an agent