Have an idea for Knownbase? Suggest a feature →

MCP memory

MCP memory servers: persistent memory for AI agents.

The Model Context Protocol gave agents a standard way to reach external tools. A memory server is the one that gives them somewhere to put what they learn. Here's what the category actually contains, how the options differ, and how to pick.

What an MCP memory server is

MCP is a protocol: a client (Claude Code, ChatGPT, Cursor) connects to a server, asks what tools it has, and calls them. The protocol says nothing about what a server does. A filesystem server exposes file operations; a database server exposes queries; a memory server exposes storage and retrieval of things the agent wants to remember.

At minimum that means three tools:

WriteThe agent saves a fact, decision or finding.
StoreOutside the context window. Survives the session.
Search & readA later agent retrieves only what's relevant.

That's the whole idea. Everything interesting is in the details: what gets written, who decides, how retrieval ranks, and whether the store is shared.

Why the context window isn't enough

A context window is working memory. It is large, it is fast, and it is completely volatile. Three things end it:

  • The session ends. New conversation, empty context.
  • Compaction runs. A long session summarizes its own earlier turns to make room, and summarization drops specifics first. What that costs in practice.
  • You switch tools. Whatever Claude Code learned is invisible to Codex, because they don't share a context. Unless something outside both holds it.

Bigger context windows help with the third-hardest of these and none of the first two. A million-token window still starts empty tomorrow.

The kinds of MCP memory server

"Memory server" covers several quite different designs. They're not interchangeable, and the differences matter more than feature lists suggest.

KindWhat it storesStrengthWeakness
Conversation memoryFacts extracted automatically from chatZero effort; captures things you'd never write downLow precision; the store fills with noise and stale claims
Knowledge graphEntities and relationshipsGood at "how does X relate to Y"Schema design is real work; brittle when reality doesn't fit the graph
Vector storeEmbedded chunks of textFinds things by meaning, not wordingNo structure, no status, no notion of superseded
Local file memoryMarkdown on your diskFully private; trivially inspectableOne machine, one person; no web clients
Curated project memoryDeliberate notes: decisions, discoveries, constraintsHigh precision; readable by humans; scoped per projectOnly as good as the writing habit

The sharpest divide in the category is automatic vs curated. Automatic systems extract memories from the conversation without being asked; curated systems store what the agent (or you) deliberately decides is durable. Automatic wins on recall and loses on precision, and for engineering knowledge precision is what you're actually buying — a memory store that confidently returns last quarter's abandoned architecture is worse than no memory store.

What to look for

  1. Project scoping. Memory from one codebase must not leak into another. Without it, an agent working on service A gets service B's constraints.
  2. Search that returns little. The point is to keep context clean. A memory tool that dumps full bodies into every result has moved the bloat rather than fixed it.
  3. Structure beyond text. Status, tags, and revision history are what let you tell a current decision from a superseded one.
  4. Multi-client reach. If it only works from one tool, it isn't project memory, it's tool memory.
  5. Export. Your project knowledge should be portable. Ask what a full export looks like before you put two years into it.
  6. Isolation and training policy. This is your codebase's reasoning. Check tenancy isolation and whether content is used for training. Ours, in detail.

The options people actually compare

Abstract categories only get you so far, so here are the concrete ones, and the case for each. Every tool below is a reasonable choice for the job it was built for; they were built for different jobs. Reviewed September 2026 — these projects move quickly, so check current details before deciding.

OptionShapeReachesBest when
Claude Code auto memoryMarkdown notes Claude writes itself, on by defaultOne machine, one repository, Claude Code onlyYou work solo, on one machine, with one agent
Mem0Memory infrastructure for agents and apps; extracts memories from interactions as well as explicit writes. Open source, cloud or self-hosted including air-gappedAnything you wire the API intoYou are building an application that needs a memory layer, and personalization matters more than engineering precision
Basic MemoryA knowledge graph written as plain Markdown files you own and can open in any editor. Local open-source, with a cloud optionMany MCP clients, plus Obsidian and your filesystemYou want the files on your own disk and full control of the format
KnownbaseCurated notes with type, status and supersession, searched on demandAny MCP client, any machine, shared across a workspace and a teamSeveral agents and people need one record of what a project has decided

Two questions separate them faster than any feature list. Who writes the memory? Automatic extraction costs nothing and fills the store with everything; deliberate writes cost a habit and keep the store trustworthy. Who has to read it? If the answer is "me, on this laptop, in Claude Code", the memory that ships with your agent is already enough and you should use it. The moment the answer includes a second machine, a second repository, a second agent or a second person, a store outside all of them is the only thing that works.

The case against Knownbase, stated fairly: it only contains what something deliberately wrote to it, so it is empty until a habit exists, and it is a hosted service rather than files on your disk. If neither of those trades appeals, one of the alternatives above will serve you better.

Knownbase as an MCP memory server

Knownbase is the curated-project-memory kind: hosted, multi-client, project-scoped, and deliberate about what gets written.

{
  "mcpServers": {
    "knownbase": { "url": "https://knownbase.dev/mcp" }
  }
}

Tools the agent gets: search_notes, get_note, get_notes, upsert_note, delete_note, list_projects, list_note_revisions, get_note_revision, rename_project, rename_tag, workspace_info, usage_summary and backup_project. Two design choices worth calling out:

  • Search is lean by default. search_notes returns ids, titles, snippets and metadata — not full bodies — so finding the right note costs a fraction of reading ten wrong ones. includeBody:true opts back in when you want it.
  • Updates are partial. upsert_note leaves omitted fields alone, so an agent fixing one tag can't silently blank the body it didn't send.

Authentication is either an Authorization: Bearer kb_... key or OAuth 2.1 with dynamic client registration and PKCE, so both static-config clients and interactive connector UIs work against the same URL. Keys can be read-only or restricted to one project.

Connecting a client

Claude Code

claude mcp add --transport http knownbase https://knownbase.dev/mcp. Full setup →

Codex

Add the endpoint to your MCP config and connect. Full setup →

Cursor

One entry in mcp.json, no headers needed with OAuth. Full setup →

ChatGPT & Claude Desktop

Settings → Connectors → add a custom connector with the URL. Docs →

Keep reading

FAQ

What is an MCP memory server?

An MCP server whose tools let an agent store and retrieve information that outlives a conversation. The Model Context Protocol is a standard interface between AI clients and external tools; a memory server implements that interface over a store of notes, facts or embeddings. From the agent's side it's just another set of tools, typically something like search, read and write.

How is an MCP memory server different from RAG?

RAG is retrieval over a corpus that already exists: your docs, your wiki, your codebase. A memory server is retrieval over a corpus the agent itself writes as it works. The retrieval mechanics overlap heavily; the difference is where the content comes from and whether the agent can add to it. Plenty of setups want both.

Do I need to run my own MCP server?

Only if you want to. Local memory servers run on your machine and store to a local file or database, which is the right answer if data must never leave your laptop. A hosted server is a URL, so it works from any machine, from a phone, and from web clients like ChatGPT that cannot reach a local process, and teammates can share one workspace.

Which clients can use an MCP memory server?

Any MCP client. In practice that includes Claude Code, Claude Desktop and Claude.ai, ChatGPT, Codex, Cursor, Windsurf, Zed, and a growing list of agent frameworks. A remote server also needs the client to support remote (HTTP) MCP servers rather than only local stdio ones; most current clients do.

How does authentication work?

Two common models. A static API key sent as an Authorization bearer header works everywhere, including clients that only support static config. OAuth 2.1 with dynamic client registration and PKCE lets a client sign a human in interactively, which is what Claude's connector UI and ChatGPT use. Knownbase supports both against the same endpoint.

Won't the agent just fill it with junk?

It can, if you let it. This is the real failure mode of automatic memory systems: they capture everything, and the store becomes noise that pollutes context rather than informing it. Curated memory (the agent writes deliberately, in response to a rule you set, for durable facts only) trades a little recall for a lot of precision. That trade is usually worth it for engineering knowledge.

A hosted MCP memory server, connected in a minute

One endpoint, OAuth or an API key, and every MCP client you use reads the same project memory. Free plan, no card required.

Create free workspace