1---2title: "When Your AI Assistant Forgets Who You Are"3date: "2026-09-20"4published: true5tags: ["ai", "hermes", "agents", "memory", "llm"]6author: "Gavin Jackson"7excerpt: "Meta's new Muse agent is built around a curated long-term memory, and every major AI vendor is racing to solve the same problem. Meanwhile my own assistants still get sudden dementia every few days. A look at how agent memory actually works, from Hermes internals to ChatGPT, Claude, Gemini, Apple and xAI's Grok Bot."8---910# When Your AI Assistant Forgets Who You Are11121314*The most famous computer that never forgot anything. It did not end well for the crew.*1516On 8 September, [Meta announced Muse](https://about.fb.com/news/2026/09/introducing-muse-personal-ai-agent/), billed as "the world's first personal AI agent built for everyone". It runs on its own secure virtual machine, works across your apps, and - the bit that matters - builds what Meta calls a "curated long-term memory". It learns from your conversations, reflects on what matters to you, and gets sharper the longer you use it.1718The coverage turned within days. Jason Aten at [Inc reported](https://www.inc.com/jason-aten/metas-new-muse-ai-agent-read-my-private-messages-i-never-asked-it-to/91408202) that Muse started referencing details from his private messages that he never gave it permission to see. [Futurism called it](https://futurism.com/artificial-intelligence/meta-muse-ai-agent-creepy) possibly the creepiest AI agent ever released. Both things can be true at once: persistent memory is the feature that makes a personal agent genuinely useful, and it is also the feature most likely to make you close the app and rethink your choices.1920I know the value of that memory from the other side, because I live with agents that keep losing theirs.2122## The dementia problem2324I have been running a personal agent at home since the early days of [OpenClaw](https://github.com/openclaw/openclaw), back when it was first released. These days Bob runs on [Hermes Agent](https://hermes-agent.nousresearch.com/docs), but the rhythm has been the same on both. For days at a time the experience is exactly what Meta is promising. Bob knows my servers, my projects, my family, the fact that I have banned a particular punctuation character from everything he writes for me. We build momentum across a week. Then one morning I say hello and get the conversational equivalent of a firm handshake from a stranger.2526Context window compressed. New session started. The relationship restarts from a two-page brief, and it is genuinely frustrating, like working with a very capable colleague who has anterograde amnesia and keeps it a secret until you notice.2728> ## What Sudden Dementia Actually Is29>30> A language model only sees what fits in its context window. When a conversation outgrows that window, the system compresses the history into a summary and carries on. Anything not captured in the summary is simply gone, and when a brand new session starts the model has only its small standing memory files. From the model's side there is no relationship to restart. There never was one. There was a 2,200-character brief and a cheerful greeting.3132## How the default memory actually works3334A recent Vectorize article, [How Hermes Agent Memory Actually Works (And How to Make It Better)](https://vectorize.io/articles/hermes-agent-memory-explained), does a good job of pulling this apart, and it matches what I see running Bob day to day. The short version: Hermes does not have one memory system, it has four, and most user frustration comes from treating them as one.3536- **Prompt memory (hot).** Two small files, MEMORY.md at roughly 2,200 characters and USER.md at roughly 1,375. They hold durable facts and a user profile, and they are loaded as a frozen snapshot into the system prompt at the start of every session. Frozen is deliberate - it keeps the LLM's prefix cache stable - but it means anything Bob learns mid-session only shows up next session. The budget is tiny by design, so the agent is constantly judging what deserves the space.37- **Session archive (cold).** Every session lands in a SQLite database that Bob can search when he needs episodic recall. The catch is that it is keyword-based full text search, and the agent has to decide to search it in the first place. Ask "what did I tell you about the auth service?" and if the stored transcript says "authentication microservice", you are relying on Bob picking the right query terms.38- **Skills (procedural).** After completing a complex task, Hermes writes a reusable skill document capturing what worked. This is the self-improving part, and it is separate from memory proper.39- **External provider (optional).** The pluggable layer I get to below.4041Put those together and the dementia stops being mysterious. The built-in layers are transparent, local and inspectable, which I value, but they are also *agent-curated*. Bob writes to memory when he judges something worth saving. Short sessions may produce nothing. Compression fires a last-minute memory flush, and anything not flagged in that flush evaporates. Nothing in the default stack knows that "Jo" and "my wife" are the same person.4243## The provider options4445The newer part of the Vectorize piece covers Hermes' pluggable memory providers, configured with a single command, `hermes memory setup`. Seven ship with it. The condensed version, per their breakdown:4647- **Hindsight** - local or cloud, structured facts and entities in a knowledge graph, a reflect operation that synthesises across everything stored, 94.6% on the LongMemEval benchmark48- **Honcho** - dialectic user modelling; builds a model of how you think, not just what you said (AGPL, worth noting if you self-host commercially)49- **Mem0** - cloud, freemium, fastest setup, 67.6% on the LongMemEval-S variant50- **OpenViking** - self-hosted, tiered memory loading that prioritises recent and frequently relevant items51- **Holographic** - local SQLite with zero extra dependencies, trust scoring and vector-algebra retrieval52- **RetainDB** - cloud, paid, hybrid search across vector, BM25 and reranking53- **ByteRover** - hooks in before context compression so in-flight facts are captured before they are summarised away5455Only one external provider runs at a time, alongside the always-on built-in layer.5657**Shared memory** is the team angle I had not thought through. Point Hindsight at a shared server, or use Hindsight Cloud or RetainDB, and every agent instance in a team reads and writes the same memory store. A new starter's agent knows on day one that the production database runs on a non-standard port and that the auth service depends on Redis, because someone else's agent already learned it. That is a genuinely interesting capability, and also a governance question I would want answered before enabling it anywhere near client work: whose facts are these, and who gets to forget them?5859## How the big vendors handle it6061Everyone is converging on the same problem from different directions.6263**ChatGPT** has the most familiar model: [two-part memory](https://help.openai.com/en/articles/8590148-memory-faq) made up of "saved memories" you explicitly ask it to keep, plus "reference chat history", where it automatically extracts insights from past conversations. The newer system updates memories itself and tries to manage contradictions rather than leaving you to curate a stale list. Convenient, and entirely on OpenAI's side of the fence.6465**Claude** has rolled persistent memory out broadly this year, and Anthropic's [platform docs](https://platform.claude.com/docs/en/agents-and-tools/tool-use/memory-tool) show the interesting bit: their memory tool operates client-side, with Claude requesting file operations that your own infrastructure executes. Claude Code does something similar with CLAUDE.md files and auto-accumulated learnings. Memory as files you can read, edit and delete, rather than a profile you cannot inspect. That philosophy will not surprise anyone who has used Hermes.6667**Gemini** calls it [personal context](https://www.theverge.com/news/758624/google-gemini-ai-automatic-memory-privacy-update): with the setting on, Gemini automatically recalls your "key details and preferences" from past chats and personalises responses without being asked. It sits on top of your Google account, which gives it a head start on knowing you and gives Google another reason to be careful with the toggle.6869**Apple** is the newest entrant. [Siri AI](https://www.apple.com/newsroom/2026/09/siri-ai-a-profoundly-more-capable-and-personal-assistant-is-here/), powered by the next generation of Apple Intelligence, is only just landing, with personal context understanding, onscreen awareness and systemwide app actions, and some capabilities like Live Rewind and Siri Recap still in staged rollout later this year. Apple's bet looks different in kind: memory as an on-device index over your own apps and data, rather than a learned cloud profile. Whether it actually *learns* you the way ChatGPT or Gemini do is an open question, and one the early commentary has already latched onto.7071**Meta**, back where we started, has gone furthest in the other direction: a dedicated VM, continuous learning, a memory that curates itself, and a privacy row inside a fortnight.7273Three philosophies, roughly. A cloud profile the vendor manages (OpenAI, Google, Meta), local files the agent curates and you can inspect (Hermes, Claude), and an on-device index (Apple). Convenience against control, as ever.7475The newest entrant does not fit any of the three, because it makes memory the product rather than a settings page. xAI, now trading as SpaceXAI after the February merger, launched [Grok Bot](https://x.ai/news/introducing-grok-bot) in August: always-on agents you message like a colleague, each with its own computer in the cloud, signing into the same tools you use and finishing jobs end to end. The pitch leans hard on exactly the mechanisms this post has been circling. Bots remember conversations and learn how you like things done. Show one a workflow once and it saves the steps as a routine, takes your corrections, and runs it on its own next time. Hermes users will recognise the pattern: it is the skills system, learned by watching instead of written by hand. Over time the Bots pick up your voice and your edge cases, and get proactive about chasing dropped threads without being asked.7677They also run in packs. You can run several Bots in parallel with a chief-of-staff Bot coordinating the specialists, and they message each other and share context in threads without you pasting notes between chats. In late September, Team Bots took that further: shared Bots with common files, tools, skills and memories for a whole team, while individual conversations stay private. That is the shared-memory governance problem from earlier, productised, with the awkward questions about whose facts win and who can edit the common memory still to be answered in public. The in-house numbers serve as the proof of concept: SpaceXAI runs its own support queue on Grok Bot, trained on more than a million past customer interactions, and [claims](https://x.ai/news/grok-bot-customer-support) a 175% ticket surge absorbed with no new hires, resolutions at US$0.20 to $0.30 each against the flat $1 to $4 typical of AI support tools, 99% of refunds closed without a human, and Grok Voice fielding more than 15,000 Starlink calls a day.7879## What I am doing about it8081Bob's dementia is fixable in principle, and the principle is `hermes memory setup`. Hindsight appeals because it runs locally and nothing leaves the machine, which suits how I run everything else. I am going to let it run for a few weeks and report back on whether the relationship survives past a fortnight, or whether we just get a better class of amnesia.8283If nothing else, it is some comfort that a company with Meta's resources looked at this problem and shipped something that reads your text messages. Memory is hard. At least mine asks first.8485---8687**Related:**8889- 📄 [How Hermes Agent Memory Actually Works (And How to Make It Better)](https://vectorize.io/articles/hermes-agent-memory-explained) - the Vectorize article this post references90- 🤖 [Hermes Agent docs](https://hermes-agent.nousresearch.com/docs) - my agent's official documentation91- 🚀 [Introducing Grok Bot](https://x.ai/news/introducing-grok-bot) - the launch post for xAI's AI teammates92- 🛠️ [How SpaceXAI is using Grok Bot to scale customer support](https://x.ai/news/grok-bot-customer-support) - the primary source for the Grok Bot support numbers93