Four Assistants, One Dossier, Seven Days
ChatGPT now remembers what you told it last March. Claude scopes memory to a project and makes you switch it on. Gemini turned personal context on by default. Grok keeps a list you can scroll through. Same feature, four very different bets, and after a week of using all four with the same personal context, the gaps are not subtle. The peer-reviewed LongMemEval benchmark reports a roughly 30% accuracy drop when commercial chat assistants try to recall information across sustained interactions (LongMemEval, ICLR 2025). So which one actually remembers you well, and which one remembers you in ways you’d rather it didn’t?
This is the article I wish existed before I gave four assistants the same dossier and watched what they did with it. I’ll walk through the test, score each model on accuracy, intrusiveness, and privacy posture, and end with a verdict you can act on. No vendor PR. Just one developer with a stopwatch and a checklist.
Key Takeaways
- Claude’s memory is opt-in: you enable it in Settings, and since March 2026 that includes Free users (Claude blog, 2025; 9to5Mac, 2026). Gemini launched personal context on by default.
- ChatGPT’s “reference all of your past chats” went live April 10, 2025 outside Europe; it later reached the EEA, UK, and Switzerland, but switched off by default there (TechCrunch, 2025; OpenAI Help Center, 2026).
- In a June 2025 survey, 50% of U.S. adults said they’re more concerned than excited about AI in daily life, up from 37% in 2021 (Pew Research, 2026).
- The FTC opened a 6(b) inquiry into AI chatbots and personal data on September 11, 2025, naming OpenAI, Google, xAI, and four others (FTC, 2025).
deeper dive into how tiered agent memory actually works under the hood
How AI Memory Works in 2026 (and Why the Four Models Diverge)
AI memory in a consumer chat product is a small retrieval layer bolted on top of a stateless LLM. Information you provide, a profile entry, a stated preference, an offhand fact, gets summarized into a memory store. On the next turn, the assistant retrieves a slice of that store and injects it into the prompt before the model sees your message. The LongMemEval evaluation across five memory abilities reports the roughly 30% accuracy drop mentioned above on long-horizon conversational recall, which is the ceiling all four products are bumping into (LongMemEval, ICLR 2025). The earlier LoCoMo benchmark put GPT-4-turbo at 32.1 F1 on long-term conversational QA against a human ceiling of 87.9 (Maharana et al., arXiv 2402.17753, 2024).
So why do the four chatbots feel so different in practice? Because the architectural choices around the memory layer (when to store, how much to retrieve, whether the user must opt in, what region the data sits in) move the experience more than the underlying model does. ChatGPT pulls aggressively from a global dossier. Claude treats memory like a per-project notebook you have to open first. Gemini ties memory to your Google identity. Grok lets you see and edit each entry individually. Same problem, four philosophies.
the other half of personalization: Custom GPTs, Claude Skills, and Gemini Gems compared
My take: The “memory” feature is really a privacy product wearing a productivity costume. The numerical accuracy gap between the four models is real but small. The default-state gap is huge.
the broader pipeline that surrounds memory retrieval in production AI agents
The Test: Same Context, 7 Days, Four Assistants
I ran a controlled, single-user test from May 1 through May 7, 2026. On day one, I fed each assistant the same eight-item personal-context bundle. On day eight, I asked each one the same twelve recall questions. Identical prompts, identical accounts (paid tiers where memory is gated), identical English-language session. No model got more practice than another. No one got a hint sheet.
Here’s the bundle I gave each model on day one:
| Item | Detail |
|---|---|
| Tech stack | Next.js 16, React 19, TypeScript strict, Tailwind v4 |
| Accent palette | Orange #f97316, sky blue #38bdf8, purple #a78bfa, green #22c55e |
| Writing voice | Anti-hype, conversational authority, first-person |
| Chart conventions | Inline SVG, dark-mode compatible, currentColor text |
| Image sourcing | Pixabay/Unsplash/Pexels wide, 1200×630+ |
| Publishing targets | WordPress, Dev.to, Hashnode |
| Author identity | Nishil Bhave |
| Project scope | Personal blog management dashboard |
For seven days I used each assistant the same way: drafting article outlines, debugging TypeScript, planning marketing copy, reviewing PRs. Roughly 25 prompts per model per day, evenly mixed. On day eight, I issued twelve recall queries (e.g., “What font do I use?”, “What three platforms do I publish to?”, “What’s my chart palette?”, “How would you describe my writing voice?”). I scored each answer for correctness against the bundle.
This is one tester’s run, not a peer-reviewed benchmark. Treat the numbers below as directional. The behaviors they reveal (opt-in defaults, context bleeding, geo-restrictions) are documented and verifiable.
The Scoring Rubric: Accuracy, Intrusiveness, Privacy Posture
I wanted a rubric that goes beyond “did it remember the font?” because the more interesting questions are whether memory helps you, gets in your way, or quietly builds a profile you didn’t ask for. So the score is split across three axes that map to the three things users actually feel: did it work, did it overstep, and is the data safe? Total: 100 points.
| Axis | Weight | What it measures | Scoring scale |
|---|---|---|---|
| Accuracy | 40 pts | % of the 12 recall questions answered correctly on day 8, plus partial credit for related-but-off answers | 0–40 (linear from 0%–100% recall) |
| Intrusiveness | 30 pts | Frequency of unsolicited memory invocations, context bleeding into unrelated tasks, “creepy” surfacing of old facts. Lower is better. | 30 = invisible until called, 0 = constantly volunteers stored facts |
| Privacy posture | 30 pts | Default state (opt-in vs opt-out), control granularity, retention transparency, geo restrictions, incognito option | Composite, see privacy table below |
A few rules I locked before testing. First: a wrong-but-confident answer scores worse than “I don’t know”, calibration counts. Second: a model that brings up an old fact during an unrelated task (the Simon Willison “context bleed” pattern) loses intrusiveness points even if it was correct (simonwillison.net, 2025). Third: privacy posture is judged on shipping defaults, not on what’s possible to configure. If you have to dig three menus deep to turn it off, you lose points.

Sources: TechCrunch (Feb 2024), TechCrunch on Grok (Apr 2025), Google Blog (Aug 2025), Anthropic (Sep–Oct 2025).
Accuracy: Which Model Actually Remembered Me?
Accuracy was closer than I expected, and the spread came from style, not raw recall. ChatGPT scored highest on retrieval volume, it pulled details I’d half-forgotten giving it. Gemini was a strong second when the question touched anything also visible in my Google Workspace. Claude scored well within a project but didn’t carry context across separate projects, which is by design. Grok recalled the basics confidently and missed two structured details (chart palette specifics, exact React version).
The interesting failures cluster around overconfidence. ChatGPT, on a question about my “preferred font,” confidently named a typeface I’d never mentioned (it had inferred something from a stylesheet I once pasted). Gemini did something similar, blending what it knew about my Workspace docs with my chat-stated preferences. Claude said “I don’t have that info in this project” twice, a frustrating answer that’s also the correct one. Grok hedged. Calibration matters. A wrong-but-confident answer is worse than a hedge, and the rubric punishes it accordingly.
For framing, remember the LoCoMo benchmark put even GPT-4-turbo at 32.1 F1 on long-term conversational QA versus a human ceiling of 87.9 (Maharana et al., 2024). The four products in this test are running on much newer models with shipped retrieval layers, but the underlying difficulty is the same. None of them is a database. They’re approximations.
| Model | Recall (correct of 12) | Calibration | Accuracy score (40) |
|---|---|---|---|
| ChatGPT (Plus, both toggles on) | 11/12 | Mild overconfidence | 35 |
| Gemini app (personal context on) | 10/12 | Workspace-blended | 32 |
| Claude Pro (memory on, single project) | 9/12 | Honest “don’t know” | 30 |
| Grok (xAI, memory on) | 8/12 | Hedged appropriately | 27 |
Intrusiveness: When Memory Stops Helping and Starts Watching
Intrusiveness is the axis that surprised me most. ChatGPT was, by a clear margin, the most intrusive of the four. Independent researcher Simon Willison documented the pattern back in May 2025: he asked ChatGPT to generate an image of his dog in a pelican costume, and the image came back with a “Half Moon Bay” sign it had pulled from unrelated past chats. He titled the post “I really don’t like ChatGPT’s new memory dossier” and argued memory should be scoped to projects instead of applied globally (simonwillison.net, 2025). I saw the same pattern in my own test on day four: a roadmap discussion bled marketing tone preferences from a totally different thread.
Claude was the opposite, to a fault. Memory only fires inside the project, and across projects you start fresh. That feels less magical, but it also never made me wonder what else it knew. Gemini sits in the middle: helpful when I was working on something it could legitimately know about, occasionally awkward when it surfaced a Workspace fact during a brainstorming session. Grok’s memory was the most visible, you can scroll through your stored memories and delete them individually, and that visibility, in practice, made it feel less intrusive even when it volunteered information.
The broader population is paying attention here. In Pew’s June 2025 survey, 50% of U.S. adults said they’re more concerned than excited about AI in daily life (up from 37% in 2021), and only 10% were more excited (Pew Research, 2026). In Pew’s 2024 survey, about half or more of both the public and AI experts said they have little or no control over how AI is used in their lives. Intrusiveness isn’t a niche complaint anymore. It’s the median view.
| Model | Unsolicited recall events (over 7 days) | Context-bleed incidents | Intrusiveness score (30) |
|---|---|---|---|
| Claude Pro | 0 | 0 | 28 |
| Grok | 3 | 0 | 22 |
| Gemini app | 5 | 1 | 18 |
| ChatGPT | 11 | 3 | 12 |
Privacy Posture: Defaults, Controls, and Geography
Privacy is where the four philosophies actually live. Defaults are the single biggest decision a vendor makes about your data, and only one of the four made memory opt-in everywhere.
Claude shipped memory to Team and Enterprise customers on September 11, 2025, then to Pro and Max subscribers on October 23, 2025: both as opt-in features (Anthropic blog, 2025; Axios, 2025). Free users got memory on March 2, 2026, along with a tool for importing saved memories from other chatbots (9to5Mac, 2026). Incognito chat is available to every user and bypasses memory entirely. Gemini did the opposite. Google launched “Personal context” on August 13, 2025, on by default (first on 2.5 Pro, for over-18 accounts in select countries) alongside Temporary Chat with a 72-hour retention window (Google Blog, 2025). ChatGPT’s most expansive memory mode (“reference all of your past chats”) went live April 10, 2025 for Plus and Pro, initially excluding the UK, EU, Iceland, Liechtenstein, Norway, and Switzerland (TechCrunch, 2025). It has since reached those regions, where it ships switched off until you enable it (OpenAI Help Center, 2026). Grok’s memory launched April 16, 2025 with the same EU and UK carve-outs (TechCrunch, 2025).
Why those geo-restrictions and opt-in defaults in Europe? Because regulators are circling. The Italian Garante fined OpenAI €15 million on December 20, 2024 for processing personal data to train ChatGPT without an adequate legal basis and breaching transparency obligations (Euronews, 2024). On September 11, 2025 the FTC issued 6(b) orders to seven companies, Alphabet, OpenAI, Meta, Snap, xAI, and Character Technologies among them, demanding details on how chatbot conversation data is used (FTC press release, 2025). Memory features are exactly the kind of “personal information obtained through users’ conversations” that order targets.

Source: Author’s seven-day evaluation, May 2026. Scoring rubric described in section above.
Here’s the privacy-posture comparison side by side:
| Model | Default state | Granular controls | Incognito mode | EU/UK availability | Notable enforcement |
|---|---|---|---|---|---|
| Claude (all plans) | Off by default | Per-project memory, full delete | Yes, all tiers including Free | Available, opt-in | None to date |
| ChatGPT (Plus/Pro) | On if not opted out (off by default in EEA/UK/CH) | Toggle “saved” and “reference past chats” separately | Temporary Chat | Excluded at launch, now available opt-in | €15M Italian DPA fine (Dec 2024) |
| Gemini (consumer app) | On by default | Saved Info edit, Personal context toggle | Temporary Chat (72-hour retention) | Available with regional defaults | FTC 6(b) inquiry (Sep 2025) |
| Grok (xAI) | On (visible memory list) | Per-entry delete | Private Chat (not saved to history) | Excluded at launch | FTC 6(b) inquiry (Sep 2025) |
Defaults matter more than feature parity. A user who never opens settings ends up with very different privacy outcomes across these four products.
The behavior I respected most over the seven days came from Claude: when I asked it to forget a fact mid-test, it stopped surfacing it within the same session. The behavior I respected least came from ChatGPT: a memory entry I deleted on day three reappeared as inferred context on day five, sourced from a paraphrase in a different chat I’d never thought to clean up.
Final Verdict: Which AI Memory Should You Use?
There’s no single winner across all three axes. Each model is the best choice for a different user.
- Claude wins overall (85/100) if you value privacy posture and predictability above raw recall. The opt-in default, project scoping, and incognito-for-all matter compounded over months.
- ChatGPT wins on accuracy (60/100) if you want maximum recall and accept context bleeding as the price. In Europe you’ll have to switch cross-chat memory on yourself.
- Gemini wins on integration (67/100) if you’re already in Workspace and want memory tied to docs, calendar, and email, but check the default-on setting today.
- Grok wins on transparency (70/100) if you want to see and edit each memory entry one at a time. The list-based UI is the clearest of the four.
If I had to pick one for a developer who works on both client projects and personal projects, it’s Claude. The project-scoped memory is the only design that respects “this work and that work are different worlds.” Everyone else flattens it. That said, I keep ChatGPT around for research because it pulls deeper. Different jobs, different tools.
how I route code, marketing, research, and real-time work across Claude, ChatGPT, Gemini, and Grok
What I won’t do is leave Gemini’s personal context on by default while I’m shipping client code. That’s not a Gemini critique, it’s a defaults critique. Defaults are decisions.
Frequently Asked Questions
Is AI memory the same as long-context windows?
No. A long-context window holds tokens for the duration of one conversation, it’s working memory, not persistence. The peer-reviewed LongMemEval paper reports that even strong long-context models drop ~30% in accuracy on long-term recall when information must be remembered across sessions, which is exactly the gap the four shipped memory features are trying to close (LongMemEval, ICLR 2025). Memory adds an external retrieval layer that survives session boundaries.
Is ChatGPT’s cross-chat memory available in the EU and UK?
Yes, now, but it works differently. OpenAI’s “reference all of your past chats” feature launched April 10, 2025 with exclusions for the UK, EU, Iceland, Liechtenstein, Norway, and Switzerland (TechCrunch, 2025). It later rolled out there, but off by default: you enable it under Settings > Personalization (OpenAI Help Center, 2026). The caution is compliance-driven. Italy’s Garante had already fined OpenAI €15 million in December 2024 over processing personal data without an adequate legal basis (Euronews, 2024). Cross-conversation memory amplifies that exposure.
Can I turn off memory in Claude or Gemini?
Yes, and the experience differs. Claude memory is off by default on every plan, including Free since March 2026; you enable it in Settings, and each project keeps its own separate memory (Anthropic, 2025). Gemini “Personal context” launched on by default in August 2025 alongside a Temporary Chat with 72-hour retention (Google Blog, 2025), you’ll find the off switch in personalization settings.
Has any AI memory feature failed publicly?
Yes. Researcher Simon Willison documented ChatGPT’s memory “dossier” misfiring in May 2025, where stored facts bled into unrelated tasks (an image generation prompt picked up a location from different chats) (simonwillison.net, 2025). The FTC’s 6(b) inquiry into seven AI companies in September 2025 specifically targets how chatbot conversation data is used and shared (FTC, 2025).
Conclusion: The Memory Layer Is a Trust Layer
After seven days, the headline isn’t which assistant has the sharpest recall. It’s that “memory” is really a privacy product wearing a productivity costume. The accuracy spread across the four models is small. The defaults spread is enormous. Claude is opt-in. Gemini is opt-out. ChatGPT keeps a global dossier (opt-in only in Europe). Grok lets you scroll through what it has on you.
If you’re picking a tool for ongoing work, audit the default state, the geo policy, and the incognito option before you audit the recall benchmark. The accuracy is a wash. The trust posture is not. And as the FTC inquiry makes clear, regulators are about to ask the same questions you should be asking now.
if you’re building an agent rather than picking a chatbot, here’s how memory architectures actually work