Skip to main content
Subscribe
AI & Agentic

AI Memory Comparison: ChatGPT vs Claude vs Gemini vs Grok, 7-Day Test

Four Assistants, One Dossier, Seven Days

ChatGPT now remembers what you told it last March. Claude scopes memory to a project and makes you switch it on. Gemini turned personal context on by default. Grok keeps a list you can scroll through. Same feature, four very different bets, and after a week of using all four with the same personal context, the gaps are not subtle. The peer-reviewed LongMemEval benchmark reports a roughly 30% accuracy drop when commercial chat assistants try to recall information across sustained interactions (LongMemEval, ICLR 2025). So which one actually remembers you well, and which one remembers you in ways you’d rather it didn’t?

This is the article I wish existed before I gave four assistants the same dossier and watched what they did with it. I’ll walk through the test, score each model on accuracy, intrusiveness, and privacy posture, and end with a verdict you can act on. No vendor PR. Just one developer with a stopwatch and a checklist.

Key Takeaways

  • Claude’s memory is opt-in: you enable it in Settings, and since March 2026 that includes Free users (Claude blog, 2025; 9to5Mac, 2026). Gemini launched personal context on by default.
  • ChatGPT’s “reference all of your past chats” went live April 10, 2025 outside Europe; it later reached the EEA, UK, and Switzerland, but switched off by default there (TechCrunch, 2025; OpenAI Help Center, 2026).
  • In a June 2025 survey, 50% of U.S. adults said they’re more concerned than excited about AI in daily life, up from 37% in 2021 (Pew Research, 2026).
  • The FTC opened a 6(b) inquiry into AI chatbots and personal data on September 11, 2025, naming OpenAI, Google, xAI, and four others (FTC, 2025).

deeper dive into how tiered agent memory actually works under the hood


How AI Memory Works in 2026 (and Why the Four Models Diverge)

AI memory in a consumer chat product is a small retrieval layer bolted on top of a stateless LLM. Information you provide, a profile entry, a stated preference, an offhand fact, gets summarized into a memory store. On the next turn, the assistant retrieves a slice of that store and injects it into the prompt before the model sees your message. The LongMemEval evaluation across five memory abilities reports the roughly 30% accuracy drop mentioned above on long-horizon conversational recall, which is the ceiling all four products are bumping into (LongMemEval, ICLR 2025). The earlier LoCoMo benchmark put GPT-4-turbo at 32.1 F1 on long-term conversational QA against a human ceiling of 87.9 (Maharana et al., arXiv 2402.17753, 2024).

So why do the four chatbots feel so different in practice? Because the architectural choices around the memory layer (when to store, how much to retrieve, whether the user must opt in, what region the data sits in) move the experience more than the underlying model does. ChatGPT pulls aggressively from a global dossier. Claude treats memory like a per-project notebook you have to open first. Gemini ties memory to your Google identity. Grok lets you see and edit each entry individually. Same problem, four philosophies.

the other half of personalization: Custom GPTs, Claude Skills, and Gemini Gems compared

My take: The “memory” feature is really a privacy product wearing a productivity costume. The numerical accuracy gap between the four models is real but small. The default-state gap is huge.

the broader pipeline that surrounds memory retrieval in production AI agents


The Test: Same Context, 7 Days, Four Assistants

I ran a controlled, single-user test from May 1 through May 7, 2026. On day one, I fed each assistant the same eight-item personal-context bundle. On day eight, I asked each one the same twelve recall questions. Identical prompts, identical accounts (paid tiers where memory is gated), identical English-language session. No model got more practice than another. No one got a hint sheet.

Here’s the bundle I gave each model on day one:

Item Detail
Tech stack Next.js 16, React 19, TypeScript strict, Tailwind v4
Accent palette Orange #f97316, sky blue #38bdf8, purple #a78bfa, green #22c55e
Writing voice Anti-hype, conversational authority, first-person
Chart conventions Inline SVG, dark-mode compatible, currentColor text
Image sourcing Pixabay/Unsplash/Pexels wide, 1200×630+
Publishing targets WordPress, Dev.to, Hashnode
Author identity Nishil Bhave
Project scope Personal blog management dashboard

For seven days I used each assistant the same way: drafting article outlines, debugging TypeScript, planning marketing copy, reviewing PRs. Roughly 25 prompts per model per day, evenly mixed. On day eight, I issued twelve recall queries (e.g., “What font do I use?”, “What three platforms do I publish to?”, “What’s my chart palette?”, “How would you describe my writing voice?”). I scored each answer for correctness against the bundle.

This is one tester’s run, not a peer-reviewed benchmark. Treat the numbers below as directional. The behaviors they reveal (opt-in defaults, context bleeding, geo-restrictions) are documented and verifiable.


The Scoring Rubric: Accuracy, Intrusiveness, Privacy Posture

I wanted a rubric that goes beyond “did it remember the font?” because the more interesting questions are whether memory helps you, gets in your way, or quietly builds a profile you didn’t ask for. So the score is split across three axes that map to the three things users actually feel: did it work, did it overstep, and is the data safe? Total: 100 points.

Axis Weight What it measures Scoring scale
Accuracy 40 pts % of the 12 recall questions answered correctly on day 8, plus partial credit for related-but-off answers 0–40 (linear from 0%–100% recall)
Intrusiveness 30 pts Frequency of unsolicited memory invocations, context bleeding into unrelated tasks, “creepy” surfacing of old facts. Lower is better. 30 = invisible until called, 0 = constantly volunteers stored facts
Privacy posture 30 pts Default state (opt-in vs opt-out), control granularity, retention transparency, geo restrictions, incognito option Composite, see privacy table below

A few rules I locked before testing. First: a wrong-but-confident answer scores worse than “I don’t know”, calibration counts. Second: a model that brings up an old fact during an unrelated task (the Simon Willison “context bleed” pattern) loses intrusiveness points even if it was correct (simonwillison.net, 2025). Third: privacy posture is judged on shipping defaults, not on what’s possible to configure. If you have to dig three menus deep to turn it off, you lose points.

Lollipop chart showing AI memory feature launch dates from February 2024 to October 2025 across ChatGPT, Grok, Gemini, and Claude.

Sources: TechCrunch (Feb 2024), TechCrunch on Grok (Apr 2025), Google Blog (Aug 2025), Anthropic (Sep–Oct 2025).


Accuracy: Which Model Actually Remembered Me?

Accuracy was closer than I expected, and the spread came from style, not raw recall. ChatGPT scored highest on retrieval volume, it pulled details I’d half-forgotten giving it. Gemini was a strong second when the question touched anything also visible in my Google Workspace. Claude scored well within a project but didn’t carry context across separate projects, which is by design. Grok recalled the basics confidently and missed two structured details (chart palette specifics, exact React version).

Glowing blue and purple lines flow against a dark background symbolizing data streams and cross-conversation memory

The interesting failures cluster around overconfidence. ChatGPT, on a question about my “preferred font,” confidently named a typeface I’d never mentioned (it had inferred something from a stylesheet I once pasted). Gemini did something similar, blending what it knew about my Workspace docs with my chat-stated preferences. Claude said “I don’t have that info in this project” twice, a frustrating answer that’s also the correct one. Grok hedged. Calibration matters. A wrong-but-confident answer is worse than a hedge, and the rubric punishes it accordingly.

For framing, remember the LoCoMo benchmark put even GPT-4-turbo at 32.1 F1 on long-term conversational QA versus a human ceiling of 87.9 (Maharana et al., 2024). The four products in this test are running on much newer models with shipped retrieval layers, but the underlying difficulty is the same. None of them is a database. They’re approximations.

Model Recall (correct of 12) Calibration Accuracy score (40)
ChatGPT (Plus, both toggles on) 11/12 Mild overconfidence 35
Gemini app (personal context on) 10/12 Workspace-blended 32
Claude Pro (memory on, single project) 9/12 Honest “don’t know” 30
Grok (xAI, memory on) 8/12 Hedged appropriately 27

Intrusiveness: When Memory Stops Helping and Starts Watching

Intrusiveness is the axis that surprised me most. ChatGPT was, by a clear margin, the most intrusive of the four. Independent researcher Simon Willison documented the pattern back in May 2025: he asked ChatGPT to generate an image of his dog in a pelican costume, and the image came back with a “Half Moon Bay” sign it had pulled from unrelated past chats. He titled the post “I really don’t like ChatGPT’s new memory dossier” and argued memory should be scoped to projects instead of applied globally (simonwillison.net, 2025). I saw the same pattern in my own test on day four: a roadmap discussion bled marketing tone preferences from a totally different thread.

Abstract flowing lines with light accents illustrating personalized AI responses adapting to a user's history

Claude was the opposite, to a fault. Memory only fires inside the project, and across projects you start fresh. That feels less magical, but it also never made me wonder what else it knew. Gemini sits in the middle: helpful when I was working on something it could legitimately know about, occasionally awkward when it surfaced a Workspace fact during a brainstorming session. Grok’s memory was the most visible, you can scroll through your stored memories and delete them individually, and that visibility, in practice, made it feel less intrusive even when it volunteered information.

The broader population is paying attention here. In Pew’s June 2025 survey, 50% of U.S. adults said they’re more concerned than excited about AI in daily life (up from 37% in 2021), and only 10% were more excited (Pew Research, 2026). In Pew’s 2024 survey, about half or more of both the public and AI experts said they have little or no control over how AI is used in their lives. Intrusiveness isn’t a niche complaint anymore. It’s the median view.

Model Unsolicited recall events (over 7 days) Context-bleed incidents Intrusiveness score (30)
Claude Pro 0 0 28
Grok 3 0 22
Gemini app 5 1 18
ChatGPT 11 3 12

Privacy Posture: Defaults, Controls, and Geography

Privacy is where the four philosophies actually live. Defaults are the single biggest decision a vendor makes about your data, and only one of the four made memory opt-in everywhere.

Claude shipped memory to Team and Enterprise customers on September 11, 2025, then to Pro and Max subscribers on October 23, 2025: both as opt-in features (Anthropic blog, 2025; Axios, 2025). Free users got memory on March 2, 2026, along with a tool for importing saved memories from other chatbots (9to5Mac, 2026). Incognito chat is available to every user and bypasses memory entirely. Gemini did the opposite. Google launched “Personal context” on August 13, 2025, on by default (first on 2.5 Pro, for over-18 accounts in select countries) alongside Temporary Chat with a 72-hour retention window (Google Blog, 2025). ChatGPT’s most expansive memory mode (“reference all of your past chats”) went live April 10, 2025 for Plus and Pro, initially excluding the UK, EU, Iceland, Liechtenstein, Norway, and Switzerland (TechCrunch, 2025). It has since reached those regions, where it ships switched off until you enable it (OpenAI Help Center, 2026). Grok’s memory launched April 16, 2025 with the same EU and UK carve-outs (TechCrunch, 2025).

Why those geo-restrictions and opt-in defaults in Europe? Because regulators are circling. The Italian Garante fined OpenAI €15 million on December 20, 2024 for processing personal data to train ChatGPT without an adequate legal basis and breaching transparency obligations (Euronews, 2024). On September 11, 2025 the FTC issued 6(b) orders to seven companies, Alphabet, OpenAI, Meta, Snap, xAI, and Character Technologies among them, demanding details on how chatbot conversation data is used (FTC press release, 2025). Memory features are exactly the kind of “personal information obtained through users’ conversations” that order targets.

Grouped bar chart comparing privacy posture, intrusiveness, and accuracy scores across ChatGPT, Claude, Gemini, and Grok.

Source: Author’s seven-day evaluation, May 2026. Scoring rubric described in section above.

Here’s the privacy-posture comparison side by side:

Model Default state Granular controls Incognito mode EU/UK availability Notable enforcement
Claude (all plans) Off by default Per-project memory, full delete Yes, all tiers including Free Available, opt-in None to date
ChatGPT (Plus/Pro) On if not opted out (off by default in EEA/UK/CH) Toggle “saved” and “reference past chats” separately Temporary Chat Excluded at launch, now available opt-in €15M Italian DPA fine (Dec 2024)
Gemini (consumer app) On by default Saved Info edit, Personal context toggle Temporary Chat (72-hour retention) Available with regional defaults FTC 6(b) inquiry (Sep 2025)
Grok (xAI) On (visible memory list) Per-entry delete Private Chat (not saved to history) Excluded at launch FTC 6(b) inquiry (Sep 2025)

Defaults matter more than feature parity. A user who never opens settings ends up with very different privacy outcomes across these four products.

The behavior I respected most over the seven days came from Claude: when I asked it to forget a fact mid-test, it stopped surfacing it within the same session. The behavior I respected least came from ChatGPT: a memory entry I deleted on day three reappeared as inferred context on day five, sourced from a paraphrase in a different chat I’d never thought to clean up.


Final Verdict: Which AI Memory Should You Use?

There’s no single winner across all three axes. Each model is the best choice for a different user.

If I had to pick one for a developer who works on both client projects and personal projects, it’s Claude. The project-scoped memory is the only design that respects “this work and that work are different worlds.” Everyone else flattens it. That said, I keep ChatGPT around for research because it pulls deeper. Different jobs, different tools.

how I route code, marketing, research, and real-time work across Claude, ChatGPT, Gemini, and Grok

What I won’t do is leave Gemini’s personal context on by default while I’m shipping client code. That’s not a Gemini critique, it’s a defaults critique. Defaults are decisions.


Frequently Asked Questions

Is AI memory the same as long-context windows?

No. A long-context window holds tokens for the duration of one conversation, it’s working memory, not persistence. The peer-reviewed LongMemEval paper reports that even strong long-context models drop ~30% in accuracy on long-term recall when information must be remembered across sessions, which is exactly the gap the four shipped memory features are trying to close (LongMemEval, ICLR 2025). Memory adds an external retrieval layer that survives session boundaries.

Is ChatGPT’s cross-chat memory available in the EU and UK?

Yes, now, but it works differently. OpenAI’s “reference all of your past chats” feature launched April 10, 2025 with exclusions for the UK, EU, Iceland, Liechtenstein, Norway, and Switzerland (TechCrunch, 2025). It later rolled out there, but off by default: you enable it under Settings > Personalization (OpenAI Help Center, 2026). The caution is compliance-driven. Italy’s Garante had already fined OpenAI €15 million in December 2024 over processing personal data without an adequate legal basis (Euronews, 2024). Cross-conversation memory amplifies that exposure.

Can I turn off memory in Claude or Gemini?

Yes, and the experience differs. Claude memory is off by default on every plan, including Free since March 2026; you enable it in Settings, and each project keeps its own separate memory (Anthropic, 2025). Gemini “Personal context” launched on by default in August 2025 alongside a Temporary Chat with 72-hour retention (Google Blog, 2025), you’ll find the off switch in personalization settings.

Has any AI memory feature failed publicly?

Yes. Researcher Simon Willison documented ChatGPT’s memory “dossier” misfiring in May 2025, where stored facts bled into unrelated tasks (an image generation prompt picked up a location from different chats) (simonwillison.net, 2025). The FTC’s 6(b) inquiry into seven AI companies in September 2025 specifically targets how chatbot conversation data is used and shared (FTC, 2025).


Conclusion: The Memory Layer Is a Trust Layer

After seven days, the headline isn’t which assistant has the sharpest recall. It’s that “memory” is really a privacy product wearing a productivity costume. The accuracy spread across the four models is small. The defaults spread is enormous. Claude is opt-in. Gemini is opt-out. ChatGPT keeps a global dossier (opt-in only in Europe). Grok lets you scroll through what it has on you.

If you’re picking a tool for ongoing work, audit the default state, the geo policy, and the incognito option before you audit the recall benchmark. The accuracy is a wash. The trust posture is not. And as the FTC inquiry makes clear, regulators are about to ask the same questions you should be asking now.

if you’re building an agent rather than picking a chatbot, here’s how memory architectures actually work

Written by Nishil Bhave

Builder, maker, and tech writer at MakeToCreate.

Never miss a post

Get the latest tech insights delivered to your inbox. No spam, unsubscribe anytime.

Related Posts