Wood Chen

Why I Dropped memory-lancedb for OpenClaw's Official Memory System After Costly Real-World Testing

4 comments1.2k views1.2k words

This post was translated from Chinese by AI. If anything reads oddly, the Chinese original is authoritative. 中文原文

A few days ago, I spent some time tinkering with OpenClaw's memory system. The reason was simple: I installed the memory-lancedb plugin after seeing someone recommend it on Douyin.
The video hyped it up, making it sound like installing it would instantly supercharge memory: long-term memory, smart recall, and a system that understands you better the more you use it. It all sounded convincing.

I was persuaded at the time. I figured the official defaults might be too conservative, and memory-lancedb could be the more powerful “advanced option.”
After actually using it, though, my conclusion was straightforward: it just wasn't good.

This wasn't a case of “trying it a few times and jumping to conclusions.”
I used it for several days, and the workload was substantial:

  • My primary models were gpt-5.4 and gpt-5.2-codex
  • Token costs averaged roughly 90 USD per day
  • Total usage had reached 700–800 million tokens

At that level of usage, problems don't stay hidden for long.
You don't need to listen to the hype. Run it for a few days, and you'll know whether it works well.

The biggest problem with memory-lancedb wasn't that it couldn't run, but that it fell far short of the hype in real-world use.
Manual memory_store and memory_recall calls worked, and the feature set looked complete on the surface. But automatic recall was where things went wrong: irrelevant recall, derailed context, and piles of information that had no business appearing in the current conversation.

At first, I suspected the embeddings, the vector dimensions, or a broken index. After investigating, the problem turned out to be fairly straightforward:

  1. memory-lancedb is more of a lightweight plugin, with very simple recall logic
  2. I had loaded quite a few memory documents directly into the vector database, making the candidate pool noisy
  3. It could retrieve information, but often retrieved things that “shouldn't appear right now”

The clearest example was this:
You ask about a current issue, and it digs up old logs, fragments of contact information, forum accounts, or even technical details from past troubleshooting sessions. The system looks busy, but it's actually polluting the context.

I did try to fix it. I went through a round of troubleshooting and adjustments:

  • Turned off memorySearch
  • Switched the memory slot to memory-lancedb
  • Reduced document chunk sizes
  • Reloaded the documents into LanceDB
  • Adjusted the number of entries injected by autoRecall

These changes weren't entirely useless; things did improve a little in the short term.
But the underlying problem never went away. The longer I used it, the more certain I became: I was forcing a plugin suited to “manual long-term memory” to do the job of the official primary memory system.

Then I looked at OpenClaw's local documentation and built-in implementation, and the direction became clear. The official approach had been spelled out all along:

  • memory-core is the default memory plugin
  • The primary tools are memory_search and memory_get
  • The default backend is builtin
  • QMD is an experimental sidecar, not a required default

One line in the official documentation stood out to me: unless you explicitly want to run QMD, stick with builtin.

That's the approach I followed when switching back:

  • Restored agents.defaults.memorySearch.enabled = true
  • Changed plugins.slots.memory to memory-core
  • Disabled memory-lancedb
  • Restarted the gateway
  • Checked memory status again

After switching back, the system immediately ran much more smoothly.
openclaw status showed:

  • plugin memory-core
  • FTS ready
  • vector ready
  • memory files were now managed by the plugin

Further retrieval tests weren't exactly brilliant, but at least the system was behaving normally again.
Queries about user preferences hit entities/wood.md and opinions.md first; queries about contacts also prioritized entity and user pages. There was still some noise farther down the results, but it was no longer the previous situation where “everything surfaced at once.”

After all this tinkering, my conclusions are simple.

1. memory-lancedb works as a lightweight plugin, not as a replacement for the official primary memory pipeline

If you just want to manually store some preferences, facts, and decisions, it can work.
But if you expect it to handle automatic recall over the long term while dumping lots of documents into it, things will sooner or later get out of hand.

2. The official defaults aren't as weak as I thought

I started with a bias: builtin seemed underpowered, while QMD looked like the “full version.” Actual use showed me that wasn't the case.
memory-core + memorySearch + builtin is already a complete setup, and it's noticeably more stable and easier to maintain.

3. It's not just the engine holding things back; the memory files matter too

After switching back to the official setup, I kept reviewing the index results and found another practical problem: the same facts were repeated across multiple files.

For example, a contact's information might appear in all of these:

  • daily logs
  • world.md
  • entities/*.md
  • users/*.md

In that situation, even a stable official memory system can end up recalling duplicate facts together.

So I did another small cleanup. I didn't overhaul the structure; I just removed the most obvious duplicates and content that shouldn't be kept long-term, such as:

  • Consolidating contact details into entity / user pages
  • Removing plaintext usernames and passwords from public long-term memory
  • De-emphasizing long-term facts in daily logs once they had been consolidated elsewhere

This step turned out to be crucial. No matter how strong the memory system is, messy source files still produce messy results.

4. Problems like these only fully reveal themselves after several days of heavy use

This experience reinforced one thing for me: you can't judge a memory system from demo videos, much less just by checking whether it runs.

In small-scale tests, plenty of solutions look usable.
But when you actually use them for several days straight, spend tens of dollars on tokens each day, and reach hundreds of millions of tokens in total, many of the “advantages” amplified in videos quickly fade, while the real problems become increasingly obvious.

That's exactly what happened with memory-lancedb for me.
As a Demo, it looked convincing.
Under heavy daily use, the noise and context pollution became increasingly frustrating.

5. My choice now is simple: return to the official default pipeline first

At least for now, I won't keep tinkering with QMD, and I won't bring memory-lancedb back as the primary system.

My principles going forward are:

  • Primary memory: memory-core
  • backend: builtin
  • Keep organizing memory files into layers, removing duplicates, and consolidating information
  • Use daily logs mainly for that day's events, and consolidate long-term facts into dedicated files wherever possible

It's not flashy, but it's stable.

One final thought.

If you're also tinkering with OpenClaw's memory and have started wondering, “Do I need to switch plugins, backends, or embeddings to fix this?”, I'd suggest pausing and checking two things first:

  1. Have you already drifted away from the official default pipeline?
  2. Have your memory files themselves become a mess?

Often, the problem isn't how powerful the model is. It's that you've gone down the wrong path.

Related posts

Comments 0