I had the retrieval working. Thousands of support articles, a vector store, semantic search — the kind of thing you can demo in twenty minutes and feel good about.
The problem surfaced later, quietly. An article that had been accurate for eighteen months became wrong overnight when a product feature changed. The retrieval system kept serving it. It was relevant. It was just wrong.
That gap — between "retrieval works" and "the system actually knows things reliably" — is where most production knowledge systems fail in production. Here's what I've learned building on both sides of it.
Retrieval Is the Easy Half (And That's Not a Criticism)
RAG is genuinely useful. If your team has thousands of internal documents and your AI can't surface anything from them, RAG solves that problem. I'm not here to argue that vector search is oversold. It works.
But RAG as shipped is a retrieval system, not a knowledge system. There's a difference.
A retrieval system answers: given this query, what documents are similar?
A knowledge system answers: given this question, what is true — and how confident should I be in that?
The distinction matters in production. When I was building the retrieval layer over a large support knowledge base (de-identified), the hard part wasn't the embedding or the similarity search. Those are solved problems with known tradeoffs. The hard part was that the retrieval layer had no concept of whether its sources were still accurate.
There are three things RAG doesn't give you out of the box:
Staleness awareness. A document ranked highly by similarity has no signal about whether it reflects current truth. If the underlying system changed and the article wasn't updated, retrieval can't detect that. It surfaces what's relevant. Relevance and accuracy are not the same thing.
Contradiction detection. When two articles give conflicting guidance, retrieval will return both and let the model sort it out. Sometimes the model handles it well. Sometimes it doesn't, and the user gets a confident, wrong answer.
Accumulation. Every query starts from zero. Retrieval doesn't learn that a certain type of question consistently returns outdated results. It doesn't update its confidence. It retrieves.
This is why I think of retrieval as the easy half — not because it's simple (building a production-quality RAG system with real performance characteristics takes real engineering), but because it's a solved problem with well-understood tools. The harder problem is what has to sit underneath it.
For a deeper look at what memory-aware systems need to track across interactions, the agent memory guide on this site covers the memory types in detail.
What Keeping Knowledge Current Actually Requires
The support platform I worked on had three kinds of content that looked identical to the retrieval layer: evergreen policy documents, procedure guides that changed with product releases, and articles written to address a specific support spike and never revisited.
The retrieval system treated all three identically. An article was an article.
What you actually need is ingestion routing — the recognition that not all content carries the same kind of truth, and different kinds of truth have different freshness requirements.
Policy documents carry high stakes when they're wrong, but they change infrequently — deliberate review cycles rather than reactive updates are appropriate here. Procedure guides are tightly coupled to product state: they go stale whenever the product changes, and their freshness is better measured by product version than by last-edited date. Surge articles — written to address a specific support spike and never revisited — are often accurate at write time and obsolete within weeks; they need explicit expiry or decay, not indefinite storage.
If your ingestion pipeline doesn't distinguish between these, your retrieval layer will confidently surface the wrong answer at the worst possible moment — when a user is confused about something that recently changed.
Beyond routing, you need contradiction detection at write time. Before a new article enters the knowledge base, something should ask: does this conflict with existing content? This is harder than it sounds. Contradictions aren't always lexically obvious — they're often semantic, and catching them reliably means embedding the new content and comparing it against nearby neighbors before committing. The knowledge graph approach — maintaining explicit relationships between articles — helps considerably here, because you can run targeted contradiction checks on connected nodes rather than exhaustive comparisons across the full corpus.
A minimal write-time check looks like this:
The check doesn't need to be exhaustive — neighboring nodes in the knowledge graph narrow the comparison space considerably.
The third piece is temporal reasoning. Knowledge isn't just true or false; it's true as of some point in time. The data model needs to distinguish between "this has always been the policy" and "this was accurate in Q3 and may have changed." That distinction has to be explicit in the data model, not implicit in the article text. If it's implicit, no downstream system can use it.
The data model shape that makes this distinction trackable:
A consolidation worker can walk this schema, flag entries whose valid_until has passed or whose confidence has fallen below threshold, and queue them for review — without touching the retrieval surface at all.
Memory Consolidation: Where Implementations Fail Silently
This is the hardest part to get right and the easiest to defer — which is why most systems never build it.
Memory consolidation is the process by which a knowledge system synthesizes what it has ingested over time into durable, structured understanding. Without it, you get retrieval that works at launch and degrades slowly. More content, more noise, more unresolved contradictions. Every query still returns something. The something is just progressively less reliable.
In a production system, consolidation involves three things:
Episodic to semantic compression. Specific articles and specific interactions eventually need to be synthesized into general principles. If many support conversations all resolved to "the user misunderstood feature X," that's a semantic fact about your product that belongs explicitly in the knowledge base — not buried across dozens of individual embeddings. The compression doesn't happen automatically. You have to build it.
Explicit conflict resolution policy. When two pieces of content contradict each other and the contradiction wasn't caught at write time, what happens? If your system has no explicit policy, the model decides — inconsistently. Explicit policy (prefer the more recent source, prefer the higher-authority source, surface both with a flag) is better than implicit behavior in every case. The failure mode isn't the model picking wrong once; it's picking wrong unpredictably, which means you can't reason about when to trust the system.
Decay and pruning. High-volume knowledge bases accumulate content that becomes low-signal over time. Without decay mechanics, old surge articles and superseded procedures sit in the vector store indefinitely, adding noise to every retrieval. Decay doesn't have to be aggressive — a freshness decay factor on similarity scores, or periodic review queues for content past a certain age, is enough. But it has to exist, because without it the knowledge base only grows in size and in noise.
The failure mode I keep seeing: teams build the ingestion and retrieval layers carefully, then skip consolidation because it isn't visible. The system performs well at launch. Six months later, quality has degraded and nobody can tell you why. No single query fails. The overall reliability just erodes.
This is the same class of problem I described in the 12-factor agent pattern: the hardest bugs in agentic systems aren't the ones that crash. They're the ones that succeed, incorrectly, silently.
What This Looks Like in Practice
The Python orchestration framework I'm building is designed specifically for this kind of pipeline work. The architecture is node-based and event-driven — each step in the ingestion and consolidation pipeline is a named, typed node with a known input and output. When something goes wrong, you can see where.
The memory layer runs on pgvector. The ingestion pipeline routes content by type, attaches freshness metadata to every write, and runs a similarity-based contradiction check before committing new content. The consolidation workers run asynchronously and feed back into the vector store without blocking retrieval. The retrieval surface is the same semantic search you'd find in any RAG system — but it's sitting on top of a data model that tracks what it knows, how confident it is, and when that knowledge was last validated.
For a look at how to architect the separation of concerns cleanly — so that background consolidation workers can update the knowledge base without disrupting the retrieval surface — the MCP triage architecture post covers that design pattern in detail.
What I've proven in production is the retrieval layer. It works. What I'm proving now is whether the memory and freshness layer underneath it makes the reliability difference I expect it to.
The pieces aren't exotic: a data model that tracks content type and freshness, an ingestion pipeline that routes and validates before writing, contradiction detection at write time, and consolidation workers that run in the background. The challenge is building all of it consistently and not deferring the hard parts until after launch — which is when deferring them hurts most.
If you're working on a similar architecture or wrestling with knowledge freshness at scale, I'd be glad to compare notes. Reach me at [email protected] or on LinkedIn.