Production Knowledge System
Architecture and in-progress build of a multi-source knowledge system: ingestion normalization, memory consolidation, and semantic queryability.
Overview
Building a production-grade knowledge system forces you to confront problems that most RAG tutorials skip entirely: how do you handle knowledge that changes over time? What happens when a new document contradicts an older one? How do you ingest from heterogeneous sources — documents, notes, conversations, code — without writing a custom parser for each? The architecture centers on three problems: ingestion normalization (a common schema regardless of source format), memory consolidation (detecting and resolving when new knowledge supersedes older entries), and query routing (combining semantic similarity with structured retrieval). The ingestion pipeline processes raw content through a normalization layer before embedding, so vector search operates over clean, standardized knowledge units rather than raw document chunks. Content-addressed storage using hashes prevents duplicate re-embedding and enables efficient incremental updates. The architecture is fully designed and documented; the core ingestion pipeline and embedding layer are built. Memory consolidation — the hardest piece — is under active development. The design work came directly from shipping a production knowledge system in a commercial context, and the personal build is the space to solve the remaining problems properly, without production pressure narrowing the solution space. This project makes a specific claim: senior agentic engineering is about knowing which problems are solved and which require real design work. The solved problems — chunking, embedding, vector retrieval — are handled by existing tooling. The unsolved problems — supersession detection, cross-source consolidation, knowledge freshness — are where the actual engineering happens. This build makes that distinction visible.
Technical Stack
AI/ML
- ▸OpenAI Embeddings
- ▸ChromaDB
- ▸Sentence Transformers
- ▸LangChain
- ▸Semantic Search
Backend
- ▸Python
- ▸FastAPI
- ▸SQLAlchemy
- ▸AsyncIO
- ▸PostgreSQL
Infrastructure
- ▸Docker
- ▸Celery
- ▸Redis
- ▸Mac Mini (self-hosted)
Key Features
Source-agnostic ingestion pipeline normalizing documents, notes, conversations, and code to a common schema before embedding
Content-addressed storage using SHA-256 hashes to skip re-embedding of exact duplicates on incremental updates
Memory consolidation engine that detects when incoming knowledge supersedes or contradicts existing entries
Hybrid semantic and keyword query routing for high-recall retrieval with precision re-ranking
Incremental indexing without full re-embedding — only changed units trigger new embedding calls
Async processing pipeline decoupling ingestion throughput from embedding latency
Cross-source contradiction detection as a first-class design concern, not an edge-case afterthought
Code Examples
Technical Challenges
Designing memory consolidation that reliably detects supersession without false positives — especially when knowledge is implicit and no document references another
Building a chunking strategy that preserves enough context for accurate embedding while keeping units small enough for precise retrieval
Handling heterogeneous source formats through a single normalized schema without losing structure meaningful to downstream retrieval
Incremental updates in vector stores that do not natively support partial re-indexing at the unit level