Rust RAG Engine
Rust RAG backend: two-stage hybrid retrieval over pgvector, streaming WebSocket chat via Actix actors, local Ollama inference.
Overview
Most RAG demos stop at pure vector search. This project goes further: a two-stage hybrid retrieval pipeline where semantic pgvector cosine search narrows the candidate set, then a keyword re-rank pass scores within that set and fuses the two signals. The result handles exact-term queries that trip up pure vector search without sacrificing recall on paraphrased ones. The whole thing is written in Rust — not because Rust is fashionable, but because the architecture calls for it: an Actix actor model per WebSocket session, a bounded-concurrency Ollama pipeline, and a pgvector integration that needs zero-overhead async. The ingestion side fetches articles from a HelpScout-style REST API, converts HTML to clean Markdown using html2md and custom scraper rules, stores canonical articles in PostgreSQL via Diesel ORM, then fans out to two parallel pipelines: a Python Flask microservice (all-MiniLM-L6-v2) generates embeddings stored in pgvector, while a bounded-concurrency Ollama pipeline generates per-article keywords and summaries. The metadata pipeline uses a Tokio Semaphore and buffer_unordered to cap in-flight local LLM calls, with structured success/partial/failure accounting and failed-ID persistence for retries. Streaming chat runs through an Actix actor model. Each WebSocket connection spawns a ChatSession actor; a central ChatServer actor routes messages and manages session state. Ollama's streaming API is consumed token-by-token and forwarded over the connection in real time — the client sees words appear as they're generated, grounded in articles retrieved by the hybrid pipeline. A round-robin Ollama load balancer backed by a Rayon thread pool is included in the codebase for multi-instance inference distribution, currently shelved pending hardware but fully wired and ready. This is a backend-only service — the unfinished frontend scaffold was removed intentionally. The value is in the infrastructure: a production-grade Rust RAG pipeline with real hybrid retrieval, streaming chat, and a local-inference-first design. Clean cargo build, zero clippy warnings, 19 tests, CI configured. Public on GitHub under bredmond1019.
Technical Stack
AI / Retrieval
- ▸pgvector (cosine similarity)
- ▸Sentence Transformers (all-MiniLM-L6-v2)
- ▸Ollama (llama3.1)
- ▸Hybrid two-stage retrieval
- ▸Keyword re-ranking
Backend
- ▸Rust (2021 edition)
- ▸Actix-web 4
- ▸Actix actors
- ▸Tokio
- ▸Diesel 2
Data
- ▸PostgreSQL
- ▸pgvector
- ▸Diesel migrations
- ▸Python Flask (embedding microservice)
Infrastructure
- ▸GitHub Actions CI
- ▸Docker (Postgres + pgvector)
Key Features
Two-stage hybrid retrieval: semantic pgvector cosine search narrows candidates, keyword re-rank fuses signals for score fusion
Actix actor model for WebSocket chat — per-session ChatSession actors managed by a central ChatServer
Token-by-token streaming from Ollama forwarded in real time over the WebSocket connection
Bounded-concurrency metadata pipeline using Tokio Semaphore and buffer_unordered to cap in-flight Ollama calls
Paginated article sync from a HelpScout-style REST API with 30-second timeout and full HTML-to-Markdown conversion
Python Flask embedding microservice running sentence-transformers all-MiniLM-L6-v2 at localhost:8080
Async job queue for background ingestion with per-job status tracking
Failed-article persistence and re-embed endpoint for recovery from partial pipeline failures
Round-robin Ollama load balancer (Rayon thread pool) for multi-instance inference distribution — included, shelved pending hardware
Code Examples
Technical Challenges
Designing the two-stage retrieval fusion so keyword re-ranking improves exact-term precision without regressing semantic recall
Building the Actix actor model so per-session chat history lives inside the actor rather than in shared external state
Implementing bounded concurrency in the metadata pipeline — Semaphore + buffer_unordered — so bulk Ollama calls don't saturate local hardware
Vendoring and patching ollama-rs 0.2.0: upstream 0.2.x→0.3.x changed the history API to require external Arc<Mutex<MessagesHistory>>, which is incompatible with the per-session actor ownership model
Writing custom scraper rules for HTML-to-Markdown conversion to preserve structure in long, nested support articles
Keeping the Python embedding microservice interface stable so embedding dimension changes can be handled through a migration rather than code churn