Skip to main content
Back to Projects

Rust RAG Engine

Rust RAG backend: two-stage hybrid retrieval over pgvector, streaming WebSocket chat via Actix actors, local Ollama inference.

RustRAGActixpgvectorOllamaWebSocketHybrid Retrieval

Overview

Most RAG demos stop at pure vector search. This project goes further: a two-stage hybrid retrieval pipeline where semantic pgvector cosine search narrows the candidate set, then a keyword re-rank pass scores within that set and fuses the two signals. The result handles exact-term queries that trip up pure vector search without sacrificing recall on paraphrased ones. The whole thing is written in Rust — not because Rust is fashionable, but because the architecture calls for it: an Actix actor model per WebSocket session, a bounded-concurrency Ollama pipeline, and a pgvector integration that needs zero-overhead async. The ingestion side fetches articles from a HelpScout-style REST API, converts HTML to clean Markdown using html2md and custom scraper rules, stores canonical articles in PostgreSQL via Diesel ORM, then fans out to two parallel pipelines: a Python Flask microservice (all-MiniLM-L6-v2) generates embeddings stored in pgvector, while a bounded-concurrency Ollama pipeline generates per-article keywords and summaries. The metadata pipeline uses a Tokio Semaphore and buffer_unordered to cap in-flight local LLM calls, with structured success/partial/failure accounting and failed-ID persistence for retries. Streaming chat runs through an Actix actor model. Each WebSocket connection spawns a ChatSession actor; a central ChatServer actor routes messages and manages session state. Ollama's streaming API is consumed token-by-token and forwarded over the connection in real time — the client sees words appear as they're generated, grounded in articles retrieved by the hybrid pipeline. A round-robin Ollama load balancer backed by a Rayon thread pool is included in the codebase for multi-instance inference distribution, currently shelved pending hardware but fully wired and ready. This is a backend-only service — the unfinished frontend scaffold was removed intentionally. The value is in the infrastructure: a production-grade Rust RAG pipeline with real hybrid retrieval, streaming chat, and a local-inference-first design. Clean cargo build, zero clippy warnings, 19 tests, CI configured. Public on GitHub under bredmond1019.

Technical Stack

AI / Retrieval

  • ▸pgvector (cosine similarity)
  • ▸Sentence Transformers (all-MiniLM-L6-v2)
  • ▸Ollama (llama3.1)
  • ▸Hybrid two-stage retrieval
  • ▸Keyword re-ranking

Backend

  • ▸Rust (2021 edition)
  • ▸Actix-web 4
  • ▸Actix actors
  • ▸Tokio
  • ▸Diesel 2

Data

  • ▸PostgreSQL
  • ▸pgvector
  • ▸Diesel migrations
  • ▸Python Flask (embedding microservice)

Infrastructure

  • ▸GitHub Actions CI
  • ▸Docker (Postgres + pgvector)

Key Features

✓

Two-stage hybrid retrieval: semantic pgvector cosine search narrows candidates, keyword re-rank fuses signals for score fusion

✓

Actix actor model for WebSocket chat — per-session ChatSession actors managed by a central ChatServer

✓

Token-by-token streaming from Ollama forwarded in real time over the WebSocket connection

✓

Bounded-concurrency metadata pipeline using Tokio Semaphore and buffer_unordered to cap in-flight Ollama calls

✓

Paginated article sync from a HelpScout-style REST API with 30-second timeout and full HTML-to-Markdown conversion

✓

Python Flask embedding microservice running sentence-transformers all-MiniLM-L6-v2 at localhost:8080

✓

Async job queue for background ingestion with per-job status tracking

✓

Failed-article persistence and re-embed endpoint for recovery from partial pipeline failures

✓

Round-robin Ollama load balancer (Rayon thread pool) for multi-instance inference distribution — included, shelved pending hardware

Code Examples

Technical Challenges

▪

Designing the two-stage retrieval fusion so keyword re-ranking improves exact-term precision without regressing semantic recall

▪

Building the Actix actor model so per-session chat history lives inside the actor rather than in shared external state

▪

Implementing bounded concurrency in the metadata pipeline — Semaphore + buffer_unordered — so bulk Ollama calls don't saturate local hardware

▪

Vendoring and patching ollama-rs 0.2.0: upstream 0.2.x→0.3.x changed the history API to require external Arc<Mutex<MessagesHistory>>, which is incompatible with the per-session actor ownership model

▪

Writing custom scraper rules for HTML-to-Markdown conversion to preserve structure in long, nested support articles

▪

Keeping the Python embedding microservice interface stable so embedding dimension changes can be handled through a migration rather than code churn

Project Outcomes

Clean — zero clippy warnings, cargo fmt pass
Build
19 passing (api_client, article, HTML conversion)
Tests
Backend-only, self-contained — fresh clone builds
Scope
Public on GitHub (bredmond1019)
Code