A fully self-hosted, GPU-split Retrieval-Augmented Generation pipeline — local LLM, local embeddings, hybrid vector search, and PDF→Markdown preprocessing — wired together and documented end to end.
This repository is the design documentation for that pipeline: an annotated architecture, a colour-coded pipeline graph, an exhaustively enumerated configuration reference, the design decisions (including the trade-offs and dead-ends), an operations runbook, and the (genericized) scripts.
Everything runs on-premises — no data leaves the host, no per-token API costs. Host names, IPs, and hardware identifiers are intentionally omitted; GPUs are referred to as GPU A (12 GB) and GPU B (8 GB).
- Pipeline at a glance
- Highlights
- Verified working
- Components
- The two data paths
- Quickstart
- Documentation
- Known limitations
- License
Source: docs/pipeline.dot · scalable: assets/pipeline.svg
- Dual-GPU split — the generation model and the embedding model live on separate cards, so ingestion (embedding-heavy) and querying (generation-heavy) never contend for VRAM.
- Hybrid retrieval — dense vectors (semantic) plus sparse BM25 (exact tokens: codes, IDs, names, acronyms), fused, on a real Qdrant server.
- Agentic RAG that actually grounds — the LLM calls a retrieval tool; a chat-template override gives an uncensored ChatML-trained model reliable tool-calling that it otherwise lacks.
- Quality-first preprocessing — PDFs are converted to clean, structured
Markdown with
markerbefore chunking, which beats raw-PDF text extraction. - Fully local & API-compatible — OpenAI/Anthropic-shaped HTTP throughout; PrivateGPT runs standalone (embedded SQLite); no Postgres/Redis/RabbitMQ.
- Operable — idempotent ingestion (sha256 manifests), systemd-managed services, a one-command rollback path, and a built-in evaluation harness to tune retrieval objectively before committing a large corpus.
Each stage was tested end to end:
- ✅ Ingest → retrieve → grounded answer through the agentic API, with citations (the LLM correctly answers only from ingested content).
- ✅ Hybrid retrieval — an exact-code query is found via the BM25/sparse side and a paraphrased query is found via the dense side.
- ✅ Embeddings — 1024-dim, L2-normalized, last-token pooled (Qwen3-Embedding).
- ✅ Tool-calling — the LLM emits well-formed tool calls (via the template fix).
- ✅ Idempotent ingestion — re-running skips unchanged files; changed files are re-ingested.
- ✅ PDF→Markdown —
markerproduces structured Markdown from source PDFs. - ✅ Eval harness — reports retrieval hit-rate and answer accuracy.
A query is a single API call; the orchestrator handles retrieval + grounding:
curl -s localhost:8001/v1/messages -H 'Content-Type: application/json' -d '{
"model":"dolphin-8b","max_tokens":300,
"messages":[{"role":"user","content":"<your question>"}],
"tools":[{"name":"semantic_search","type":"semantic_search_v1",
"context":[{"type":"ingested_artifact",
"context_filter":{"collection":"handbook"}}],
"inputSchema":{"type":"object",
"properties":{"query":{"type":"string"}},
"required":["query"]}}]
}'
# -> a grounded answer drawn from the "handbook" collection, with sources.| Layer | Software | Model | Where | Port |
|---|---|---|---|---|
| Generation LLM | llama.cpp llama-server |
Dolphin-2.9.4-llama3.1-8b Q8_0 | GPU A (12 GB) | 8081 |
| Embeddings | llama.cpp llama-server |
Qwen3-Embedding-0.6B f16 (1024-dim) | GPU B (8 GB) | 8082 |
| RAG orchestrator | PrivateGPT (private-gpt, uv tool) |
— (middleware) | CPU/RAM | 8001 |
| Vector store | Qdrant server (podman) + fastembed BM25 | — | disk/RAM | 6333 |
| Doc/index store | SQLite (embedded) | — | disk | — |
| PDF preprocessing | marker (marker-pdf, dedicated venv) |
Surya models | CPU | — |
1 — Ingestion (offline). Drop files in documents_raw/<collection>/ →
preprocess.py converts PDFs to clean Markdown with marker → bulk-ingest.py
sends each file to PrivateGPT, which chunks it, embeds the chunks dense
(Qwen3-Embedding) and sparse (fastembed BM25), and stores both in Qdrant.
One collection per top-level subfolder.
2 — Query (online, agentic). A question goes to PrivateGPT → it prompts the
Dolphin LLM with a tool spec → the LLM calls semantic_search → PrivateGPT
runs a hybrid retrieval against Qdrant → the top-k chunks are returned to the
LLM → it produces a grounded answer with citations.
Full runbook: docs/operations.md. The short version:
# ingest
cp mydocs/*.pdf ~/pgpt/documents_raw/handbook/
~/marker-venv/bin/python ~/pgpt/preprocess.py # PDF → Markdown
python3 ~/pgpt/bulk-ingest.py # chunk + embed + store
# ask
xdg-open http://localhost:8001/ui # or POST /v1/messages (above)
# measure quality before scaling up
python3 ~/pgpt/eval.py ~/pgpt/evalset.jsonl| Doc | What's in it |
|---|---|
| docs/architecture.md | Components, data flow, GPU/VRAM allocation, request lifecycle |
| docs/configuration-reference.md | Every env var, CLI flag, port, path, and model parameter |
| docs/design-decisions.md | Why each choice was made — the trade-offs and the dead-ends |
| docs/operations.md | Install, start/stop, rollback, ingestion workflow, troubleshooting |
| scripts/ | Genericized serve wrappers, preprocess/ingest/eval drivers, systemd units |
- No reranker — PrivateGPT v2 has no cross-encoder rerank stage; precision is
managed with hybrid retrieval and a tuned
top_k. (details) - Small-model tool-calling is chatty — the 8B occasionally issues several retrieval calls before answering; correct, but a larger model is crisper.
- Hybrid requires a Qdrant server — the embedded Qdrant client can't do BM25 hybrid; a containerized server is used instead. (details)
Apache 2.0. See LICENSE.
Proudly Made in Nebraska. Go Big Red! 🌽 https://xkcd.com/2347/
