Local-first semantic search over your files. Index a directory and query it with natural language — all embeddings are computed on your machine using native CoreML or ONNX Runtime, with no external API calls.
Install (choose one):
# Via npm (prebuilt binary, no Rust toolchain needed)
npm install -g @mathew-cf/rag-cli
# Via Cargo (builds from source)
cargo install rag-cli
# From the repo
cargo install --path .Use:
# Index a directory
rag index ./my-project
# Search it
rag search "how does authentication work"
# Pre-cache the embedding model (optional — makes first use faster)
rag downloadPrebuilt binaries are published for macOS ARM64, macOS x86_64, Linux x86_64, Linux ARM64, and Windows x86_64. On any other platform, fall back to cargo install rag-cli.
With a <path>, recursively discovers text files under it, chunks them,
computes embeddings, and writes the index to disk. With no path, reads a
config file and builds every index it declares.
| Flag | Default | Description |
|---|---|---|
-o, --output <dir> |
.rag |
Where to store the index |
-m, --model <id> |
sentence-transformers/all-MiniLM-L6-v2 |
HuggingFace model ID |
--chunk-size <n> |
512 |
Chunk size in characters |
--chunk-overlap <n> |
64 |
Overlap between consecutive chunks |
--ext <list> |
— | Extra file extensions to index, beyond the built-in allowlist (comma-separated or repeated), e.g. --ext mdx,rst |
--exclude <list> |
— | Directory/file specs to skip (comma-separated or repeated) |
--include <list> |
— | Normally-skipped directories to index anyway, e.g. --include dist |
--no-ignore |
off | Include files ignored by .gitignore, .ignore, and Git excludes |
--hidden |
off | Include hidden source directories |
-c, --config <file> |
rag.toml |
Config file to build from when no path is given |
--only <list> |
— | In config mode, build only these named indexes (skip slow ones you didn't change) |
An --exclude/--include spec without a / matches any path component by
name (e.g. changelog skips every changelog/ directory). A spec containing a
/ is treated as a relative-path prefix (e.g. src/content/changelog skips
only that one). --include re-enables directories that are skipped by default
(node_modules, dist, build, vendor, target, hidden dirs, …).
File discovery also respects .gitignore, .ignore, .rgignore, Git excludes, and global
ignore rules by default. --no-ignore disables those ignore files; --hidden
includes hidden directories. These two switches work for a single path or all
selected rag.toml entries. Per-entry no_ignore and hidden values override
the global config defaults; CLI switches can turn either behavior on for a run.
Re-running rag index on the same directory performs incremental indexing —
only changed or new files are re-embedded. File changes are detected using
blake3 content hashes. If you change
the model or chunk settings — or when a rag-cli upgrade changes the embedding
backend/precision or on-disk format — the entire index is rebuilt automatically.
A no-change run checks the small metadata file and does not load or rewrite the
full index. Changed files reuse embeddings for chunk text already present in the
index. Search reads matching chunks from the original files, so those files must
remain available and unchanged until the index is rebuilt.
To build a whole set of indexes with one command — instead of a shell script
that calls rag index once per directory — declare them in a rag.toml and run
rag index with no path. Global keys at the top are defaults; each [[index]]
may override them. Paths and output dirs are resolved relative to the config
file, including an explicit output value. Relative source paths remain
relative in meta.json (and therefore in JSON search results), so committed
indexes do not contain machine-specific absolute paths.
# Global defaults (all optional)
model = "sentence-transformers/all-MiniLM-L6-v2"
chunk_size = 512
chunk_overlap = 64
no_ignore = false # optional; default respects ignore files
hidden = false # optional; default skips hidden directories
[search]
hybrid = true # optional default for rag search
semantic_weight = 1.0 # relative RRF contribution
keyword_weight = 2.0
[[index]]
name = "docs" # output defaults to .rag/<name>
path = "docs/src/content" # relative to this config file
extensions = ["mdx"] # extra extensions beyond the built-in allowlist
exclude = ["changelog"] # skip these dirs/files
# include = ["dist"] # re-include normally-skipped dirs
# no_ignore = true # override the global ignore-file default
# hidden = true # override the global hidden-dir default
# output = ".rag/docs" # override the default output dir
[[index]]
name = "reference"
path = "reference/md"With this config in the current directory:
| Command | Default scope | One-off selection |
|---|---|---|
rag index |
Build every [[index]] |
--only docs |
rag search "query" |
Search every configured index | --only docs or repeat --index |
rag search "query" --hybrid |
Semantic and live keyword ranking over every configured index | --only docs |
rag keyword -e term |
Scan every configured source directory live | --only docs or give a path |
rag info |
Show every configured index | --only docs |
rag search and rag info require every selected index to have been built.
After rag index --only docs, use --only docs for those commands until the
other indexes are built. rag keyword scans source files and works before indexing.
rag index looks for rag.toml then .rag.toml in the current directory, or
use --config <file>. To rebuild just some of the declared indexes, use
--only: rag index --only docs.
The optional [search] table sets defaults for rag search without changing
the stored indexes. CLI search options override these defaults.
Indexes can live in the root repository while their source directories are Git submodules. This keeps generated data out of the submodules and lets each corpus update independently:
[[index]]
name = "vendor-docs"
path = "vendor/docs" # submodule
output = ".rag/vendor-docs" # root-repository indexThe same config can be searched as a federation with one query embedding:
rag search "cache behavior" # auto-discovers rag.toml
rag search "cache behavior" --only docs,reference
rag search "cache behavior" --config path/to/rag.tomlEmbeds your query and returns the most similar chunks by cosine similarity.
Add --hybrid to combine semantic ranking with live keyword matches from the
indexed source files. Hybrid result scores are reciprocal-rank-fusion scores,
not cosine similarities. Without extra options, the two rankings have equal
weight and keywords are extracted from the query. Hybrid search considers
files recorded in the index; run rag index after adding files or changing
file-selection rules so they join the hybrid corpus.
| Flag | Default | Description |
|---|---|---|
-i, --index <dir> |
auto | Index directory to search; repeat to federate several indexes |
-c, --config <file> |
auto | Search indexes declared by a config file |
--only <list> |
— | With an explicit or discovered config, search only these named indexes |
-k, --top-k <n> |
5 |
Number of results across all indexes |
-m, --model <id> |
(from index) | Override embedding model; must match the indexes |
--group-by-source |
off | Return at most one result from each source file |
--hybrid |
off or rag.toml |
Combine semantic and live keyword ranking across selected indexes |
--no-hybrid |
off | Use semantic ranking only, overriding rag.toml |
--semantic-weight <n> |
1 or rag.toml |
Relative contribution of semantic ranking; implies --hybrid |
--keyword-weight <n> |
1 or rag.toml |
Relative contribution of keyword ranking; implies --hybrid |
--keyword <text> |
(query terms) | Literal keyword or phrase; repeat to match any; implies --hybrid |
--full |
off | Show full chunk text instead of truncated preview |
--json |
off | Output compact JSON (for piping to LLMs or other tools) |
With neither --index nor --config, search discovers rag.toml or
.rag.toml in the current directory and searches its indexes. If neither file
exists, it falls back to .rag. Explicit options always take precedence.
Use --group-by-source when broad source coverage is more useful than several
high-scoring chunks from one file:
rag search "cache behavior" --top-k 5 --group-by-source
rag search "cache behavior" --hybrid --top-k 5
rag search "why do cache entries expire" --keyword TTL --keyword eviction --keyword-weight 2Weights must be positive finite numbers; their ratio determines the relative
influence of each ranking. --keyword replaces the automatically extracted
terms, so a phrase such as --keyword "cache miss" matches those words together.
The command searches every index in rag.toml unless --only narrows the set.
[search] can set hybrid, semantic_weight, and keyword_weight for routine
queries. --hybrid, --keyword, or either weight flag enables hybrid search
for one command; --no-hybrid switches back to semantic search. An explicit
--index uses the named index paths without loading rag.toml defaults.
With --group-by-source, rag-cli considers a bounded window of up to 10 times --top-k
semantic candidates in each index and keeps each source's best occurrence from
that window. Federated results are then grouped globally. Sources under absolute
root_dir values overlap by canonical root and source path, so only the best
result is retained. A relative root_dir (such as .) remains scoped to its
index for source grouping, so unrelated same-named files from different indexes
stay separate. Per-index source candidates are also overfetched
(up to the same bound) so the global merge can refill slots removed by overlap.
If the bounded windows contain fewer than k distinct canonical sources, fewer
than k results are returned. Identical chunk text occurring in several files
can represent each file. Without this flag, search behavior is unchanged: it
ranks unique chunk bodies per index and resolves each result to its first
occurrence, without cross-index source grouping.
Repeat --index to search existing indexes without merging or rebuilding them:
rag search "cache behavior" \
-i .rag/product-docs \
-i .rag/api-reference \
-i .rag/examplesA config may build independent indexes with different models, but searching
those indexes together is not possible: federated indexes must use the same
model and embedding dimensions. Use --only to select a compatible subset.
Compatible indexes are loaded and searched sequentially, keeping peak memory
near the largest index
rather than the sum of all indexes. Each index contributes its local top-k, then
rag-cli computes the global top-k. Exact-text deduplication remains per-index;
identical text stored in different indexes may appear more than once.
Federated JSON results retain source, score, byte_offset, and text, and
also include index, root_dir, and (for config entries) index_name. These
fields let callers resolve a relative source path against the correct corpus.
When an index was built from a relative CLI or rag.toml path, root_dir
preserves that relative path rather than exposing the builder's absolute path.
Search live files without loading an embedding model. Patterns passed with -e
are ORed and treated as regular expressions by default. Add -F for literal
strings, -i to ignore case, and -l to print only matching file paths.
--glob accepts ripgrep-style include and exclude patterns and overrides
ignore-file rules as in ripgrep. By default the walker respects .gitignore,
.ignore, .rgignore, Git excludes, and hidden-file rules; --no-ignore and --hidden
relax those rules separately. With no path, rag keyword
searches every source directory in rag.toml or .rag.toml; use --only to
select names or --config to choose another file. If there is no config, it
searches the current directory. An explicit path searches just that path.
rag keyword -e retry -e backoff docs --glob '*.md' -l
rag keyword -F -i -e 'error.code' sessions --glob '*.jsonl' -l
rag keyword -e cache --only docs,reference -l # source dirs from rag.tomlConfig-based keyword search scans live files under the selected source paths
using each entry's extensions, include/exclude, and ignore settings. It does
not require the indexes to be built. --glob narrows that configured corpus;
it can override ignore files but not an entry's explicit exclude rule. An
explicit path scans all file types unless --glob narrows them. Results
identify files by path, including their source directory.
Exit status is 0 when files match, 1 when none match, and 2 on an error.
Prints index metadata: format, model, chunk and unique-text counts, duplicate count, source file count, index size, etc.
| Flag | Default | Description |
|---|---|---|
-i, --index <dir> |
auto | One index directory to inspect |
-c, --config <file> |
auto | Show indexes declared by a config file |
--only <list> |
— | With an explicit or discovered config, show only these indexes |
With no options, info discovers rag.toml/.rag.toml and reports every
configured index, falling back to .rag when no config exists.
Apple Silicon builds use a native FP16 CoreML model with pooling and normalization fused into the compiled graph. Other platforms use ONNX Runtime on CPU with architecture-tuned int8 weights. The backend is selected at build time; no GPU toolkit or runtime flags are required.
On Apple Silicon, native CoreML substantially reduces indexing time and memory compared with the int8 ONNX path. The downloaded CoreML artifact is pinned to an immutable Hugging Face revision. The ONNX Runtime library remains statically linked for non-Apple builds, so there is nothing to install separately.
Persisted vectors use F16 independently of int8 model inference. Exact duplicate chunk text is embedded once. The index stores one vector and content hash per unique chunk, plus compact source and byte-range records for each occurrence. Search decodes F16 values for exact cosine scoring, then reads the winning chunks from their source files and checks that those files have not changed.
rag-cli downloads model weights directly from HuggingFace over HTTPS on first
use, then caches them locally in the standard HuggingFace Hub layout
(~/.cache/huggingface/hub). We do this instead of using the hf-hub crate because hf-hub uses a bundled
certificate store and does not respect the system root CA certificates. That
makes it fail in environments with custom CA roots (corporate proxies, internal
TLS inspection, etc.). rag-cli uses native-tls, which delegates to the OS certificate store,
so it works in those environments without extra configuration.
The native CoreML backend on Apple Silicon currently supports the default MiniLM model only. Non-Apple ONNX builds retain model overrides when the requested repository publishes the expected architecture-specific ONNX artifact.
Run rag download to prefetch models. With no --model or --config, it
auto-discovers rag.toml/.rag.toml, deduplicates the models used by its
entries, and downloads each one. Use --only <names> to prefetch a subset,
--config <file> to select another config, or --model <id> for one explicit
model. If no config exists, the built-in default model is downloaded.
You can control the cache location:
# Via flag
rag --cache-dir /path/to/cache index ./docs
# Via environment variable
export RAG_CACHE_DIR=/path/to/cache
rag index ./docs
# Or use HF_HOME (standard HuggingFace convention)
export HF_HOME=/path/to/hf
rag index ./docsrag-cli is aimed at prose, docs, and config — not source code. It indexes:
- Docs / prose:
.md,.txt,.tex,.org,.rst - Data:
.csv,.tsv,.log - Config:
.toml,.conf,.cfg,.ini,.env,.tf,.hcl,.nix - Markup:
.xml,.html - Schemas:
.sql,.proto,.graphql - Build:
Dockerfile,Makefile,.cmake
Files like Makefile, Dockerfile, LICENSE, README, and .gitignore are
recognized by name.
Programming language sources (.rs, .py, .ts, .go, .js, shell scripts,
etc.), .json, .yaml/.yml, and .css/.scss are intentionally not
indexed — they tend to drown out useful matches with boilerplate. Index your
code with a code-aware tool instead.
To index an extension that isn't in the allowlist (for example .mdx), add it
with --ext (or an entry's extensions in rag.toml): rag index ./docs --ext mdx.
Hidden directories, node_modules, target, __pycache__, vendor, dist,
and build are skipped automatically. Skip more with --exclude, or force a
skipped directory back in with --include.
- Discover text files recursively, skipping binary and vendored content
- Chunk each file into overlapping segments (~512 chars), breaking at paragraph or line boundaries when possible
- Deduplicate and embed each unique chunk body once, in batches of 128,
using
all-MiniLM-L6-v2with mean pooling and L2 normalization - Store one F16 vector and content hash per unique chunk, plus compact
source/byte-range occurrence records, in
.rag/index.bin; inference remains int8 with f32 output, so F16 applies only to persisted vectors - Search by embedding the query with the same model and ranking unique chunk bodies by exact cosine similarity, then reading matches from source files
Apache-2.0