Sitelet https://github.com/mathew-cf/rag-cli/tree/main
Skip to content

Latest commit

 

History

62 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

rag-cli

Local-first semantic search over your files. Index a directory and query it with natural language — all embeddings are computed on your machine using native CoreML or ONNX Runtime, with no external API calls.

Quick start

Install (choose one):

# Via npm (prebuilt binary, no Rust toolchain needed)
npm install -g @mathew-cf/rag-cli

# Via Cargo (builds from source)
cargo install rag-cli

# From the repo
cargo install --path .

Use:

# Index a directory
rag index ./my-project

# Search it
rag search "how does authentication work"

# Pre-cache the embedding model (optional — makes first use faster)
rag download

Supported npm platforms

Prebuilt binaries are published for macOS ARM64, macOS x86_64, Linux x86_64, Linux ARM64, and Windows x86_64. On any other platform, fall back to cargo install rag-cli.

Commands

rag index [path]

With a <path>, recursively discovers text files under it, chunks them, computes embeddings, and writes the index to disk. With no path, reads a config file and builds every index it declares.

Flag Default Description
-o, --output <dir> .rag Where to store the index
-m, --model <id> sentence-transformers/all-MiniLM-L6-v2 HuggingFace model ID
--chunk-size <n> 512 Chunk size in characters
--chunk-overlap <n> 64 Overlap between consecutive chunks
--ext <list> — Extra file extensions to index, beyond the built-in allowlist (comma-separated or repeated), e.g. --ext mdx,rst
--exclude <list> — Directory/file specs to skip (comma-separated or repeated)
--include <list> — Normally-skipped directories to index anyway, e.g. --include dist
--no-ignore off Include files ignored by .gitignore, .ignore, and Git excludes
--hidden off Include hidden source directories
-c, --config <file> rag.toml Config file to build from when no path is given
--only <list> — In config mode, build only these named indexes (skip slow ones you didn't change)

An --exclude/--include spec without a / matches any path component by name (e.g. changelog skips every changelog/ directory). A spec containing a / is treated as a relative-path prefix (e.g. src/content/changelog skips only that one). --include re-enables directories that are skipped by default (node_modules, dist, build, vendor, target, hidden dirs, …). File discovery also respects .gitignore, .ignore, .rgignore, Git excludes, and global ignore rules by default. --no-ignore disables those ignore files; --hidden includes hidden directories. These two switches work for a single path or all selected rag.toml entries. Per-entry no_ignore and hidden values override the global config defaults; CLI switches can turn either behavior on for a run.

Re-running rag index on the same directory performs incremental indexing — only changed or new files are re-embedded. File changes are detected using blake3 content hashes. If you change the model or chunk settings — or when a rag-cli upgrade changes the embedding backend/precision or on-disk format — the entire index is rebuilt automatically. A no-change run checks the small metadata file and does not load or rewrite the full index. Changed files reuse embeddings for chunk text already present in the index. Search reads matching chunks from the original files, so those files must remain available and unchanged until the index is rebuilt.

Config file (rag.toml)

To build a whole set of indexes with one command — instead of a shell script that calls rag index once per directory — declare them in a rag.toml and run rag index with no path. Global keys at the top are defaults; each [[index]] may override them. Paths and output dirs are resolved relative to the config file, including an explicit output value. Relative source paths remain relative in meta.json (and therefore in JSON search results), so committed indexes do not contain machine-specific absolute paths.

# Global defaults (all optional)
model = "sentence-transformers/all-MiniLM-L6-v2"
chunk_size = 512
chunk_overlap = 64
no_ignore = false             # optional; default respects ignore files
hidden = false                # optional; default skips hidden directories

[search]
hybrid = true                  # optional default for rag search
semantic_weight = 1.0          # relative RRF contribution
keyword_weight = 2.0

[[index]]
name = "docs"                  # output defaults to .rag/<name>
path = "docs/src/content"      # relative to this config file
extensions = ["mdx"]           # extra extensions beyond the built-in allowlist
exclude = ["changelog"]        # skip these dirs/files
# include = ["dist"]           # re-include normally-skipped dirs
# no_ignore = true            # override the global ignore-file default
# hidden = true               # override the global hidden-dir default
# output = ".rag/docs"         # override the default output dir

[[index]]
name = "reference"
path = "reference/md"

With this config in the current directory:

Command Default scope One-off selection
rag index Build every [[index]] --only docs
rag search "query" Search every configured index --only docs or repeat --index
rag search "query" --hybrid Semantic and live keyword ranking over every configured index --only docs
rag keyword -e term Scan every configured source directory live --only docs or give a path
rag info Show every configured index --only docs

rag search and rag info require every selected index to have been built. After rag index --only docs, use --only docs for those commands until the other indexes are built. rag keyword scans source files and works before indexing.

rag index looks for rag.toml then .rag.toml in the current directory, or use --config <file>. To rebuild just some of the declared indexes, use --only: rag index --only docs. The optional [search] table sets defaults for rag search without changing the stored indexes. CLI search options override these defaults.

Indexes can live in the root repository while their source directories are Git submodules. This keeps generated data out of the submodules and lets each corpus update independently:

[[index]]
name = "vendor-docs"
path = "vendor/docs"          # submodule
output = ".rag/vendor-docs"   # root-repository index

The same config can be searched as a federation with one query embedding:

rag search "cache behavior"                  # auto-discovers rag.toml
rag search "cache behavior" --only docs,reference
rag search "cache behavior" --config path/to/rag.toml

rag search <query>

Embeds your query and returns the most similar chunks by cosine similarity. Add --hybrid to combine semantic ranking with live keyword matches from the indexed source files. Hybrid result scores are reciprocal-rank-fusion scores, not cosine similarities. Without extra options, the two rankings have equal weight and keywords are extracted from the query. Hybrid search considers files recorded in the index; run rag index after adding files or changing file-selection rules so they join the hybrid corpus.

Flag Default Description
-i, --index <dir> auto Index directory to search; repeat to federate several indexes
-c, --config <file> auto Search indexes declared by a config file
--only <list> — With an explicit or discovered config, search only these named indexes
-k, --top-k <n> 5 Number of results across all indexes
-m, --model <id> (from index) Override embedding model; must match the indexes
--group-by-source off Return at most one result from each source file
--hybrid off or rag.toml Combine semantic and live keyword ranking across selected indexes
--no-hybrid off Use semantic ranking only, overriding rag.toml
--semantic-weight <n> 1 or rag.toml Relative contribution of semantic ranking; implies --hybrid
--keyword-weight <n> 1 or rag.toml Relative contribution of keyword ranking; implies --hybrid
--keyword <text> (query terms) Literal keyword or phrase; repeat to match any; implies --hybrid
--full off Show full chunk text instead of truncated preview
--json off Output compact JSON (for piping to LLMs or other tools)

With neither --index nor --config, search discovers rag.toml or .rag.toml in the current directory and searches its indexes. If neither file exists, it falls back to .rag. Explicit options always take precedence.

Use --group-by-source when broad source coverage is more useful than several high-scoring chunks from one file:

rag search "cache behavior" --top-k 5 --group-by-source
rag search "cache behavior" --hybrid --top-k 5
rag search "why do cache entries expire" --keyword TTL --keyword eviction --keyword-weight 2

Weights must be positive finite numbers; their ratio determines the relative influence of each ranking. --keyword replaces the automatically extracted terms, so a phrase such as --keyword "cache miss" matches those words together. The command searches every index in rag.toml unless --only narrows the set. [search] can set hybrid, semantic_weight, and keyword_weight for routine queries. --hybrid, --keyword, or either weight flag enables hybrid search for one command; --no-hybrid switches back to semantic search. An explicit --index uses the named index paths without loading rag.toml defaults.

With --group-by-source, rag-cli considers a bounded window of up to 10 times --top-k semantic candidates in each index and keeps each source's best occurrence from that window. Federated results are then grouped globally. Sources under absolute root_dir values overlap by canonical root and source path, so only the best result is retained. A relative root_dir (such as .) remains scoped to its index for source grouping, so unrelated same-named files from different indexes stay separate. Per-index source candidates are also overfetched (up to the same bound) so the global merge can refill slots removed by overlap. If the bounded windows contain fewer than k distinct canonical sources, fewer than k results are returned. Identical chunk text occurring in several files can represent each file. Without this flag, search behavior is unchanged: it ranks unique chunk bodies per index and resolves each result to its first occurrence, without cross-index source grouping.

Repeat --index to search existing indexes without merging or rebuilding them:

rag search "cache behavior" \
  -i .rag/product-docs \
  -i .rag/api-reference \
  -i .rag/examples

A config may build independent indexes with different models, but searching those indexes together is not possible: federated indexes must use the same model and embedding dimensions. Use --only to select a compatible subset. Compatible indexes are loaded and searched sequentially, keeping peak memory near the largest index rather than the sum of all indexes. Each index contributes its local top-k, then rag-cli computes the global top-k. Exact-text deduplication remains per-index; identical text stored in different indexes may appear more than once.

Federated JSON results retain source, score, byte_offset, and text, and also include index, root_dir, and (for config entries) index_name. These fields let callers resolve a relative source path against the correct corpus. When an index was built from a relative CLI or rag.toml path, root_dir preserves that relative path rather than exposing the builder's absolute path.

rag keyword

Search live files without loading an embedding model. Patterns passed with -e are ORed and treated as regular expressions by default. Add -F for literal strings, -i to ignore case, and -l to print only matching file paths. --glob accepts ripgrep-style include and exclude patterns and overrides ignore-file rules as in ripgrep. By default the walker respects .gitignore, .ignore, .rgignore, Git excludes, and hidden-file rules; --no-ignore and --hidden relax those rules separately. With no path, rag keyword searches every source directory in rag.toml or .rag.toml; use --only to select names or --config to choose another file. If there is no config, it searches the current directory. An explicit path searches just that path.

rag keyword -e retry -e backoff docs --glob '*.md' -l
rag keyword -F -i -e 'error.code' sessions --glob '*.jsonl' -l
rag keyword -e cache --only docs,reference -l   # source dirs from rag.toml

Config-based keyword search scans live files under the selected source paths using each entry's extensions, include/exclude, and ignore settings. It does not require the indexes to be built. --glob narrows that configured corpus; it can override ignore files but not an entry's explicit exclude rule. An explicit path scans all file types unless --glob narrows them. Results identify files by path, including their source directory.

Exit status is 0 when files match, 1 when none match, and 2 on an error.

rag info

Prints index metadata: format, model, chunk and unique-text counts, duplicate count, source file count, index size, etc.

Flag Default Description
-i, --index <dir> auto One index directory to inspect
-c, --config <file> auto Show indexes declared by a config file
--only <list> — With an explicit or discovered config, show only these indexes

With no options, info discovers rag.toml/.rag.toml and reports every configured index, falling back to .rag when no config exists.

Hardware acceleration

Apple Silicon builds use a native FP16 CoreML model with pooling and normalization fused into the compiled graph. Other platforms use ONNX Runtime on CPU with architecture-tuned int8 weights. The backend is selected at build time; no GPU toolkit or runtime flags are required.

On Apple Silicon, native CoreML substantially reduces indexing time and memory compared with the int8 ONNX path. The downloaded CoreML artifact is pinned to an immutable Hugging Face revision. The ONNX Runtime library remains statically linked for non-Apple builds, so there is nothing to install separately.

Persisted vectors use F16 independently of int8 model inference. Exact duplicate chunk text is embedded once. The index stores one vector and content hash per unique chunk, plus compact source and byte-range records for each occurrence. Search decodes F16 values for exact cosine scoring, then reads the winning chunks from their source files and checks that those files have not changed.

Model management

rag-cli downloads model weights directly from HuggingFace over HTTPS on first use, then caches them locally in the standard HuggingFace Hub layout (~/.cache/huggingface/hub). We do this instead of using the hf-hub crate because hf-hub uses a bundled certificate store and does not respect the system root CA certificates. That makes it fail in environments with custom CA roots (corporate proxies, internal TLS inspection, etc.). rag-cli uses native-tls, which delegates to the OS certificate store, so it works in those environments without extra configuration.

The native CoreML backend on Apple Silicon currently supports the default MiniLM model only. Non-Apple ONNX builds retain model overrides when the requested repository publishes the expected architecture-specific ONNX artifact.

Run rag download to prefetch models. With no --model or --config, it auto-discovers rag.toml/.rag.toml, deduplicates the models used by its entries, and downloads each one. Use --only <names> to prefetch a subset, --config <file> to select another config, or --model <id> for one explicit model. If no config exists, the built-in default model is downloaded.

You can control the cache location:

# Via flag
rag --cache-dir /path/to/cache index ./docs

# Via environment variable
export RAG_CACHE_DIR=/path/to/cache
rag index ./docs

# Or use HF_HOME (standard HuggingFace convention)
export HF_HOME=/path/to/hf
rag index ./docs

Supported file types

rag-cli is aimed at prose, docs, and config — not source code. It indexes:

  • Docs / prose: .md, .txt, .tex, .org, .rst
  • Data: .csv, .tsv, .log
  • Config: .toml, .conf, .cfg, .ini, .env, .tf, .hcl, .nix
  • Markup: .xml, .html
  • Schemas: .sql, .proto, .graphql
  • Build: Dockerfile, Makefile, .cmake

Files like Makefile, Dockerfile, LICENSE, README, and .gitignore are recognized by name.

Programming language sources (.rs, .py, .ts, .go, .js, shell scripts, etc.), .json, .yaml/.yml, and .css/.scss are intentionally not indexed — they tend to drown out useful matches with boilerplate. Index your code with a code-aware tool instead.

To index an extension that isn't in the allowlist (for example .mdx), add it with --ext (or an entry's extensions in rag.toml): rag index ./docs --ext mdx.

Hidden directories, node_modules, target, __pycache__, vendor, dist, and build are skipped automatically. Skip more with --exclude, or force a skipped directory back in with --include.

How it works

  1. Discover text files recursively, skipping binary and vendored content
  2. Chunk each file into overlapping segments (~512 chars), breaking at paragraph or line boundaries when possible
  3. Deduplicate and embed each unique chunk body once, in batches of 128, using all-MiniLM-L6-v2 with mean pooling and L2 normalization
  4. Store one F16 vector and content hash per unique chunk, plus compact source/byte-range occurrence records, in .rag/index.bin; inference remains int8 with f32 output, so F16 applies only to persisted vectors
  5. Search by embedding the query with the same model and ranking unique chunk bodies by exact cosine similarity, then reading matches from source files

License

Apache-2.0

About

Local-first RAG CLI powered by candle for semantic search over your files

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages