Sitelet https://github.com/reflex-search/reflex/pull/51
Skip to content

v2.1.0: incremental index, auto-update of stale indexes, 10-tool MCP surface - #51

Merged
therecluse26 merged 46 commits into
mainfrom
feature/auto-update
Sep 30, 2026
Merged

therecluse26 merged 46 commits into
mainfrom
feature/auto-update

Conversation

@therecluse26

@therecluse26 therecluse26 commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Reflex 2.1.0. The index is now updated in place instead of rebuilt, every command brings a stale index up to date before it answers, and the MCP server's tool list is a quarter of its former size.

1. Incremental index

rfx index no longer rewrites the whole index after a change.

  • Files whose size and mtime match the index are not read. A change is published as a small delta (a recent tier folded into a delta tier) plus tombstones over the base, named by .reflex/manifest.json. Past 2000 files or 5 % of the corpus, the delta is merged into a new base.
  • File ids are stable (INSERT … ON CONFLICT(path)), so symbols, dependencies and other branches' rows survive an update. files.walk_seq keeps every output in walk order.
  • An id-free planning size keeps query planning identical to a fresh build.
  • Indexer::update_paths(root, paths) updates named paths without walking the tree.
  • Kubernetes (27k files): a 1-file edit re-indexes in 0.3 s (was 8.5 s and ~1 GB of memory); nothing changed takes 0.19 s; a cold build 7.5 s (was 8.5 s). The library path takes 51–53 ms for a 1-file edit.
  • Equivalence: an updated index answers exactly like a fresh build of the same tree (golden battery on four corpora, property test, crash test at 11 write points).

2. Auto-update: no rfx index after an edit

Every command that reads the index (rfx query, deps, analyze, stats, context, list-files, ask, snapshot, pulse, interactive mode, rfx mcp tools, rfx serve) updates a stale index before it answers, and builds a missing one.

  • A search runs next to the freshness check, so a fresh index costs nothing extra. A stale one is updated (only the changed paths) and searched again.
  • --no-update (on every command, and rfx mcp / rfx serve) gives the old behaviour.
  • An update that cannot run (read-only .reflex/, another version's cache in a server) never fails the command: it answers from the current index, marked stale, with the reason in warnings.
  • Kubernetes, load ~20: rfx query after a 1-file edit 0.16–0.18 s; MCP search_code after an edit 138–161 ms. latency_budget +1.6 %, green.
  • Fixed on the way: .gitignore / .reflex/config.toml edits are now stale; a tracked file that .gitignore ignores is no longer reported as added; switching back to an indexed branch no longer reports fresh while the index holds the other branch; compaction no longer hides deleted files; index_project, POST /index, rfx watch, interactive mode and rfx ask now read .reflex/config.toml; rfx watch updates only the changed paths.

3. MCP: 10 tools, smaller schemas (⚠️ breaking)

Claude Code carries every listed tool schema on every turn, and that prefix was the whole token gap to Grep in long sessions.

  • tools/list 44 KB → 10.6 KB, 17 tools → 10:
    • count_occurrences → search_code with mode: "count" (now also returns files)
    • get_dependents / get_transitive_deps → get_dependencies with reverse: true / depth: N
    • find_hotspots, find_circular, find_unused, find_islands, analyze_summary → analyze with kind
  • The eight old names still work, unlisted, and answer with a deprecation warning. Clients that allow-list tools by name must add mcp__reflex__analyze.
  • The server no longer tells agents to call check_index_status / index_project. Every JSON-object answer carries status and can_trust_results.
  • list_locations takes preview: true (each matching line, 120 chars).
  • The example MCP configs set "alwaysLoad": true (schemas at session start, no ToolSearch turn).

Measured against built-in Grep

benches/efficacy/session_bench.py runs many questions in one Claude Code session (lookups, definitions, renames with lookups before and after), graded against ripgrep on the final tree. Cost vs Grep, schemas loaded at session start:

before this PR now
Sonnet 5, 12 questions 1.36× 1.05×
Sonnet 5, 50 questions 1.10–1.14× 1.11×
Opus 5.5, 12 questions 1.18× 1.15×
Opus 5.5, 50 questions — 0.89×

Accuracy 0.998–1.000 in every arm, including lookups right after edits.

Compatibility

  • The first rfx index (or any command) after upgrading rebuilds the index once (cache format changed). While a delta is live, an older rfx stops with "Cache appears to be corrupted" instead of answering from stale stores.
  • MCP: see the breaking change above.

Details: CHANGELOG.md, .context/INCREMENTAL_INDEX_RESEARCH.md, .context/AUTO_UPDATE_RESEARCH.md, .context/PERFORMANCE_RESEARCH.md.

🤖 Generated with Claude Code

therecluse26 and others added 30 commits September 29, 2026 00:00
The goal is token parity or better with built-in Grep on plain searches,
not routing them away from Reflex. CLAUDE.md states it; TODO.md lists the
steps (drop the pre-search status nudge, shrink the tool surface and test
alwaysLoad, trim reply payload), each re-measured with run-ref222.sh.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Backlog is now three sections: Grep parity (A: drop the status-check
nudge, B: shrink tools + alwaysLoad, D: trim payload), capabilities as
parameters on existing tools (enclosing symbol, many patterns per call,
co-occurrence, changed-files filter, callers-of-callers depth), and
other (auto-reindex, workspace resolution, multi-repo, reflexd, LSP).
Query result caching is dropped: latency was never the measured cost.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Replaces 1.7.2 policy 1 (honest staleness over auto-refresh) by user
decision: rfx mcp gets an in-process watcher and a bounded pre-query
catch-up, and only then drops the check_index_status nudge. A one-file
edit rebuilds the Reflex repo in 0.66 s (measured 2026-09-29); large trees
still need the incremental index path.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A one-file edit on a Kubernetes clone reprocessed all 27,448 files
(re-extract, full files-table rewrite, all dependencies; 27.7 s under
load, ~8 s idle), so auto-refresh without an incremental path cannot be
near-instant on large trees. Step A now includes a delta segment,
tombstones and background compaction, with a < 100 ms target.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Any change rebuilds content.bin and trigrams.bin from every file, and has
since v0.2.0; only the no-change shortcut and the hash-keyed symbol cache
are incremental. The watcher docs, rfx index --force / rfx watch help,
ai-agent-integration.md and API.md said otherwise.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…al builds

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
INSERT OR REPLACE into files gives every file a new id, and the cascade
deletes every symbols row (verified: 282 rows -> 0 after a one-file edit).
Corrects the claim added in e3ff843 that unchanged files keep symbols.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… writer docs

INCREMENTAL_INDEX_RESEARCH.md records how rfx index rebuilds today (two
id spaces joined by path, INSERT OR REPLACE wiping the symbol cache, the
two-rename race) and a staged design: stable metadata first, then one
delta segment with tombstones and a manifest publish point.

Also: max_posting_list_entries is dead configuration (never applied), so
the 'drops files past the cap' bug is removed; the Writers section now
describes TrigramIndexBuilder; new open bugs from the investigation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
benches/incremental/golden.sh runs a fixed battery (rfx query variants, deps,
analyze, stats, list-files, context, every MCP tool over stdio) against pinned
scratch copies of the efficacy corpora and tests/corpus, normalizing only
timings, timestamps and cache sizes. reference.sha256 records the pre-change
(2.0.3) outputs. perf.sh times cold / nothing-changed / 1-file-edit indexing
with peak RSS and load average.

Three 2.0.3 outputs are nondeterministic (HashMap order); the harness compares
them as sets and TODO.md records them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e config walk

The ~850-line reclassify/resolve block of index_with_callback moves unchanged
into src/dependency_resolve.rs (ResolverContext: reclassify, resolve_import,
resolve_export, resolve_file_imports), so an update can resolve one file's
imports, or re-resolve stored rows, without a full pass.

The seven resolver-config finders (go.mod, Maven/Gradle, Python, gemspec,
Cargo.toml, composer.json, tsconfig.json) each walked the whole tree. One walk
with their shared WalkBuilder settings now feeds new parse_*_from functions and
keeps every finder's own rules: the vendor and venv skips, the root Cargo.toml
gate, name-only matching for Cargo.toml and tsconfig.json, and which kinds fail
on a walk error. A unit test checks it against the seven finders.

Golden battery identical to 2.0.3 on all four corpora; dependency_equivalence
and the full debug suite pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… in the battery

Adds rfx pulse map (mermaid, d2, zoom), pulse glossary --no-llm --json, pulse
model --json, pulse changelog --no-llm, rfx snapshot, and rfx deps --depth 2 in
tree and table form: they read files and dependency rows in id or row order,
which stable ids could change. The pulse map edge order and the transitive
table are random in 2.0.3 and are compared as sets. Three captures of the
pre-change binary are identical; reference.sha256 is regenerated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ema change

meta.db now keeps one row per path with a stable id:
- files are upserted (ON CONFLICT(path) DO UPDATE ... RETURNING id), so the
  cascades no longer wipe the symbol cache, other branches' file_branches rows
  and exports, or null importers' resolved ids on every reindex;
- rows for files no longer in the tree are deleted by set difference;
- a full build clears and rewrites every dependency and export row in walk
  order (exports have no key, so stale rows could otherwise survive).

Outputs that listed files in id order keep the order a fresh build gives:
files.walk_seq records walk position, and get_dependents, find_unused, hotspot
ties, island adjacency, the Pulse glossary hotspots and the transitive tree sort
by it. The transitive table's 'File ID' column prints the 1-based walk position,
which is the id a fresh build assigns.

The symbol cache is keyed by files.hash (the stored bytes) in the query path
and the background pass; cleanup_stale also drops versions no file or branch
holds. The Pulse glossary reads only current-version symbols; the snapshot
fingerprint hashes files rows (identical to before on a fresh build).

The schema-hash check now reads the stored hash before init() stamps it, so a
cache from other code is rebuilt in full (the check was dead). init() no longer
resets total_files / last_compaction, so the unlocked background compaction no
longer starts on every command; it now also skips rfx index and holds
index.lock. A second build-time hash (EXTRACTION_HASH: src/parsers, line_filter,
dependency_resolve) clears the symbol cache when extraction code changes.
build.rs also covers src/trigram_build.rs.

Golden battery identical to 2.0.3; the update-and-revert flow matches a fresh
build of the same directory on all four corpora; tests/stable_ids.rs added.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ndencies

rfx index now compares each discovered file's (size, mtime) with its files
row and reads only the files whose stat differs; the nothing-changed path no
longer reads or hashes the tree (nor NUL-sniffs unchanged non-code files).
Deleted files are the rows the walk no longer finds. git state and the
resolver-config walk run in parallel with discovery, and stats reuse the
branch instead of running git again.

A run that changes content still rewrites both stores from every file (stage
1 replaces that with a delta), but meta.db work is proportional to the change,
in one transaction (src/meta_update.rs): rows of added, modified and touched
files; deletions by id; walk positions that moved (longest increasing run kept,
gaps filled); dirty flags that flipped; this branch's hash rows. Imports are
extracted only for files whose bytes changed; their dependency and export rows
are replaced. When files are added or removed, every other stored import and
export is resolved again (suffix matching makes resolution depend on the whole
path set). A digest of every resolver config input (stored as
resolver_config_digest) triggers a full dependency pass when any config,
including a tsconfig alias, changes.

The race threshold for recorded mtimes is now the mtime of .reflex/.index-run,
written at run start on the file system's own clock.

Golden battery identical to 2.0.3; update-and-revert flow equals a fresh build
on all four corpora; tests/incremental_meta.rs added (mutation-checked).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ader

Every index run now writes its stores as generation files (content.<g>.bin,
trigrams.<g>.bin) and publishes them by renaming .reflex/manifest.json into
place, the single commit point. meta.db follows in one transaction that
records the same generation (statistics.index_generation); a run that stops
between the two is caught by the next one (generations differ -> rebuild),
and meanwhile the changed files read as stale. Files no manifest names any more
are deleted on the next publish (the previous generation is kept for readers
that just read the old manifest). This closes the two-rename race in which a
reader could pair a new trigrams.bin with an old content.bin.

content.bin / trigrams.bin remain as hard links to the current base for older
binaries and tools; where hard links fail they are removed.

Every reader of the stores goes through src/snapshot.rs::IndexSnapshot (live
ids, content, paths, context lines, candidate lookups): the query engine,
OpenIndex (whose registry fingerprint now stamps the manifest), the background
symbol pass, Pulse extraction, rfx query's grouped JSON, validate(), stats()
sizes and rfx stats' trigram count. Error texts keep the logical names
content.bin / trigrams.bin. Symbol parses are cached, and the fingerprint memo
kept, only while meta.db's generation matches the snapshot's.

Golden battery identical to 2.0.3; update-and-revert flow equals a fresh build
on all four corpora; full debug suite passes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
search_candidates stops intersecting early, and its order and stop rule used
each trigram's on-disk byte size. Those bytes include every block's file-id
delta, so an index made of a base plus later updates could never plan exactly
like a full build of the same tree, and neither could two full builds of the
same tree in different directory orders.

The planner now uses the planning size: for each file block, 1 +
len(varint(n_lines << 1)) + the line-delta varints, i.e. the V4 encoding with
the file-id delta counted as one byte. It adds up per file, so base - deleted +
added gives exactly what a full build would have. It equals the on-disk size
whenever every file-id delta is below 128.

The builder computes it per record while encoding (partial record header +4
bytes; partials are temporary), the merge sums it, and rfx index writes
trigrams.<g>.plan (u32 per directory entry) next to trigrams.<g>.bin; the
manifest names it and the snapshot attaches it. An index without it plans by
bytes, as before.

The lazy search paths (search_candidates, search_candidates_fold,
exotic_fold_lines) now run through one planner over 'lists made of parts'
(plan_intersection, plan_fold_intersection, union_lists), ready for a delta
segment and tombstones.

Golden battery identical to 2.0.3 on all four corpora (the approved checkpoint:
0 differences in approx_total, total_is_exact, has_more, excluded_by_default);
update flow equals a fresh build; full debug suite passes; builder tests check
the planning size of every list and equality with bytes below 128 files.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ing the stores

After a change that is not a full-rebuild case (schema, extraction code or a
resolver config changed), rfx index now reads only the added and modified
files and publishes:

- delta.<g>.content.bin / delta.<g>.trigrams.bin / delta.<g>.plan: the read
  files plus the previous delta's files that are still present and unchanged,
  in walk order, with local ids after the base's;
- tombstones in the manifest: base ids superseded or gone;
- delta.<g>.tomb: the planning size of every tombstoned posting, carried
  forward, so base - tomb + delta equals a full build's planning size exactly;
- live trigram count and live corpus bytes, as a full build would record them.

The base is not rewritten. Publish order: delta files, unlink content.bin /
trigrams.bin (an older binary then stops with CacheCorrupted instead of
serving the base alone), manifest, one meta.db transaction with the
generation, invalidate, remove files named by neither the current nor the
previous manifest. When the delta is empty again, the fixed names come back as
hard links.

A delta past 2000 files or 5 % of the live corpus bytes is merged: the run
falls back to a full build (Indexer::set_merge_limits overrides the limits for
tests; no config key or env var).

IndexSnapshot reads base + delta through one composite list per trigram; a
trigram whose every file is tombstoned has no list (a full build has none, and
the fold planner counts variants). Live corpus bytes come from the content
entry table, not the content pages.

Golden battery identical to 2.0.3 on all four corpora, fresh and after the
scripted edit-and-revert updates (the updated index keeps its delta through
the whole battery; golden.sh now fails if the generation falls back to 1).
New tests: tests/incremental_delta.rs (layout, old names hidden, every change
kind vs a fresh build, merge limit + hard links, cleanup, no-op run) and a
snapshot unit test comparing every trigram's planning size and candidates with
a fresh build (fails without the no-list rule). Full release suite passes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A caller that knows what changed (a watcher, an editor) names the paths;
update_paths brings the index to what Indexer::index would build without
walking the tree:

- the named paths (files or directories) are walked through the same walker,
  entering only them and their ancestors, so every ignore rule applies as in
  a full walk; rows at and under them come from index lookups;
- a moved HEAD adds `git diff --name-only` paths; `git status` runs on the
  named paths only, on a thread (decision 5: branches.is_dirty can stay true
  after a revert until the next rfx index);
- walk positions: an edited file keeps its position while its neighbours
  still bracket it (by readdir order, src/walk_order.rs); a new or moved file
  is placed by a binary search over files.walk_seq (new index
  idx_files_walk_seq);
- resolver configs come from .reflex/resolver-configs.json, which every walk
  now writes; a named ignore file, .reflex/config.toml, resolver config, a new
  directory holding one, a branch change, or the merge limit falls back to
  Indexer::index.

rfx index and update_paths now publish a delta through one change-set path
(publish_delta): it writes only the named rows, sets only their branch rows
(rfx index still re-syncs all), and folds the branch row and the total_files,
schema and extraction stamps into the meta.db transaction (three commits
fewer). stats_on_branch skips its two debug-only scans unless debug logging
is on.

Kubernetes 1-file edit through update_paths: 73-103 ms at load 9 (was
260-500 ms before the change set); an idle measurement follows with the perf
gates.

Tests: tests/incremental_equivalence.rs (seeded random edits, adds,
deletes, renames, directory renames, atomic saves, resolver-config and
.gitignore edits, ignored/hidden/binary files, commits and branch switches,
each followed by index or update_paths, compared with a fresh build: query
battery, dependency and export rows, analyses, walk order, snapshot shape;
4x18 steps by default, 40x30 ignored: 629 direct updates, 175 fallbacks, all
equal); tests/incremental_crash.rs (the process exits at each of 11 write
points of a delta, a library update and a full build; the cache answers as
the old or the new tree, never fresh on old content, and recovers, also when
the tree is reverted before recovery; fails without the generation check).
Golden battery identical on four corpora, fresh and after updates; full
release suite passes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…es, open retry

- tests/incremental_concurrency.rs: a writer moves a tree through six states
  (index and update_paths, every third publish a merge into a new base) while
  three reader threads and a reader child process query in a loop; every
  answer is some state's answer, and no query fails (about 13k answers a run).
- tests/incremental_cross_version.rs (ignored; needs the 2.0.3 binary at
  /scratch/cache/rfx-pre-incremental): 2.0.3 reads a base-only cache as
  stale, stops with "corrupted" on a live delta instead of answering from the
  base, and after a 2.0.3 index run (which replaces the hard-linked
  content.bin, leaving content.<g>.bin intact) this binary reports stale and
  its next index run rebuilds.
- IndexSnapshot::open reads the manifest through open_reading, so a unit test
  can show the NotFound retry opening the next manifest and the error after
  the last retry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
When the delta would pass its limits (or the dependency rows are all
rewritten) and the published stores match meta.db, the new base takes each
unchanged or touched file's text from the stores instead of the disk, and its
hash from its row; only added and modified files are read. Same text, same
walk order, same builder: the base is byte-identical to a build that reads
every file (tests/incremental_delta.rs compares the three store files with a
fresh build's). A test makes an unchanged file unreadable: the merge still
indexes it, and fails when the merge reads from disk.

publish_delta checks the merge limits from stat sizes and entry-table
lengths before it reads anything, so a merge no longer reads the changed
files twice.

Kubernetes, 1500 Go files edited (16.9 MB, past the 12.3 MB limit): merge
3.42 s against a 7.38 s cold build (load 8-9; 25,948 files from the stores).
Idle numbers follow with the perf gates.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… delta

With one delta segment every update rebuilt it whole: on Kubernetes a
1-file update_paths with a 1000-file (11 MB) delta took 191-207 ms (load
9-13), 128-134 ms of it writing the delta - past the 100 ms line the plan
set for tiering.

Now the snapshot is base + delta + recent:
- each update rebuilds only the small recent segment (the files it read plus
  the recent files it did not touch);
- past 256 files or 1/16 of the delta byte limit (Indexer::set_recent_limits
  for tests), the recent segment is folded with the delta's live files into
  a new delta;
- tombstones are global ids over base and delta; delta.<g>.dtomb holds the
  planning sizes of dead delta postings (dropped at a fold), so every
  trigram's live planning size stays exactly a fresh build's;
- the live trigram count is updated from the trigrams an update touches
  (dead files, replaced and new segment) instead of scanning every dead
  posting;
- the merge limits are checked on the whole (delta + recent) before any read.

Manifest format 2 (recent, tomb_delta). build.rs now hashes src/snapshot.rs
and src/meta_update.rs into the cache schema.

The same update now takes 71 ms (load 15; the recent segment's write is
19-21 ms). Query cost with a 1000-file delta live stayed within noise
(examples/delta_threshold_timing.rs), so no skip pointers.

Tests: a snapshot unit test walks two folds and a delta tombstone while a
recent segment is live, comparing every trigram's planning size, the live
counts and candidates with a fresh build; some property-test seeds fold
every 2 files (40x30 ignored run: 630 direct updates, all equal); the crash
test gains a fold-every-update mode. Golden battery identical on four
corpora, fresh and after updates; full release suite passes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Library path (update_paths), Kubernetes 1-file edit, load 7: 73-110 ms ->
49-52 ms warm.
- The path resolver of the last publish stays in the process, keyed by a
  random publish_id the manifest now carries (a generation number repeats
  after the cache is cleared), and is patched with the added and deleted
  paths instead of reloading 27k paths.
- git status of the named paths is joined only when meta.db is written, so
  it runs while the stores are written; if it fails, every named path counts
  as dirty (a superset, safe for freshness).
- unlink_fixed_names no longer fsyncs the directory when the names were
  already gone.

rfx index with nothing changed, Kubernetes, load 3-5: 0.82-0.95 s (2.0.3)
-> 186-192 ms, against a 142-148 ms walk:
- the branch whose rows the last run synced is recorded
  (statistics.synced_branch, in the syncing transaction); on it, the run
  neither loads the branch's hashes nor re-syncs every row, and counts come
  from the rows themselves;
- a cache this binary completed a run on skips the schema transaction in
  init() (config.toml is still recreated when missing);
- the stored rows load on a thread during the walk (the NUL check of changed
  non-code files moved after the walk, in the pool); the df check too;
- the branch row and total_files are written in the refresh transaction;
  statistics come from the rows already loaded (no query); each walked
  path's row is looked up once; plan_walk_seq returns early when the order
  is unchanged;
- the symbol-pass yield polls from 5 ms up to 100 ms (was 100 ms).
stats_synced (files-only counts) replaces the three joins wherever the
branch was just synced; results are the same.

Golden battery identical on four corpora, fresh and after updates;
property test 40x30 (636 direct updates, all equal to fresh builds); full
release suite passes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… changelog

- CLAUDE.md: the .reflex/ layout (manifest, generation files, delta tiers,
  resolver-configs.json, .index-run), an "Incremental updates" section, the
  symbol pass and change-detection notes.
- docs/ARCHITECTURE.md: cache table, versioning (both hashes), the indexing
  pipeline with the delta and merge paths and update_paths.
- .context/BINARY_FORMAT_RESEARCH.md §4: manifest format 2, RFPL planning
  sizes, RFTB tombstone sizes, fixed names and older binaries.
- .context/INCREMENTAL_INDEX_RESEARCH.md: "As built" (what changed from the
  plan and why, what was tried and dropped).
- .context/PERFORMANCE_RESEARCH.md: the Kubernetes A/B gates, the library
  path, the merge, the two tiers, the live-delta query cost (against a fresh
  build of the same tree), latency_budget over 12 runs.
- CHANGELOG Unreleased; TODO: decisions and open follow-ups.
- examples/delta_threshold_timing.rs: per-phase medians and --fresh-too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Measured on Kubernetes (27,448 files) against 2.0.3, alternating runs:
- rfx index with nothing changed: peak RSS 45.6 MiB (+25 % over 2.0.3's
  36.4) -> 34.5-34.8 MB against 37.3-37.8 MB. Two causes:
  - stores_intact opened the whole snapshot (mapped stores, walked the path
    tables: an 8 MB transient peak). It now checks the manifest, the file
    sizes and each store's 32-byte header (snapshot::check_published);
    whatever opens the stores next validates them in full.
  - the walk kept a full std::fs::Metadata (~150 B) and an absolute PathBuf
    per file. It keeps FileStat (size + mtime, 24 B) and the relative path;
    readers build root.join(rel) when they read a file.
- a change past the merge limit (1,500 files): 1,138 MiB (+8 % over 2.0.3's
  full rebuild, 1,052) -> 947 MiB against 1,056. The merge read unchanged
  files' text from the mapped old stores, whose pages stayed resident;
  after each batch it now drops them (MADV_DONTNEED on the read-only
  mappings; the page cache keeps them).
- cold build: median 1,030 MiB against 1,059 (8 runs each); unchanged path.

examples/update_paths_timing.rs prints the process's RSS: a process running
update_paths levels off at 61-67 MiB over 150 edit/revert/query rounds.

Golden battery identical on four corpora, fresh and after updates; property
test 40x30 (628 direct updates, all equal); full release suite passes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-cache race

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…dex run

- CacheManager::effective_index_config: config.toml plus --languages; rfx index
  saves its --languages (statistics.languages_override) so later runs index the
  same languages. index_project, POST /index, rfx watch, interactive mode and
  rfx ask used IndexConfig::default() and ignored config.toml.
- The version-mismatch self-heal in rfx index keeps --languages.
- BackgroundIndexer::spawn_detached replaces two copies of the spawn code.
- LOCK_WAIT_FOREVER; index_project and POST /index wait for index.lock.
- QueryEngine's open index can be reset; the freshness memo is not stored when
  an index write invalidated it during the check (epoch).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s are not misreported

- An index run records the blake3 of .reflex/config.toml, the root ignore files
  and every ignore file git lists as dirty (statistics.rule_files). The check
  compares them and lists a changed one under files_modified: a .gitignore or
  config edit changes WHICH files are indexed, and was never stale.
- A new file counts as added only when the walk reaches it (a narrowed walk of
  its ancestors): a tracked file that .gitignore ignores was reported added by
  every check, on a fresh build too.
- Compaction no longer deletes the rows of missing files. It left them in the
  stores while the check, which compares rows with the disk, lost the deletion.
  It now only runs VACUUM when meta.db has free pages; files_removed is 0.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
reflex::auto_update::update_if_stale asks the freshness check for a plan
(query::update_plan): nothing, update_paths on the listed paths, or a full
index run (many changes, a rule-file edit, a format change). No index is built.
A cache another released version wrote is rebuilt only when the caller asks
(the CLI); servers skip it. Every other failure returns Skipped(reason), never
an error. One update per workspace per process; index.lock is waited for as
long as it takes; a symbol pass that keeps making progress is waited out; the
same plan over the same bytes is not retried after it failed to make the
index fresh. Not wired into any command yet.

The check now compares the tree with the commit of the LAST index run (the
synced branch), not the current branch's row: after switching back to an
indexed branch, 2.0.3 reported fresh while the index held the other branch.

tests/auto_update.rs: an agent-style sequence (edit, add, delete, rename,
revert, commit, branch switch and back, .gitignore and config edits, 150 new
files) where every answer equals a fresh build of the same tree: 15/15.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- QueryEngine::with_update: search_with_metadata runs the search and the
  freshness check together as before; a stale verdict triggers update_if_stale
  and a second search (at most two updates per call). search, AST search,
  --symbols and list_by_kind update first. The engine drops its open index
  before an update. QueryEngine::new keeps answering from the index as it is.
- CLI: a global --no-update flag. rfx query and interactive mode search through
  cli::engine; stats, list-files, analyze, deps, ask, context, snapshot and
  pulse update before they run. A missing index is built (stderr says so).
  rfx query --ast --json reports the real status instead of fresh.
- rfx mcp [--no-update]: search tools update through the engine; every other
  tool except index_project and check_index_status updates at the
  handle_call_tool chokepoint; a skipped update is a warning.
  run_mcp_server_io_with runs a test server with an update.
- rfx serve [--no-update]: /query and /stats update; blocking work runs in
  spawn_blocking.
- rfx watch passes the collected paths to update_paths and reacts to
  ignore-file and config edits.
- timings.update_us when an update ran.
- tests/auto_update_front_ends.rs: rfx query / deps / stats and MCP tools after
  an edit and with no index, --no-update, and a guard that fails on any
  QueryEngine::new in src/ outside the known front ends.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… scripts

benches/incremental/auto_update.sh and mcp_edit_latency.py time a query after
an edit (new process and rfx mcp session), and the commands that now check
before they run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
therecluse26 and others added 16 commits September 30, 2026 00:05
… git status

After update_paths, only the updated paths are compared with the index; when
all match, fresh becomes the memoised verdict for the rest of the original
check's window (the terms the 1 s memo already has). The second full check
was a git status on every stale query.

Kubernetes, load 18-24: MCP edit-then-search_code 200-340 ms -> 138-161 ms
(update step 160-200 ms -> 53-68 ms); rfx query after an edit in a new process
0.23-0.44 s -> 0.16-0.18 s. The timing script prints the engine's phases.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
golden.sh's index() ran `rfx index` in the directory the script was started
from, so the update half never indexed the edited trees: every battery command
failed with "Index not found" on both sides, and identical errors compared
equal. The 2026-09-29 "identical after scripted updates" result was vacuous.
The fresh index of that comparison is now parked outside the tree (an
untracked .reflex-inc/ made git call the tree dirty).

Rerun 2026-09-30: 0 of 76 outputs differ on each of the four corpora, for the
incremental branch (c2e0de2) and for auto-update; each side has the one
failing output the 2.0.3 reference has (q_two_char).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The index is updated before every call, so the server instructions and every
tool description now say so: no FRESHNESS paragraph asks for index_project and
a retry, no "Index not found → index_project" sentence remains, and the
check_index_status / index_project descriptions call them rarely needed (a probe
that never updates; a forced run). The paragraph is shorter in 14 descriptions.

Every JSON-object answer now carries status and can_trust_results (added from
the memoised verdict when the tool lacked them): list_locations,
count_occurrences, find_references, mode: count, and the structural tools.
Array answers and the path-keyed get_transitive_deps are unchanged.

Tests: every such answer says true with auto-update and false (stale) without;
tools/list and the instructions no longer contain the old advice.
README, CLAUDE.md, the cheatsheet, CHANGELOG and TODO follow.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Removing "Only fall back to Grep/Glob after index_project has been called and
the tool still fails" (9e30ff5) dropped Reflex adoption in the efficacy
harness from 37/72 (2.0.3) to 2/72: Claude Code defers the tool schemas, so
the instructions are all the agent reads before choosing Grep or a ToolSearch.

Pilots, 3 find-all tasks x 4 trials, Opus 5.5, Claude Code 2.1.284, trials
that called Reflex: old text (c2e0de2) 8/12; "only fall back to Grep/Glob if a
Reflex tool fails" 4/12; "if a Reflex tool fails, retry it once; only fall
back to Grep/Glob after the retry also fails" (this commit) 8/12.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… and adoption

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…n the harness

README and docs/ai-agent-integration.md register Reflex with
`claude mcp add-json ... "alwaysLoad": true` (claude mcp add has no flag for it),
so Claude Code loads the tool schemas at session start instead of deferring
them behind a ToolSearch turn.

benches/efficacy: arm Beager = arm B with alwaysLoad in its MCP config;
analyze.py and plots.py know it. Measured 2026-09-30 (9 find-all tasks x 8,
three arms at once), vs Grep: Sonnet 5 tokens 1.63 -> 1.53, cost 1.77 -> 1.20;
Opus 5.5 tokens 1.68 -> 1.82, cost 1.32 -> 0.98; turns equal to Grep on both.
The 44 KB tools/list on every turn is the remaining gap.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…question sessions

benches/efficacy/session_bench.py runs many questions in ONE Claude Code
session (stream-json input, one message after the previous answer), each trial
on a fresh corpus copy, and grades every answer against ripgrep on the final
tree. tasks/sessions.json: investigate, edit (renames with lookups before and
after) and 50-question sessions per corpus.

Results (2026-09-30), cost vs Grep with alwaysLoad: 1.36x Sonnet / 1.18x Opus
over 12 questions, 1.10x over 50 (Sonnet; cheaper than Grep on tokio). Per
query Reflex costs what Grep costs; the gap is the ~16K-token schema prefix on
every turn. Accuracy 1.000 for Reflex throughout, including lookups after edits
(auto-update). With deferred schemas agents mostly skip Reflex in long sessions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…KB -> 10.6 KB)

Claude Code carries every listed tool schema on every turn; in long sessions
that ~16K-token prefix was the whole token gap between Reflex and Grep
(session_bench.py, 2026-09-30).

- count_occurrences -> search_code mode:"count" (count mode now also returns
  files); get_dependents / get_transitive_deps -> get_dependencies with
  reverse:true / depth:N; find_hotspots, find_circular, find_unused,
  find_islands, analyze_summary -> analyze {kind}. enable_structural_tools now
  hides only analyze.
- The eight old names still work, unlisted (LEGACY_TOOLS); object answers carry a
  deprecation warning naming the replacement.
- Descriptions and parameter text are short; the matching, coverage and
  freshness rules moved into the server instructions (said once per session;
  budget 1700 -> 2400 chars).
- tests/mcp_tool_surface.rs: the list, a 14 KB size guard, analyze kinds,
  get_dependencies reverse/depth, the old names, count-mode files.
- Harness allowlist names analyze; docs follow.

BREAKING CHANGE: tools/list no longer lists the eight merged tools; clients that
allow-list tools by name must add mcp__reflex__analyze.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The short descriptions (703d421) lost the old steering: in 50-question
sessions agents answered "where does X occur" with search_code and
find_references (1.3-1.9K chars each) instead of list_locations (~0.7K), and
tool output per session rose 43K -> 76K chars, which every later turn re-reads.
One sentence in the instructions and a pointer in two descriptions bring it
back (43 list_locations calls per 50-question session).

Sonnet 5, cost vs Grep with alwaysLoad: 12-question sessions 1.19x -> 1.05x
(cache-weighted tokens 1.23x -> 1.07x); 50-question sessions 1.14x -> 1.11x.
Accuracy 1.000.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The root line is the random temp directory name; one ending in c
(.tmpE0KsIc/) still matched "ends with c/".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each location carries preview: the matching line, trimmed and cut to 120
characters. In short sessions Opus called the preview-less list_locations once,
re-ran the search with grep to see the lines, and stayed on grep. The
instructions and the description mention the option; preview and reverse are
coerced from strings like the other flags. Smoke test (Opus, two 12-question
sessions): one session used Reflex throughout (8 calls with preview), one
drifted to grep after a count cross-check.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@therecluse26 therecluse26 changed the title docs: record the Grep-parity goal for Reflex MCP v2.1.0: incremental index, auto-update of stale indexes, 10-tool MCP surface Sep 30, 2026
@therecluse26
therecluse26 merged commit da89213 into main Sep 30, 2026
6 of 7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant