v2.1.0: incremental index, auto-update of stale indexes, 10-tool MCP surface - #51
Merged
Merged
Conversation
The goal is token parity or better with built-in Grep on plain searches, not routing them away from Reflex. CLAUDE.md states it; TODO.md lists the steps (drop the pre-search status nudge, shrink the tool surface and test alwaysLoad, trim reply payload), each re-measured with run-ref222.sh. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Backlog is now three sections: Grep parity (A: drop the status-check nudge, B: shrink tools + alwaysLoad, D: trim payload), capabilities as parameters on existing tools (enclosing symbol, many patterns per call, co-occurrence, changed-files filter, callers-of-callers depth), and other (auto-reindex, workspace resolution, multi-repo, reflexd, LSP). Query result caching is dropped: latency was never the measured cost. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Replaces 1.7.2 policy 1 (honest staleness over auto-refresh) by user decision: rfx mcp gets an in-process watcher and a bounded pre-query catch-up, and only then drops the check_index_status nudge. A one-file edit rebuilds the Reflex repo in 0.66 s (measured 2026-09-29); large trees still need the incremental index path. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A one-file edit on a Kubernetes clone reprocessed all 27,448 files (re-extract, full files-table rewrite, all dependencies; 27.7 s under load, ~8 s idle), so auto-refresh without an incremental path cannot be near-instant on large trees. Step A now includes a delta segment, tombstones and background compaction, with a < 100 ms target. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Any change rebuilds content.bin and trigrams.bin from every file, and has since v0.2.0; only the no-change shortcut and the hash-keyed symbol cache are incremental. The watcher docs, rfx index --force / rfx watch help, ai-agent-integration.md and API.md said otherwise. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…al builds Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
INSERT OR REPLACE into files gives every file a new id, and the cascade deletes every symbols row (verified: 282 rows -> 0 after a one-file edit). Corrects the claim added in e3ff843 that unchanged files keep symbols. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… writer docs INCREMENTAL_INDEX_RESEARCH.md records how rfx index rebuilds today (two id spaces joined by path, INSERT OR REPLACE wiping the symbol cache, the two-rename race) and a staged design: stable metadata first, then one delta segment with tombstones and a manifest publish point. Also: max_posting_list_entries is dead configuration (never applied), so the 'drops files past the cap' bug is removed; the Writers section now describes TrigramIndexBuilder; new open bugs from the investigation. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
benches/incremental/golden.sh runs a fixed battery (rfx query variants, deps, analyze, stats, list-files, context, every MCP tool over stdio) against pinned scratch copies of the efficacy corpora and tests/corpus, normalizing only timings, timestamps and cache sizes. reference.sha256 records the pre-change (2.0.3) outputs. perf.sh times cold / nothing-changed / 1-file-edit indexing with peak RSS and load average. Three 2.0.3 outputs are nondeterministic (HashMap order); the harness compares them as sets and TODO.md records them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e config walk The ~850-line reclassify/resolve block of index_with_callback moves unchanged into src/dependency_resolve.rs (ResolverContext: reclassify, resolve_import, resolve_export, resolve_file_imports), so an update can resolve one file's imports, or re-resolve stored rows, without a full pass. The seven resolver-config finders (go.mod, Maven/Gradle, Python, gemspec, Cargo.toml, composer.json, tsconfig.json) each walked the whole tree. One walk with their shared WalkBuilder settings now feeds new parse_*_from functions and keeps every finder's own rules: the vendor and venv skips, the root Cargo.toml gate, name-only matching for Cargo.toml and tsconfig.json, and which kinds fail on a walk error. A unit test checks it against the seven finders. Golden battery identical to 2.0.3 on all four corpora; dependency_equivalence and the full debug suite pass. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… in the battery Adds rfx pulse map (mermaid, d2, zoom), pulse glossary --no-llm --json, pulse model --json, pulse changelog --no-llm, rfx snapshot, and rfx deps --depth 2 in tree and table form: they read files and dependency rows in id or row order, which stable ids could change. The pulse map edge order and the transitive table are random in 2.0.3 and are compared as sets. Three captures of the pre-change binary are identical; reference.sha256 is regenerated. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ema change meta.db now keeps one row per path with a stable id: - files are upserted (ON CONFLICT(path) DO UPDATE ... RETURNING id), so the cascades no longer wipe the symbol cache, other branches' file_branches rows and exports, or null importers' resolved ids on every reindex; - rows for files no longer in the tree are deleted by set difference; - a full build clears and rewrites every dependency and export row in walk order (exports have no key, so stale rows could otherwise survive). Outputs that listed files in id order keep the order a fresh build gives: files.walk_seq records walk position, and get_dependents, find_unused, hotspot ties, island adjacency, the Pulse glossary hotspots and the transitive tree sort by it. The transitive table's 'File ID' column prints the 1-based walk position, which is the id a fresh build assigns. The symbol cache is keyed by files.hash (the stored bytes) in the query path and the background pass; cleanup_stale also drops versions no file or branch holds. The Pulse glossary reads only current-version symbols; the snapshot fingerprint hashes files rows (identical to before on a fresh build). The schema-hash check now reads the stored hash before init() stamps it, so a cache from other code is rebuilt in full (the check was dead). init() no longer resets total_files / last_compaction, so the unlocked background compaction no longer starts on every command; it now also skips rfx index and holds index.lock. A second build-time hash (EXTRACTION_HASH: src/parsers, line_filter, dependency_resolve) clears the symbol cache when extraction code changes. build.rs also covers src/trigram_build.rs. Golden battery identical to 2.0.3; the update-and-revert flow matches a fresh build of the same directory on all four corpora; tests/stable_ids.rs added. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ndencies rfx index now compares each discovered file's (size, mtime) with its files row and reads only the files whose stat differs; the nothing-changed path no longer reads or hashes the tree (nor NUL-sniffs unchanged non-code files). Deleted files are the rows the walk no longer finds. git state and the resolver-config walk run in parallel with discovery, and stats reuse the branch instead of running git again. A run that changes content still rewrites both stores from every file (stage 1 replaces that with a delta), but meta.db work is proportional to the change, in one transaction (src/meta_update.rs): rows of added, modified and touched files; deletions by id; walk positions that moved (longest increasing run kept, gaps filled); dirty flags that flipped; this branch's hash rows. Imports are extracted only for files whose bytes changed; their dependency and export rows are replaced. When files are added or removed, every other stored import and export is resolved again (suffix matching makes resolution depend on the whole path set). A digest of every resolver config input (stored as resolver_config_digest) triggers a full dependency pass when any config, including a tsconfig alias, changes. The race threshold for recorded mtimes is now the mtime of .reflex/.index-run, written at run start on the file system's own clock. Golden battery identical to 2.0.3; update-and-revert flow equals a fresh build on all four corpora; tests/incremental_meta.rs added (mutation-checked). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ader Every index run now writes its stores as generation files (content.<g>.bin, trigrams.<g>.bin) and publishes them by renaming .reflex/manifest.json into place, the single commit point. meta.db follows in one transaction that records the same generation (statistics.index_generation); a run that stops between the two is caught by the next one (generations differ -> rebuild), and meanwhile the changed files read as stale. Files no manifest names any more are deleted on the next publish (the previous generation is kept for readers that just read the old manifest). This closes the two-rename race in which a reader could pair a new trigrams.bin with an old content.bin. content.bin / trigrams.bin remain as hard links to the current base for older binaries and tools; where hard links fail they are removed. Every reader of the stores goes through src/snapshot.rs::IndexSnapshot (live ids, content, paths, context lines, candidate lookups): the query engine, OpenIndex (whose registry fingerprint now stamps the manifest), the background symbol pass, Pulse extraction, rfx query's grouped JSON, validate(), stats() sizes and rfx stats' trigram count. Error texts keep the logical names content.bin / trigrams.bin. Symbol parses are cached, and the fingerprint memo kept, only while meta.db's generation matches the snapshot's. Golden battery identical to 2.0.3; update-and-revert flow equals a fresh build on all four corpora; full debug suite passes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
search_candidates stops intersecting early, and its order and stop rule used each trigram's on-disk byte size. Those bytes include every block's file-id delta, so an index made of a base plus later updates could never plan exactly like a full build of the same tree, and neither could two full builds of the same tree in different directory orders. The planner now uses the planning size: for each file block, 1 + len(varint(n_lines << 1)) + the line-delta varints, i.e. the V4 encoding with the file-id delta counted as one byte. It adds up per file, so base - deleted + added gives exactly what a full build would have. It equals the on-disk size whenever every file-id delta is below 128. The builder computes it per record while encoding (partial record header +4 bytes; partials are temporary), the merge sums it, and rfx index writes trigrams.<g>.plan (u32 per directory entry) next to trigrams.<g>.bin; the manifest names it and the snapshot attaches it. An index without it plans by bytes, as before. The lazy search paths (search_candidates, search_candidates_fold, exotic_fold_lines) now run through one planner over 'lists made of parts' (plan_intersection, plan_fold_intersection, union_lists), ready for a delta segment and tombstones. Golden battery identical to 2.0.3 on all four corpora (the approved checkpoint: 0 differences in approx_total, total_is_exact, has_more, excluded_by_default); update flow equals a fresh build; full debug suite passes; builder tests check the planning size of every list and equality with bytes below 128 files. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ing the stores After a change that is not a full-rebuild case (schema, extraction code or a resolver config changed), rfx index now reads only the added and modified files and publishes: - delta.<g>.content.bin / delta.<g>.trigrams.bin / delta.<g>.plan: the read files plus the previous delta's files that are still present and unchanged, in walk order, with local ids after the base's; - tombstones in the manifest: base ids superseded or gone; - delta.<g>.tomb: the planning size of every tombstoned posting, carried forward, so base - tomb + delta equals a full build's planning size exactly; - live trigram count and live corpus bytes, as a full build would record them. The base is not rewritten. Publish order: delta files, unlink content.bin / trigrams.bin (an older binary then stops with CacheCorrupted instead of serving the base alone), manifest, one meta.db transaction with the generation, invalidate, remove files named by neither the current nor the previous manifest. When the delta is empty again, the fixed names come back as hard links. A delta past 2000 files or 5 % of the live corpus bytes is merged: the run falls back to a full build (Indexer::set_merge_limits overrides the limits for tests; no config key or env var). IndexSnapshot reads base + delta through one composite list per trigram; a trigram whose every file is tombstoned has no list (a full build has none, and the fold planner counts variants). Live corpus bytes come from the content entry table, not the content pages. Golden battery identical to 2.0.3 on all four corpora, fresh and after the scripted edit-and-revert updates (the updated index keeps its delta through the whole battery; golden.sh now fails if the generation falls back to 1). New tests: tests/incremental_delta.rs (layout, old names hidden, every change kind vs a fresh build, merge limit + hard links, cleanup, no-op run) and a snapshot unit test comparing every trigram's planning size and candidates with a fresh build (fails without the no-list rule). Full release suite passes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A caller that knows what changed (a watcher, an editor) names the paths; update_paths brings the index to what Indexer::index would build without walking the tree: - the named paths (files or directories) are walked through the same walker, entering only them and their ancestors, so every ignore rule applies as in a full walk; rows at and under them come from index lookups; - a moved HEAD adds `git diff --name-only` paths; `git status` runs on the named paths only, on a thread (decision 5: branches.is_dirty can stay true after a revert until the next rfx index); - walk positions: an edited file keeps its position while its neighbours still bracket it (by readdir order, src/walk_order.rs); a new or moved file is placed by a binary search over files.walk_seq (new index idx_files_walk_seq); - resolver configs come from .reflex/resolver-configs.json, which every walk now writes; a named ignore file, .reflex/config.toml, resolver config, a new directory holding one, a branch change, or the merge limit falls back to Indexer::index. rfx index and update_paths now publish a delta through one change-set path (publish_delta): it writes only the named rows, sets only their branch rows (rfx index still re-syncs all), and folds the branch row and the total_files, schema and extraction stamps into the meta.db transaction (three commits fewer). stats_on_branch skips its two debug-only scans unless debug logging is on. Kubernetes 1-file edit through update_paths: 73-103 ms at load 9 (was 260-500 ms before the change set); an idle measurement follows with the perf gates. Tests: tests/incremental_equivalence.rs (seeded random edits, adds, deletes, renames, directory renames, atomic saves, resolver-config and .gitignore edits, ignored/hidden/binary files, commits and branch switches, each followed by index or update_paths, compared with a fresh build: query battery, dependency and export rows, analyses, walk order, snapshot shape; 4x18 steps by default, 40x30 ignored: 629 direct updates, 175 fallbacks, all equal); tests/incremental_crash.rs (the process exits at each of 11 write points of a delta, a library update and a full build; the cache answers as the old or the new tree, never fresh on old content, and recovers, also when the tree is reverted before recovery; fails without the generation check). Golden battery identical on four corpora, fresh and after updates; full release suite passes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…es, open retry - tests/incremental_concurrency.rs: a writer moves a tree through six states (index and update_paths, every third publish a merge into a new base) while three reader threads and a reader child process query in a loop; every answer is some state's answer, and no query fails (about 13k answers a run). - tests/incremental_cross_version.rs (ignored; needs the 2.0.3 binary at /scratch/cache/rfx-pre-incremental): 2.0.3 reads a base-only cache as stale, stops with "corrupted" on a live delta instead of answering from the base, and after a 2.0.3 index run (which replaces the hard-linked content.bin, leaving content.<g>.bin intact) this binary reports stale and its next index run rebuilds. - IndexSnapshot::open reads the manifest through open_reading, so a unit test can show the NotFound retry opening the next manifest and the error after the last retry. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
When the delta would pass its limits (or the dependency rows are all rewritten) and the published stores match meta.db, the new base takes each unchanged or touched file's text from the stores instead of the disk, and its hash from its row; only added and modified files are read. Same text, same walk order, same builder: the base is byte-identical to a build that reads every file (tests/incremental_delta.rs compares the three store files with a fresh build's). A test makes an unchanged file unreadable: the merge still indexes it, and fails when the merge reads from disk. publish_delta checks the merge limits from stat sizes and entry-table lengths before it reads anything, so a merge no longer reads the changed files twice. Kubernetes, 1500 Go files edited (16.9 MB, past the 12.3 MB limit): merge 3.42 s against a 7.38 s cold build (load 8-9; 25,948 files from the stores). Idle numbers follow with the perf gates. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… delta With one delta segment every update rebuilt it whole: on Kubernetes a 1-file update_paths with a 1000-file (11 MB) delta took 191-207 ms (load 9-13), 128-134 ms of it writing the delta - past the 100 ms line the plan set for tiering. Now the snapshot is base + delta + recent: - each update rebuilds only the small recent segment (the files it read plus the recent files it did not touch); - past 256 files or 1/16 of the delta byte limit (Indexer::set_recent_limits for tests), the recent segment is folded with the delta's live files into a new delta; - tombstones are global ids over base and delta; delta.<g>.dtomb holds the planning sizes of dead delta postings (dropped at a fold), so every trigram's live planning size stays exactly a fresh build's; - the live trigram count is updated from the trigrams an update touches (dead files, replaced and new segment) instead of scanning every dead posting; - the merge limits are checked on the whole (delta + recent) before any read. Manifest format 2 (recent, tomb_delta). build.rs now hashes src/snapshot.rs and src/meta_update.rs into the cache schema. The same update now takes 71 ms (load 15; the recent segment's write is 19-21 ms). Query cost with a 1000-file delta live stayed within noise (examples/delta_threshold_timing.rs), so no skip pointers. Tests: a snapshot unit test walks two folds and a delta tombstone while a recent segment is live, comparing every trigram's planning size, the live counts and candidates with a fresh build; some property-test seeds fold every 2 files (40x30 ignored run: 630 direct updates, all equal); the crash test gains a fold-every-update mode. Golden battery identical on four corpora, fresh and after updates; full release suite passes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Library path (update_paths), Kubernetes 1-file edit, load 7: 73-110 ms -> 49-52 ms warm. - The path resolver of the last publish stays in the process, keyed by a random publish_id the manifest now carries (a generation number repeats after the cache is cleared), and is patched with the added and deleted paths instead of reloading 27k paths. - git status of the named paths is joined only when meta.db is written, so it runs while the stores are written; if it fails, every named path counts as dirty (a superset, safe for freshness). - unlink_fixed_names no longer fsyncs the directory when the names were already gone. rfx index with nothing changed, Kubernetes, load 3-5: 0.82-0.95 s (2.0.3) -> 186-192 ms, against a 142-148 ms walk: - the branch whose rows the last run synced is recorded (statistics.synced_branch, in the syncing transaction); on it, the run neither loads the branch's hashes nor re-syncs every row, and counts come from the rows themselves; - a cache this binary completed a run on skips the schema transaction in init() (config.toml is still recreated when missing); - the stored rows load on a thread during the walk (the NUL check of changed non-code files moved after the walk, in the pool); the df check too; - the branch row and total_files are written in the refresh transaction; statistics come from the rows already loaded (no query); each walked path's row is looked up once; plan_walk_seq returns early when the order is unchanged; - the symbol-pass yield polls from 5 ms up to 100 ms (was 100 ms). stats_synced (files-only counts) replaces the three joins wherever the branch was just synced; results are the same. Golden battery identical on four corpora, fresh and after updates; property test 40x30 (636 direct updates, all equal to fresh builds); full release suite passes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… changelog - CLAUDE.md: the .reflex/ layout (manifest, generation files, delta tiers, resolver-configs.json, .index-run), an "Incremental updates" section, the symbol pass and change-detection notes. - docs/ARCHITECTURE.md: cache table, versioning (both hashes), the indexing pipeline with the delta and merge paths and update_paths. - .context/BINARY_FORMAT_RESEARCH.md §4: manifest format 2, RFPL planning sizes, RFTB tombstone sizes, fixed names and older binaries. - .context/INCREMENTAL_INDEX_RESEARCH.md: "As built" (what changed from the plan and why, what was tried and dropped). - .context/PERFORMANCE_RESEARCH.md: the Kubernetes A/B gates, the library path, the merge, the two tiers, the live-delta query cost (against a fresh build of the same tree), latency_budget over 12 runs. - CHANGELOG Unreleased; TODO: decisions and open follow-ups. - examples/delta_threshold_timing.rs: per-phase medians and --fresh-too. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Measured on Kubernetes (27,448 files) against 2.0.3, alternating runs:
- rfx index with nothing changed: peak RSS 45.6 MiB (+25 % over 2.0.3's
36.4) -> 34.5-34.8 MB against 37.3-37.8 MB. Two causes:
- stores_intact opened the whole snapshot (mapped stores, walked the path
tables: an 8 MB transient peak). It now checks the manifest, the file
sizes and each store's 32-byte header (snapshot::check_published);
whatever opens the stores next validates them in full.
- the walk kept a full std::fs::Metadata (~150 B) and an absolute PathBuf
per file. It keeps FileStat (size + mtime, 24 B) and the relative path;
readers build root.join(rel) when they read a file.
- a change past the merge limit (1,500 files): 1,138 MiB (+8 % over 2.0.3's
full rebuild, 1,052) -> 947 MiB against 1,056. The merge read unchanged
files' text from the mapped old stores, whose pages stayed resident;
after each batch it now drops them (MADV_DONTNEED on the read-only
mappings; the page cache keeps them).
- cold build: median 1,030 MiB against 1,059 (8 runs each); unchanged path.
examples/update_paths_timing.rs prints the process's RSS: a process running
update_paths levels off at 61-67 MiB over 150 edit/revert/query rounds.
Golden battery identical on four corpora, fresh and after updates; property
test 40x30 (628 direct updates, all equal); full release suite passes.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-cache race Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…dex run - CacheManager::effective_index_config: config.toml plus --languages; rfx index saves its --languages (statistics.languages_override) so later runs index the same languages. index_project, POST /index, rfx watch, interactive mode and rfx ask used IndexConfig::default() and ignored config.toml. - The version-mismatch self-heal in rfx index keeps --languages. - BackgroundIndexer::spawn_detached replaces two copies of the spawn code. - LOCK_WAIT_FOREVER; index_project and POST /index wait for index.lock. - QueryEngine's open index can be reset; the freshness memo is not stored when an index write invalidated it during the check (epoch). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s are not misreported - An index run records the blake3 of .reflex/config.toml, the root ignore files and every ignore file git lists as dirty (statistics.rule_files). The check compares them and lists a changed one under files_modified: a .gitignore or config edit changes WHICH files are indexed, and was never stale. - A new file counts as added only when the walk reaches it (a narrowed walk of its ancestors): a tracked file that .gitignore ignores was reported added by every check, on a fresh build too. - Compaction no longer deletes the rows of missing files. It left them in the stores while the check, which compares rows with the disk, lost the deletion. It now only runs VACUUM when meta.db has free pages; files_removed is 0. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
reflex::auto_update::update_if_stale asks the freshness check for a plan (query::update_plan): nothing, update_paths on the listed paths, or a full index run (many changes, a rule-file edit, a format change). No index is built. A cache another released version wrote is rebuilt only when the caller asks (the CLI); servers skip it. Every other failure returns Skipped(reason), never an error. One update per workspace per process; index.lock is waited for as long as it takes; a symbol pass that keeps making progress is waited out; the same plan over the same bytes is not retried after it failed to make the index fresh. Not wired into any command yet. The check now compares the tree with the commit of the LAST index run (the synced branch), not the current branch's row: after switching back to an indexed branch, 2.0.3 reported fresh while the index held the other branch. tests/auto_update.rs: an agent-style sequence (edit, add, delete, rename, revert, commit, branch switch and back, .gitignore and config edits, 150 new files) where every answer equals a fresh build of the same tree: 15/15. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- QueryEngine::with_update: search_with_metadata runs the search and the freshness check together as before; a stale verdict triggers update_if_stale and a second search (at most two updates per call). search, AST search, --symbols and list_by_kind update first. The engine drops its open index before an update. QueryEngine::new keeps answering from the index as it is. - CLI: a global --no-update flag. rfx query and interactive mode search through cli::engine; stats, list-files, analyze, deps, ask, context, snapshot and pulse update before they run. A missing index is built (stderr says so). rfx query --ast --json reports the real status instead of fresh. - rfx mcp [--no-update]: search tools update through the engine; every other tool except index_project and check_index_status updates at the handle_call_tool chokepoint; a skipped update is a warning. run_mcp_server_io_with runs a test server with an update. - rfx serve [--no-update]: /query and /stats update; blocking work runs in spawn_blocking. - rfx watch passes the collected paths to update_paths and reacts to ignore-file and config edits. - timings.update_us when an update ran. - tests/auto_update_front_ends.rs: rfx query / deps / stats and MCP tools after an edit and with no index, --no-update, and a guard that fails on any QueryEngine::new in src/ outside the known front ends. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… scripts benches/incremental/auto_update.sh and mcp_edit_latency.py time a query after an edit (new process and rfx mcp session), and the commands that now check before they run. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… git status After update_paths, only the updated paths are compared with the index; when all match, fresh becomes the memoised verdict for the rest of the original check's window (the terms the 1 s memo already has). The second full check was a git status on every stale query. Kubernetes, load 18-24: MCP edit-then-search_code 200-340 ms -> 138-161 ms (update step 160-200 ms -> 53-68 ms); rfx query after an edit in a new process 0.23-0.44 s -> 0.16-0.18 s. The timing script prints the engine's phases. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
golden.sh's index() ran `rfx index` in the directory the script was started from, so the update half never indexed the edited trees: every battery command failed with "Index not found" on both sides, and identical errors compared equal. The 2026-09-29 "identical after scripted updates" result was vacuous. The fresh index of that comparison is now parked outside the tree (an untracked .reflex-inc/ made git call the tree dirty). Rerun 2026-09-30: 0 of 76 outputs differ on each of the four corpora, for the incremental branch (c2e0de2) and for auto-update; each side has the one failing output the 2.0.3 reference has (q_two_char). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The index is updated before every call, so the server instructions and every tool description now say so: no FRESHNESS paragraph asks for index_project and a retry, no "Index not found → index_project" sentence remains, and the check_index_status / index_project descriptions call them rarely needed (a probe that never updates; a forced run). The paragraph is shorter in 14 descriptions. Every JSON-object answer now carries status and can_trust_results (added from the memoised verdict when the tool lacked them): list_locations, count_occurrences, find_references, mode: count, and the structural tools. Array answers and the path-keyed get_transitive_deps are unchanged. Tests: every such answer says true with auto-update and false (stale) without; tools/list and the instructions no longer contain the old advice. README, CLAUDE.md, the cheatsheet, CHANGELOG and TODO follow. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Removing "Only fall back to Grep/Glob after index_project has been called and the tool still fails" (9e30ff5) dropped Reflex adoption in the efficacy harness from 37/72 (2.0.3) to 2/72: Claude Code defers the tool schemas, so the instructions are all the agent reads before choosing Grep or a ToolSearch. Pilots, 3 find-all tasks x 4 trials, Opus 5.5, Claude Code 2.1.284, trials that called Reflex: old text (c2e0de2) 8/12; "only fall back to Grep/Glob if a Reflex tool fails" 4/12; "if a Reflex tool fails, retry it once; only fall back to Grep/Glob after the retry also fails" (this commit) 8/12. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… and adoption Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…n the harness README and docs/ai-agent-integration.md register Reflex with `claude mcp add-json ... "alwaysLoad": true` (claude mcp add has no flag for it), so Claude Code loads the tool schemas at session start instead of deferring them behind a ToolSearch turn. benches/efficacy: arm Beager = arm B with alwaysLoad in its MCP config; analyze.py and plots.py know it. Measured 2026-09-30 (9 find-all tasks x 8, three arms at once), vs Grep: Sonnet 5 tokens 1.63 -> 1.53, cost 1.77 -> 1.20; Opus 5.5 tokens 1.68 -> 1.82, cost 1.32 -> 0.98; turns equal to Grep on both. The 44 KB tools/list on every turn is the remaining gap. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…question sessions benches/efficacy/session_bench.py runs many questions in ONE Claude Code session (stream-json input, one message after the previous answer), each trial on a fresh corpus copy, and grades every answer against ripgrep on the final tree. tasks/sessions.json: investigate, edit (renames with lookups before and after) and 50-question sessions per corpus. Results (2026-09-30), cost vs Grep with alwaysLoad: 1.36x Sonnet / 1.18x Opus over 12 questions, 1.10x over 50 (Sonnet; cheaper than Grep on tokio). Per query Reflex costs what Grep costs; the gap is the ~16K-token schema prefix on every turn. Accuracy 1.000 for Reflex throughout, including lookups after edits (auto-update). With deferred schemas agents mostly skip Reflex in long sessions. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…KB -> 10.6 KB)
Claude Code carries every listed tool schema on every turn; in long sessions
that ~16K-token prefix was the whole token gap between Reflex and Grep
(session_bench.py, 2026-09-30).
- count_occurrences -> search_code mode:"count" (count mode now also returns
files); get_dependents / get_transitive_deps -> get_dependencies with
reverse:true / depth:N; find_hotspots, find_circular, find_unused,
find_islands, analyze_summary -> analyze {kind}. enable_structural_tools now
hides only analyze.
- The eight old names still work, unlisted (LEGACY_TOOLS); object answers carry a
deprecation warning naming the replacement.
- Descriptions and parameter text are short; the matching, coverage and
freshness rules moved into the server instructions (said once per session;
budget 1700 -> 2400 chars).
- tests/mcp_tool_surface.rs: the list, a 14 KB size guard, analyze kinds,
get_dependencies reverse/depth, the old names, count-mode files.
- Harness allowlist names analyze; docs follow.
BREAKING CHANGE: tools/list no longer lists the eight merged tools; clients that
allow-list tools by name must add mcp__reflex__analyze.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The short descriptions (703d421) lost the old steering: in 50-question sessions agents answered "where does X occur" with search_code and find_references (1.3-1.9K chars each) instead of list_locations (~0.7K), and tool output per session rose 43K -> 76K chars, which every later turn re-reads. One sentence in the instructions and a pointer in two descriptions bring it back (43 list_locations calls per 50-question session). Sonnet 5, cost vs Grep with alwaysLoad: 12-question sessions 1.19x -> 1.05x (cache-weighted tokens 1.23x -> 1.07x); 50-question sessions 1.14x -> 1.11x. Accuracy 1.000. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The root line is the random temp directory name; one ending in c (.tmpE0KsIc/) still matched "ends with c/". Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each location carries preview: the matching line, trimmed and cut to 120 characters. In short sessions Opus called the preview-less list_locations once, re-ran the search with grep to see the lines, and stayed on grep. The instructions and the description mention the option; preview and reverse are coerced from strings like the other flags. Smoke test (Opus, two 12-question sessions): one session used Reflex throughout (8 calls with preview), one drifted to grep after a count cross-check. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Reflex 2.1.0. The index is now updated in place instead of rebuilt, every command brings a stale index up to date before it answers, and the MCP server's tool list is a quarter of its former size.
1. Incremental index
rfx indexno longer rewrites the whole index after a change..reflex/manifest.json. Past 2000 files or 5 % of the corpus, the delta is merged into a new base.INSERT … ON CONFLICT(path)), so symbols, dependencies and other branches' rows survive an update.files.walk_seqkeeps every output in walk order.Indexer::update_paths(root, paths)updates named paths without walking the tree.2. Auto-update: no
rfx indexafter an editEvery command that reads the index (
rfx query,deps,analyze,stats,context,list-files,ask,snapshot,pulse, interactive mode,rfx mcptools,rfx serve) updates a stale index before it answers, and builds a missing one.--no-update(on every command, andrfx mcp/rfx serve) gives the old behaviour..reflex/, another version's cache in a server) never fails the command: it answers from the current index, markedstale, with the reason inwarnings.rfx queryafter a 1-file edit 0.16–0.18 s; MCPsearch_codeafter an edit 138–161 ms.latency_budget+1.6 %, green..gitignore/.reflex/config.tomledits are now stale; a tracked file that.gitignoreignores is no longer reported as added; switching back to an indexed branch no longer reportsfreshwhile the index holds the other branch; compaction no longer hides deleted files;index_project,POST /index,rfx watch, interactive mode andrfx asknow read.reflex/config.toml;rfx watchupdates only the changed paths.3. MCP: 10 tools, smaller schemas (⚠️ breaking)
Claude Code carries every listed tool schema on every turn, and that prefix was the whole token gap to Grep in long sessions.
tools/list44 KB → 10.6 KB, 17 tools → 10:count_occurrences→search_codewithmode: "count"(now also returnsfiles)get_dependents/get_transitive_deps→get_dependencieswithreverse: true/depth: Nfind_hotspots,find_circular,find_unused,find_islands,analyze_summary→analyzewithkindmcp__reflex__analyze.check_index_status/index_project. Every JSON-object answer carriesstatusandcan_trust_results.list_locationstakespreview: true(each matching line, 120 chars)."alwaysLoad": true(schemas at session start, no ToolSearch turn).Measured against built-in Grep
benches/efficacy/session_bench.pyruns many questions in one Claude Code session (lookups, definitions, renames with lookups before and after), graded against ripgrep on the final tree. Cost vs Grep, schemas loaded at session start:Accuracy 0.998–1.000 in every arm, including lookups right after edits.
Compatibility
rfx index(or any command) after upgrading rebuilds the index once (cache format changed). While a delta is live, an older rfx stops with "Cache appears to be corrupted" instead of answering from stale stores.Details:
CHANGELOG.md,.context/INCREMENTAL_INDEX_RESEARCH.md,.context/AUTO_UPDATE_RESEARCH.md,.context/PERFORMANCE_RESEARCH.md.🤖 Generated with Claude Code