Repository navigation
test: differential testing against the legacy ledger machine - #195
Conversation
ea0bfdc to
5f539ad
Compare
5f539ad to
d0e34a7
Compare
Runs generated programs through both the numscript interpreter and a vendored copy of the ledger repo's `internal/machine`, and compares what moved. - `internal/oracle/` is that vendored copy, in its own module so it cannot drift into the interpreter's dependency graph. `smoke_test.go` checks the vendoring itself did not change behaviour. `DIVERGENCES.md` records every difference from ledger main and why it is there; `DIFFERENCES-BY-EXAMPLE.md` is the same list as runnable scripts. - `internal/difftest/` is the harness, also its own module: `Compare` (what counts as a real difference vs a posting-granularity one), the seeded `TestDifferentialSweep` gate, `TestKnownBugRepros`/`TestKnownOpenDivergences`, and `FuzzDiff` with its corpus. CI gets a `Difftest` job, because both are separate modules and the root `go test ./...` reaches neither. It runs the deterministic sweep, not the fuzzer: `-fuzztime` expiry surfaces as a failure with no divergence behind it. The sweep is green -- 0 divergence classes and 0 tolerated over 3000 seeds. Getting there meant changing the oracle's `kept` to consume its funding, the way the interpreter does and unlike ledger. That is a deliberate loss of independence on `kept` and is written up as DIVERGENCES.md §2 ②.
d0e34a7 to
734d4f6
Compare
Drops the running commentary about the generator's Haskell ancestry (Gen.hs/Numscript.hs/Utils.hs, QuickCheck's `frequency`, `Ratio Int`) — none of it is a fact about this code — along with comments that restate the field they sit on and rationale that had grown into prose. Kept: what cannot be read off the code — why `max`-clause amounts are never negative, why both pools are sorted before being indexed, why the var-collision shapes are biased for. `GenerateScriptAST`, `ToBuilder` and `ToBuilderScript` were exported but used only inside the package; `GenerateScript` and `RandFromBytes` are the entry points. Unexported. Also removes references to DIFFTEST_HANDOFF.md, which is not in the repo, and deletes a superseded duplicate doc comment on genPresetMetadata. internal/gen landed in #194 before this cleanup was ready, hence the separate commit.
DIVERGENCES.md and DIFFERENCES-BY-EXAMPLE.md covered the same four differences at two levels of detail, and the second was written on top of the first rather than replacing it, so both had to be kept in sync by hand. One file now. 472 lines down to 244. Each difference gets a runnable script, what all three engines do with it, why, and the open question. The vendoring mechanics and the drift check move to the end, where they belong. Section numbers become #1-#4, which is what the code comments now cite.
49d470a to
541deb1
Compare
…stead The oracle carried a negative-amount guard on OP_TAKE that ledger does not have. That made vm/machine.go a behaviour fork, and hid the cost: the oracle silently stopped being ledger on every script with a negative send amount. Copy ledger exactly instead, and absorb the difference in Compare under a named tolerance. Both engines reject those scripts and no money moves either way; only the error differs (`insufficient funds` vs `Cannot send negative amount`). The tolerance is one-directional: the interpreter naming a negative amount while the oracle blames the funds. The reverse is DIVERGENCES.md #4, where the oracle rejects and the interpreter silently zeroes the clause and runs short, and that must stay a mismatch. TestMissingFundsClassificationMismatchStillCaught caught the symmetric version of this and is why it is not. The sweep prints how often it fires -- 161 of 3000 scripts -- so the cost is countable on every run instead of invisible. vm/machine.go is now byte-identical to ledger apart from the vendored types, leaving `kept` as the oracle's only behavioural divergence. Also records in DIVERGENCES.md #4 why ledger PR #2079 was closed: clamping a negative `max` needs a new opcode, and adding one to the machine is too aggressive for what it buys.
…rements The doc said the generator rarely produces a save-overdraw followed by a bounded-overdraft draw on the same account. Measured over the same 3000 seeds it produces the ingredients constantly and the combination never: 686 saves, 2284 bounded overdrafts, 130 on a shared account, 7 in order, 2 that also overdraw, 0 complete. The reason is structural, so say it: statements are sampled independently, and a shape needing two of them to hit the same (account, asset) in order is quadratically suppressed.
|
internal/gen/gen.go:1 🔵 [suggestion] internal/gen edits are unmentioned in the PR description The diff rewrites comments across internal/gen and unexports internal/gen/ast.go:168 🟡 [minor] VarDecl.AccountAsVar comment lost its placeholder and is now garbled The reworded comment reads |
61b1ca0 to
6b3fc6a
Compare
|
numscript.go:83 🟡 [minor] Exported The description states "No interpreter, API or storage change", yet the diff adds the exported alias |
… negative-max gap Compare tolerated any one-sided runtime failure once both sides agreed on missing-funds classification, not just the known negative `max` clause divergence (DIVERGENCES.md #4). A real bug where one engine committed money the other refused to move, for an unrelated reason, would have passed through silently. Add SideResult.NegativeMaxReject, set by run_oracle.go when the failure is ledger's OP_TAKE_MAX guard rejecting a negative max clause (untyped upstream, so matched by its fixed error text rather than by type), and require it before tolerating the asymmetry. Any other one-sided failure is now a mismatch. Renamed the tolerance from "one-sided runtime failure" to "negative max clause" to match, and updated DIVERGENCES.md and the tests that asserted the old name.
What changed
Runs generated programs through both the numscript interpreter and a vendored
copy of the ledger repo's
internal/machine, and compares what moved.internal/oracle/— that vendored copy, in its own module so it cannotdrift into the interpreter's dependency graph.
smoke_test.gochecks thevendoring itself did not change behaviour.
DIVERGENCES.mdrecords everydifference from ledger main and why it is there.
internal/difftest/— the harness, also its own module:Compare(whatcounts as a real difference vs a posting-granularity one), the seeded
TestDifferentialSweepgate,TestKnownBugRepros/TestKnownOpenDivergences,and
FuzzDiffwith its corpus.Difftestjob. Both are separate modules, so the rootgo test ./...reaches neither and without this the harness would sit in therepo and never run. It runs the deterministic sweep, not the fuzzer:
-fuzztimeexpiry surfaces as a test failure with no divergence behind it,which belongs in a scheduled job, not a PR gate.
Why
Third of three PRs replacing #191. Stacked on #194 (generator) → #193 (builder).
It exists so the funds-engine swap (#190) has an independent baseline rather
than being reviewed on assertion alone.
Base is
feat/program-generator— retargets as the stack merges.Risk
MEDIUM, and not for the usual reason: nothing here ships. The risk is that
the harness looks more authoritative than it is. Two things bound it, both
written down rather than discovered later:
colors. A green sweep says nothing about them.
ledger main, in
DIVERGENCES.md— theOP_TAKEnegative-amount guard(Setup sentry #1; upstream agrees it is a bug, PR #2060 closed unmerged) and
keptconsuming its funding (Better cli #2).
Validation
go build ./...,go vet ./..., gofmt clean in allthree modules
go test ./...ininternal/oracleandinternal/difftestTestDifferentialSweepgreen —0 divergence classes over 3000 seeds (372 b-side compile rejections,
30 b-side resolve rejections, 106 tolerated under the
negative-amount-vs-missing-funds tolerance)
Architecture / behavior impact
Two new Go modules, neither reachable from the root module's build. One new CI
job. No interpreter, API or storage change.
Review focus
DIVERGENCES.md#2 is where judgment is most valuable. The sweep wentgreen partly by changing the oracle: ledger does not consume a
keptfunding —it returns to the pool and funds the next destination — while the interpreter
does. The oracle now consumes it, matching the interpreter, at two sites in
script/compiler/destination.go.That removed 7 failing and 85 tolerated scripts and let
Compare'skept source attributiontolerance be deleted outright — a real gain, since itwould have swallowed a genuinely wrong source in any script mentioning
keptand short-circuited before metadata was compared.
The cost is that the oracle can no longer catch a
keptregression in theinterpreter.
TestOracleKeptComplexkeeps ledger's original expectations inits comment as the compensating control. Worth disagreeing with if you think
the oracle should have stayed faithful and the sweep stayed red.
Second:
Compare's remaining relaxations. Each one is checking power traded forsignal, and they are the difference between a harness that finds bugs and one
that agrees with itself.
Known concerns
TestKnownOpenDivergencespins one numscript↔ledger difference that is realand undecided — negative bounded overdraft cap (#5). It asserts the divergence
is still there, so it fails loudly if anyone changes the semantics without
updating the doc. Not an oracle defect and doesn't block this PR.
Two other real, undecided differences exist and are pinned elsewhere:
savepast the balance (#3,
TestKnownBugRepros) and source-side negativemax(#4,TestSourceSideNegativeMaxClauseTolerated/TestMissingFundsClassificationMismatchStillCaught).