A local CLI for logging predictions before the outcome and checking whether your "80% sure" comes true 80% of the time.
-
Updated
Oct 5, 2026 - Rust
A local CLI for logging predictions before the outcome and checking whether your "80% sure" comes true 80% of the time.
Runs ~20 options strategies against live market data in shadow mode, records every hypothetical fill under worst/base/optimistic assumptions, and grades each with anytime-valid e-processes. Places no orders.
Repository for the paper "Auditing Pay-Per-Token in Large Language Models", AISTATS'26
This repository contains the code for the paper "Optimizing Social Utility in Sequential Experiments".
A kernel-userland protocol enforcing information-theoretic bounds on AI adaptivity leakage, benchmark gaming, and capability spillover.
Measure your agent harness, find where it wastes the model, and prove the fix worked. Harness-agnostic, agent-agnostic, zero dependencies. Reference implementation of HTP-1.
Publicly verifiable proof that an LLM endpoint actually spent the compute you paid for — zero provider cooperation. Closes the effort gap (Hollow-LLM) and the transferability gap (IRIS).
Benchmark for statistically valid AI scientist systems, using audit-closed protocols, transparency logs, and sequential inference to prevent false discoveries in autonomous research agents.
Anonymous code and result release for block-calibrated false-discovery control in open-vocabulary detection
Codebase of Concise and Logically Consistent Conformal Sets for Neuro-Symbolic Concept-Based Models, accepted at NeurIPS2025
MSc thesis: designing e-value evidence around a decision-maker's loss, with an application to Polymarket prices for the 2026 World Cup final.
Windows instrument for fine-tuning runs: every sample drawn, a report that leads with its answer, and a local-model workbench that builds formula tools and tests what each knob does, with e-value evidence across folders.
Research implementation of anytime-valid group-invariance tests with kernel, adaptive Gaussian, and neural statistics.
Label-free drift attribution (model / noise / world / annotator) with anytime-valid guarantees and a frozen pre-registration — closes the zero-label misattribution gap 0.50→0.00 on the recoverable cause; negatives reported as measured boundaries.
Reproducibility artifact for The Price of Peeking: Anytime-Valid Leakage Detection on ML-KEM EM Traces.
To associate your repository with the e-values topic, visit your repo's landing page and select "manage topics."