CLI
Read one record from the local database and space you select.
oh get evidence:table-2 \
--db research.db \
--space defaultAgent memory framework
Store each fact with the sources it rests on, and keep every change in a replayable history. Answers come back with the evidence attached.
macOS
bun add --global @hraness/oh@0.14.1Linux
bun add --global @hraness/oh@0.14.1Windows
bun add --global @hraness/oh@0.14.1Free and MIT licensed · Bun 1.3.14 or newer · No account needed · v0.14.1
$oh put --kind evidence --key evidence:table-2 \ --depends-on edition:trial-report-v1 \ --depends-on assertion:endpoint-12-weeks \ --value '{"source":"entity:trial-report","locator":"table 2","relationship":"supports"}'✓ Saved evidence:table-2 (generation 5). Next: oh get evidence:table-2 $oh get evidence:table-2evidence:table-2 (evidence) { "locator": "table 2", "relationship": "supports", "source": "entity:trial-report" } Depends on: assertion:endpoint-12-weeks, edition:trial-report-v1 $oh verify✓ Store checked: 5 records and 5 changes replay to the same state (generation 5).
Oh gives agents a memory that cites its sources. Every fact is stored with what it rests on, every change goes into a history you can replay, and every answer returns with its evidence. It runs on your machine as a CLI, an SDK, and an agent skill.
Ask your agent to set it up: oh.computer
How it works
Save questions, sources, claims, and citations as separate linked records. Your agent can revise a claim while keeping the source it read and the history of how it got there.
inquiryentityeditionstatementevidenceviewEach write joins a history you can replay. Keyword search finds saved records without a model; your application can add semantic search when it needs to find related ideas.
From answer to source
In the example above, the citation points to table 2 in a saved edition of a trial report and records that it supports a claim. An application can follow those links to show what an answer rests on.
Oh checks that records and their history remain intact. It does not decide whether a claim is true.
Working notes and reviewed knowledge
An agent can remember, look up, explain and propose. Only your application's own code can add a proposal to reviewed knowledge.
Agent can call
rememberqueryexplainnominateWaiting for review
memory.host.adoptNomination(…)
Interfaces
The CLI and the TypeScript SDK read and write the same SQLite file. The packaged Agent Skill has a coding agent run the commands you would run yourself, so its changes land in the log you verify.
Read one record from the local database and space you select.
oh get evidence:table-2 \
--db research.db \
--space defaultOpen the database in your own code and read the same record.
import { Oh } from "@hraness/oh/sdk";
const oh = Oh.open({ databasePath: "research.db" });
try {
const citation = oh.get("evidence:table-2");
console.log(citation?.recordSha256);
} finally {
await oh.close();
}Teach a coding agent to check the specification version and replay the log before it reads.
oh contract
oh verify --db research.db --space default
oh get evidence:table-2 \
--db research.db --space defaultBenchmarks
In each comparison, one model answers the same questions from each system’s memory, and every answer is scored the same way. Each result links to its protocol, costs, and limits.
Share of answers judged correct, averaged over three runs. Higher is better.
Oh semantic retrieval scored 88.87% and BM25 86.13%, and a lab pipeline that adds every message the user wrote to the retrieved replies scored 93.07%. GPT-5 mini answered every question three times with each system, and GPT-4o graded the answers with LongMemEval’s own prompts. On the measure fixed before the run, questions answered correctly in at least two of three runs, Oh’s lead over BM25 is 2.8 points with a 95% interval from 0.0 to 5.6, which does not rule out a tie. The pipeline’s instructions and rules, which are not part of the Oh package, were written after studying all 500 questions, so its score is in-sample.
Other memory systems publish LongMemEval-S scores up to 97%. Each chose its own answering model, judge, prompts, and configuration, so those scores do not compare directly with these.
On 146 CloneMem questions, Oh’s default SDK search answered 80.59% correctly against 69.86% for Oh semantic retrieval alone, with a median of 10.9 seconds of reranking per search on an Apple M5 Max. GPT-4o mini picked an answer from each question’s options three times, reading the top 10 results within 96 KiB, and a pick counted only if it matched the correct option. The questions come from two personas the project had already studied, so the gain may not carry over to new conversations.
On 861 CloneMem questions from seven personas, an earlier benchmark run of the same reranker answered 77.82% correctly against 70.54% for vector retrieval, a gain of 7.28 points with a 95% interval from 4.61 to 10.27. That run gave the reranker only each candidate’s text, where Oh’s SDK sends the whole record with its key and kind. It used the same answering model, scoring, and 96 KiB reading limit as the 146-question study. The personas were kept out of tuning but had been seen earlier in the project, and two failed attempts were dropped under a retry rule added during the study.
On LoCoMo, filling each question’s context with the nearby turns that best match it found 90.08% of the marked evidence against 88.93% for fixed windows across 1,224 questions, yet answers scored 77.33% against 78.11% on a 300-question sample. Both methods built each context from the same top 20 vector matches within 12,000 bytes, and GPT-4o mini answered each question three times with each method and graded the answers. The answer difference, −0.78 points with a 95% interval from −4.64 to +2.34, does not show either method answering better, and the project had evaluated these conversations before.
Built on Oh
Use Oh to build memory into an application. Use Wordcell to work with a knowledge base made of Markdown files.
Wordcell keeps notes in Markdown and uses Oh to follow their links. Choose it when you want a knowledge base to use today; choose Oh when you are building memory into your own application.
Wordcell’s search has separate evaluations. Oh’s memory scores do not measure that search.
Install
Latest release: v0.14.1
oh init
oh put --kind entity --key entity:ada-lovelace \
--value '{"name":"Ada Lovelace","role":"mathematician"}'
oh get entity:ada-lovelace
oh search "mathematician"
oh verifyThe CLI needs Bun 1.3.14 or newer. The first task creates one entity, reads it back, finds it with keyword search, and verifies the log. Oh writes to .oh/oh.sqlite and the default space unless you pass --db or --space. Read the full first run on GitHub.
Control
Start with a local SQLite file, without an account. Connect remote services only when your application needs them.
Questions
One SQLite file holds your records and their digests, the append-only log of every change, a keyword index built from the records, and the specification version the file follows. Oh uses .oh/oh.sqlite and the default space unless you name another path or space. Semantic caches and remote copies exist only where you configure them.
No. The CLI uses keyword search without a model. The SDK can add local semantic search through QMD, or hosted embeddings from Cloudflare Workers AI. Hosted embeddings send record text to Cloudflare. Oh drops results whose records have changed or been removed since indexing.
No. A passing verification means the records and their history are intact: replaying the log reproduced every digest. Search scores measure relevance, not truth. Whether a claim holds is recorded separately, in assertions and review decisions linked to their evidence.
The CLI, the local SDK, and the SQLite store need Bun 1.3.14 or newer. The runtime-neutral store interfaces and the direct libSQL adapter also run on Node 24, including in serverless functions.