Sitelet https://oh.computer/
Oh
Install Oh
Theme
Appearance

Agent memory framework

Agent memory that shows its work.

Store each fact with the sources it rests on, and keep every change in a replayable history. Answers come back with the evidence attached.

macOS

Terminal
bun add --global @hraness/oh@0.14.1
Apple silicon and Intel · Requires Bun 1.3.14+

Free and MIT licensed · Bun 1.3.14 or newer · No account needed · v0.14.1

$ oh put --kind evidence --key evidence:table-2 \
    --depends-on edition:trial-report-v1 \
    --depends-on assertion:endpoint-12-weeks \
    --value '{"source":"entity:trial-report","locator":"table 2","relationship":"supports"}'
✓ Saved evidence:table-2 (generation 5).
Next: oh get evidence:table-2

$ oh get evidence:table-2
evidence:table-2 (evidence)
{
  "locator": "table 2",
  "relationship": "supports",
  "source": "entity:trial-report"
}
Depends on: assertion:endpoint-12-weeks, edition:trial-report-v1

$ oh verify
✓ Store checked: 5 records and 5 changes replay to the same state (generation 5).

Oh gives agents a memory that cites its sources. Every fact is stored with what it rests on, every change goes into a history you can replay, and every answer returns with its evidence. It runs on your machine as a CLI, an SDK, and an agent skill.

Ask your agent to set it up: oh.computer

Keep the source with the claim.

Save questions, sources, claims, and citations as separate linked records. Your agent can revise a claim while keeping the source it read and the history of how it got there.

Question inquiry
Save what you are trying to find out, along with the investigation that follows.
Source entity
Identify the paper, dataset, person, or system you are researching, even if its title or URL changes.
Capture edition
Record the edition or extract you read, separate from the source as it looks today.
Claim statement
Write down the claim itself, and keep who accepts it and the evidence for it in separate records.
Citation evidence
Point to a passage, table, or observation, and record how it bears on a stance toward a claim, such as support or contradiction.
Artifact view
Build a brief or answer that keeps links to the records it draws on.

Each write joins a history you can replay. Keyword search finds saved records without a model; your application can add semantic search when it needs to find related ideas.

Open the evidence behind an answer.

In the example above, the citation points to table 2 in a saved edition of a trial report and records that it supports a claim. An application can follow those links to show what an answer rests on.

Step 1 of 3. Read the saved review brief: 12 weeks.

Oh checks that records and their history remain intact. It does not decide whether a claim is true.

Your agent proposes. Your app decides.

An agent can remember, look up, explain and propose. Only your application's own code can add a proposal to reviewed knowledge.

Interfaces

Work with the same records from a terminal, TypeScript, or an agent.

The CLI and the TypeScript SDK read and write the same SQLite file. The packaged Agent Skill has a coding agent run the commands you would run yourself, so its changes land in the log you verify.

CLI

Read one record from the local database and space you select.

oh get evidence:table-2 \
  --db research.db \
  --space default

Run the first task

TypeScript SDK

Open the database in your own code and read the same record.

import { Oh } from "@hraness/oh/sdk";

const oh = Oh.open({ databasePath: "research.db" });
try {
  const citation = oh.get("evidence:table-2");
  console.log(citation?.recordSha256);
} finally {
  await oh.close();
}

Read the SDK guide

Agent Skill

Teach a coding agent to check the specification version and replay the log before it reads.

oh contract
oh verify --db research.db --space default
oh get evidence:table-2 \
  --db research.db --space default

Read the Agent Skill

Oh’s semantic search scored 88.87% on LongMemEval-S.

In each comparison, one model answers the same questions from each system’s memory, and every answer is scored the same way. Each result links to its protocol, costs, and limits.

LongMemEval-S, all 500 questions

Share of answers judged correct, averaged over three runs. Higher is better.

LongMemEval-S, all 500 questions, share of answers judged correct, zero to one hundred percent
Lab pipeline on Oh and BM25 retrieval93.07%Every user message plus top-ranked replies within 180,000 bytes, with re-reads chosen by rules
Oh semantic retrieval88.87%Up to 100 turns within 96,000 bytes
BM25 keyword retrieval86.13%Up to 100 turns within 96,000 bytes

Oh semantic retrieval scored 88.87% and BM25 86.13%, and a lab pipeline that adds every message the user wrote to the retrieved replies scored 93.07%. GPT-5 mini answered every question three times with each system, and GPT-4o graded the answers with LongMemEval’s own prompts. On the measure fixed before the run, questions answered correctly in at least two of three runs, Oh’s lead over BM25 is 2.8 points with a 95% interval from 0.0 to 5.6, which does not rule out a tie. The pipeline’s instructions and rules, which are not part of the Oh package, were written after studying all 500 questions, so its score is in-sample.

Other memory systems publish LongMemEval-S scores up to 97%. Each chose its own answering model, judge, prompts, and configuration, so those scores do not compare directly with these.

Results on CloneMem and LoCoMo

On 146 CloneMem questions, Oh’s default SDK search answered 80.59% correctly against 69.86% for Oh semantic retrieval alone, with a median of 10.9 seconds of reranking per search on an Apple M5 Max. GPT-4o mini picked an answer from each question’s options three times, reading the top 10 results within 96 KiB, and a pick counted only if it matched the correct option. The questions come from two personas the project had already studied, so the gain may not carry over to new conversations.

On 861 CloneMem questions from seven personas, an earlier benchmark run of the same reranker answered 77.82% correctly against 70.54% for vector retrieval, a gain of 7.28 points with a 95% interval from 4.61 to 10.27. That run gave the reranker only each candidate’s text, where Oh’s SDK sends the whole record with its key and kind. It used the same answering model, scoring, and 96 KiB reading limit as the 146-question study. The personas were kept out of tuning but had been seen earlier in the project, and two failed attempts were dropped under a retry rule added during the study.

On LoCoMo, filling each question’s context with the nearby turns that best match it found 90.08% of the marked evidence against 88.93% for fixed windows across 1,224 questions, yet answers scored 77.33% against 78.11% on a 300-question sample. Both methods built each context from the same top 20 vector matches within 12,000 bytes, and GPT-4o mini answered each question three times with each method and graded the answers. The answer difference, −0.78 points with a 95% interval from −4.64 to +2.34, does not show either method answering better, and the project had evaluated these conversations before.

All memory studies and protocols

Wordcell uses Oh to query the graph of your Markdown notes.

Use Oh to build memory into an application. Use Wordcell to work with a knowledge base made of Markdown files.

Wordcell keeps notes in Markdown and uses Oh to follow their links. Choose it when you want a knowledge base to use today; choose Oh when you are building memory into your own application.

Wordcell’s search has separate evaluations. Oh’s memory scores do not measure that search.

Install

Install and start with a local database.

Latest release: v0.14.1

  • macOS
  • Linux
  • Windows
oh init
oh put --kind entity --key entity:ada-lovelace \
  --value '{"name":"Ada Lovelace","role":"mathematician"}'
oh get entity:ada-lovelace
oh search "mathematician"
oh verify

The CLI needs Bun 1.3.14 or newer. The first task creates one entity, reads it back, finds it with keyword search, and verifies the log. Oh writes to .oh/oh.sqlite and the default space unless you pass --db or --space. Read the full first run on GitHub.

Control

Keep your memory on your machine.

Start with a local SQLite file, without an account. Connect remote services only when your application needs them.

Local by default
Your records, the log of every change, and the keyword index live in a SQLite file you choose. Semantic search caches are derived from the records and can be rebuilt.
Remote services are opt-in
Hosted embeddings, network sync, and a remote libSQL database are used only when you configure them. Sync sends operations, never search vectors.

Questions

What to know before you install.

What is stored, and where?

One SQLite file holds your records and their digests, the append-only log of every change, a keyword index built from the records, and the specification version the file follows. Oh uses .oh/oh.sqlite and the default space unless you name another path or space. Semantic caches and remote copies exist only where you configure them.

Is semantic search required?

No. The CLI uses keyword search without a model. The SDK can add local semantic search through QMD, or hosted embeddings from Cloudflare Workers AI. Hosted embeddings send record text to Cloudflare. Oh drops results whose records have changed or been removed since indexing.

Does a passing verification mean a claim is true?

No. A passing verification means the records and their history are intact: replaying the log reproduced every digest. Search scores measure relevance, not truth. Whether a claim holds is recorded separately, in assertions and review decisions linked to their evidence.

Where can I run it?

The CLI, the local SDK, and the SQLite store need Bun 1.3.14 or newer. The runtime-neutral store interfaces and the direct libSQL adapter also run on Node 24, including in serverless functions.