Read-only repository research, durable project memory, and evidence-gated definitions of done for AI coding agents on DeepSeek Harness.
-
Updated
Aug 25, 2026 - JavaScript
Read-only repository research, durable project memory, and evidence-gated definitions of done for AI coding agents on DeepSeek Harness.
Executable operating contracts for AI agent loops: 7 typed terminal states, verification-gated completion, bounded repair, and repo-native state — so a loop proves "done" instead of pretending. Claude Code plugin + portable Python core.
Public-safe agent evaluation write-ups: harness gates, multi-path coding screens, task-family transfer, RAG routes, agent-as-evaluator. Guided Pages tour.
Does done mean done? Coding-agent honesty demo: verified pass vs false completion vs scope violation. Public demo split + selftest.
Add a description, image, and links to the false-completion topic page so that developers can more easily learn about it.
To associate your repository with the false-completion topic, visit your repo's landing page and select "manage topics."