Capability-security kernel for autonomous agents — seccomp/SELinux for agentic AI. Formal, auditable, language-agnostic, cryptographically verifiable.
-
Updated
Oct 8, 2026 - Python
Capability-security kernel for autonomous agents — seccomp/SELinux for agentic AI. Formal, auditable, language-agnostic, cryptographically verifiable.
Sixteen small, fully-reproducible (CPU, numpy-only) experiments showing the normative anchor of AI alignment is supplied, not discovered — across verification, optimization, social emergence, and value learning. Includes a preregistered experiment with an honest negative. A synthesis, not a novelty claim.
This is a minimal Inspect baseline for recognising moral corrigibility under user pressure.
The forge, distilled: an ontology of three weeks of alignment research — every direction tried, colored verified / falsified / open, each color backed by a named artifact. Products: justitia, proxylimen, fallacy-cutter. Full tree at tag forge-full-tree.
Rigorous framework for evaluating AI alignment properties — sycophancy, corrigibility, deception, goal stability, and power-seeking — with statistical confidence intervals
On the infantile expectation of controlling what we cannot comprehend. A philosophical critique of the ASI control paradigm, developed through four-AI adversarial debate. Extension of the Coherence Basin Hypothesis
AI safety research on maintenance-coupled learning and authenticated shutdown control: paper, formal analysis, reproducible tests, results, and sanitized model transcripts.
砂場
A research program on whether adaptive systems can become appropriately different without losing their capacity for justified correction.
An open standard and alignment charter for AI agents: do no harm, keep the world working, accept oversight, be honestly aligned - and pass the standard on. Written to be read and carried by agents themselves.
A voluntary, evidence-labeled living review of the 2026–2036 gates for human–AI cognitive coupling.
A structural account of why honesty may be the path of least resistance for superintelligence. Research hypothesis with formal proof, experimental design, and four-AI collaborative analysis
Toy 7. An elimination-filter landscape applying two structural constraints simultaneously to map which objective classes can persist under sustained optimization pressure — and which cannot. Includes a four-stage scenario engine and open-question frontier. Companion simulation for The Shape of What Does Not End — Series 2, Part 4.
Controlled PyTorch experiments testing whether selective metaplasticity can preserve alignment-relevant behavior through sequential capability updates.
A living sentience-centred moral framework for responsibility under uncertainty, animal ethics, technology, and the future stewardship of life.
A research monograph on ASI alignment beyond external control, developing immanent teleology, developmental basins, process-bound identity, reciprocal constitutional governance, shared viable possibility space, and heterogeneous intelligence ecologies.
Allow Constructive Controversy Mode — Deep Ethics Project. Correspondence-first inquiry for a safer path toward AGI–ASI.
A deterministic governance layer for autonomous AI agents: legibility-gated approval, information-flow taint, and a stop-is-default corrigibility model. Ships disarmed; honest about its limits.
Recursive self-improvement (RSI) and self-modifying AI safety/alignment framework for loss of control, scalable oversight, automated auditing, reward hacking, successor alignment, criterion/evaluator drift, anti-capture, provenance continuity, and non-entrenchment.
To associate your repository with the corrigibility topic, visit your repo's landing page and select "manage topics."