SourceBraid — Weave the web into Markdown.
SourceBraid saves articles, research papers, wiki pages, GitHub Gists, and PDF documents as durable Markdown in a private GitHub repository. Metadata lives in YAML frontmatter, relevant images become local repository assets, and every source is added to a searchable index.
SourceBraid does more than bookmark URLs. It prepares each source using the richest trustworthy representation available, keeps its provenance, and leaves you with ordinary files and Git history that remain useful without SourceBraid.
SourceBraid combines capture clients, a private GitHub repository as the durable source of truth, and two archive-access paths. The local Codex plugin supports search and archive management. The hosted, read-only ChatGPT integration is prepared for deployment and review; it is not yet a live public service. The Chrome extension and iOS app write captures directly to the configured GitHub repository. Hosted ChatGPT requests will pass through SourceBraid's Cloudflare Worker without creating a permanent content database.
The local SQLite index is only a rebuildable search cache. Markdown files and Git history remain authoritative.
flowchart TD
A["Web page, wiki, Gist, arXiv paper, or PDF"] --> B{"Capture client"}
B -->|Chrome| C["Browser extension"]
B -->|iOS| D["App and Share Extension"]
C --> E["Select the best extraction adapter"]
D --> E
E --> F["Normalize content, create frontmatter, and save images"]
F --> G{"PDF conversion required?"}
G -->|No| H["Store Markdown, assets, and URL-hash metadata shard"]
G -->|Yes| I["Store PDF, placeholder, and metadata"]
I --> J["GitHub Action converts the PDF with Docling"]
J --> H
H --> K["Private GitHub repository as source of truth"]
K --> L{"Local search index available?"}
L -->|No| M["One-time index build"]
L -->|Yes| N["Compare remote head and Git blob SHAs"]
N -->|Changed| O["Download only new or changed files"]
N -->|Unchanged| P["Use the existing index"]
M --> Q["SQLite index with FTS5"]
O --> Q
P --> Q
Q --> R["Local Codex: search, fetch, list, or safely delete"]
K --> S["Prepared hosted MCP: read-only GitHub access"]
S --> T["ChatGPT after deployment and review"]
The workflow in detail:
- Capture: Open SourceBraid on the current page or share content from iOS. Add tags and personal notes before saving.
- Extract: SourceBraid selects the strongest available adapter. Structured sources such as arXiv, Azure DevOps, Gists, and native Markdown take priority over generic DOM extraction.
- Prepare: Content becomes portable Markdown. SourceBraid adds YAML frontmatter, resolves relative links, and stores relevant images next to the document so the clip remains readable without the original page.
- Commit: Documents, assets, and metadata are written through the GitHub Contents API. Metadata is partitioned into up to 256 JSONL shards by URL hash. Normal Git commits make every change inspectable and recoverable.
- Finish PDFs: When no suitable HTML representation exists, the original PDF remains in the repository. A GitHub Action uses Docling to create the final Markdown, extract figures, and replace the pending placeholder.
- Index locally: On first use, the Codex plugin builds a local SQLite FTS5 index. Later updates compare the stored commit and Git blob SHAs, processing only new, changed, or deleted files.
- Use: Local Codex searches the index and supports guarded deletion with a preview and exact confirmation. Its last synchronized index remains readable when GitHub is unavailable. The prepared ChatGPT service searches and fetches through GitHub with read-only access, explicit coverage limits, and no local index or deletion tools.
Keeping the GitHub archive separate from the local search cache matters for large collections: a normal query does not need to reopen thousands of Markdown files. A full pass is needed only for the first build, an explicit rebuild, or index repair.
| Source or format | Preferred extraction | Markdown result | Images and attachments | Fallback |
|---|---|---|---|---|
| arXiv paper | Experimental full-paper arXiv HTML | Sections, prose, tables, citations, and LaTeX formulas; authors, arXiv ID/version, DOI, categories, and journal reference in frontmatter | Figures are copied into the asset folder and linked relatively | Download the PDF and convert it with Docling |
| Remote or local PDF | Original PDF plus asynchronous Docling workflow in GitHub Actions | Reading order, tables, OCR text, and referenced figures; starts as pending, then becomes finished Markdown |
The original remains as source.pdf; extracted figures sit beside it |
Local PDFs require Chrome's Allow access to file URLs setting; encrypted or session-only PDFs are unsupported |
| Azure DevOps Wiki | Authenticated Wiki REST API returns source Markdown | Azure macros are normalized, Mermaid remains a mermaid code block, internal wiki links become absolute |
Protected attachments are loaded through the still-authenticated source tab | Rendered .markdown-content area |
| GitHub Gist | GitHub Gist API, with the configured token for private Gists | A single Markdown file directly; multiple files as sections; source code in language-tagged fences | Public images directly, protected GitHub images through the signed-in Gist tab | Revision-specific URLs keep their revision |
| Native Markdown | HTTP response to Accept: text/markdown, for example from Hashnode or appropriately configured Cloudflare sites |
Source frontmatter and duplicate H1 removed; relative links made absolute | Relevant images are stored locally and linked relatively | Continue through dedicated APIs, then DOM extraction |
| WordPress | WordPress REST endpoint discovered from page metadata | Article content converted from structured API data | Relevant article images stored locally | Visible page content |
| Forem / DEV | Forem API with source Markdown | Normalized Markdown without site chrome | Relevant images stored locally | Visible page content |
| Ghost | Configured Ghost Content API | Structured post content with canonical URL validation | Relevant images stored locally | Visible page content |
| Blogger | Blogger API using detected blog and post IDs | Structured article content | Relevant images stored locally | Visible page content |
| Google DeepMind blog | Article sections from the page DOM | Full post without the cover controls or related-post cards | Article images stored locally | Generic visible-page extraction |
| JSON Feed, RSS, or Atom | Feed announced by the HTML page | Full feed content when available | Relevant images stored locally | Visible page content |
| Generic HTML page | Visible DOM, preferring article, main, or [role="main"] |
Headings, paragraphs, links, lists, quotes, code, and tables | Content-relevant images stored locally | body as the final fallback |
SourceBraid always uses the strongest available content source. For HTML pages, adapters run in this order:
- arXiv HTML
- Azure DevOps Wiki
- GitHub Gist
- Native Markdown
- WordPress REST
- Forem / DEV API
- Ghost Content API
- Blogger API
- Google DeepMind blog DOM
- JSON Feed, RSS, or Atom
- Visible DOM
The first matching, validated source wins. SourceBraid then normalizes the Markdown, downloads images, writes YAML frontmatter, and updates the index.
Markdown files are stored through the GitHub Contents API:
web-clips/YYYY/MM/YYYY-MM-DD-domain-title-urlhash.md
Related assets live under:
web-clips/YYYY/MM/assets/YYYY-MM-DD-domain-title-urlhash/
Markdown image references are relative to this asset directory. PDF sources
also retain the original as source.pdf.
SourceBraid maintains a URL-hash-sharded metadata index:
web-clips/index/00.jsonl
...
web-clips/index/ff.jsonl
The same URL always maps to the same shard, avoiding a rewrite of the entire
metadata collection on each capture. Existing archives with
web-clips/index.jsonl remain compatible and can be migrated atomically through
the plugin. Each entry includes title, canonical URL, repository path, capture
date, optional publication and modification dates, tags, source type,
extraction method, capture timestamp, and saved image paths.
date and the YYYY/MM path use the local capture date; a source's publication
date remains separate in published.
An arXiv abstract page such as https://arxiv.org/abs/2311.02462 can be saved
directly. SourceBraid prefers the experimental HTML version of the full paper,
converts it to Markdown, and keeps research metadata. The PDF does not need to
be downloaded or opened manually.
If no HTML version exists, the extension uploads the PDF in the background and the Docling workflow converts it automatically.
SourceBraid accepts remote HTTP(S) PDFs and local .pdf files opened in Chrome.
For local files, enable Allow access to file URLs in SourceBraid's extension
details at chrome://extensions. Without that permission, the extension shows
a concrete instruction instead of producing an empty HTML clip.
A PDF capture initially stores:
web-clips/YYYY/MM/assets/CLIP-SLUG/source.pdf
The extension creates a pending Markdown entry and metadata record. The final
PDF commit starts .github/workflows/convert-pdfs.yml, which:
- installs Docling on a GitHub runner;
- extracts reading order, tables, OCR text, and figures;
- replaces the pending Markdown while preserving notes and frontmatter;
- marks the matching metadata entry as complete; and
- retains the original PDF beside extracted assets.
GitHub Actions needs write access to repository contents. The workflow has a 45-minute timeout, and individual PDFs are limited to 25 MB by browser and GitHub API constraints. Rerun a conversion under Actions → Convert PDFs to Markdown → Run workflow.
If another clip is saved to the same branch during conversion, the workflow refreshes its branch and retries a rejected push up to five times. Concurrent SourceBraid uploads are therefore not lost to a temporary Git ref race.
SourceBraid reads source Markdown through the authenticated Azure DevOps Wiki
API. If that fails, it converts only the rendered .markdown-content area —
not navigation, headers, or unrelated Azure DevOps UI.
Protected attachment URLs may need the browser's signed-in session, so SourceBraid loads images sequentially through the open source tab, stores them in the asset folder, and rewrites links to relative repository paths. Keep the source tab open until capture completes. The frontmatter records the organization, project, wiki ID, page ID, page path, and revision when available.
A one-file Markdown Gist becomes the document body directly. Multi-file Gists become one document with a section per filename; non-Markdown files remain in language-tagged code fences.
Public Gists work anonymously. For private Gists, SourceBraid uses the configured GitHub token when it has Gist read permission. Signed-in GitHub image assets can be loaded through the still-open Gist tab.
Until the Chrome Web Store listing is live, load the reviewed extension from this checkout:
- Open
chrome://extensions. - Enable Developer mode.
- Select Load unpacked.
- Choose
chrome-extension/sourcebraid. - Open a supported source and select the SourceBraid icon.
- Configure the private GitHub repository, optionally add tags or notes, and choose Save to GitHub.
After setup, GitHub settings stay collapsed behind the settings icon. If only
the GitHub upload fails after successful extraction, the popup offers a
Download Fallback. Before an upload starts, SourceBraid verifies that the
configured repository exists and is accessible to the token; the popup reports
an explicit error when GitHub returns 404 Not Found.
Before the first capture, SourceBraid shows what page data is read and where it goes, and requires affirmative consent. Export Plugin Config never includes the GitHub token or optional source API credentials.
With the GitHub CLI installed and authenticated (gh auth login), one command
creates or initializes the private archive for the authenticated account:
python3 scripts/setup_github.pyThe default target is AUTHENTICATED_USER/sourcebraid-private. The script
refuses public repositories, preserves existing files, enables GitHub Actions,
uploads only the allowlisted PDF support files, and writes a token-free
sourcebraid-config.json. Use --repo OWNER/NAME for another private target or
--dry-run to preview the operation.
Release builds also provide the same workflow as one standalone file. Build it locally with:
python3 scripts/build_setup_package.py
python3 dist/sourcebraid-github-setup-v1.0.1.py --helpUse a fine-grained personal access token restricted to exactly one private repository:
Contents: Read and write
Workflows: Read and write
Workflows is needed only for PDF support. On the first PDF upload, SourceBraid
installs the bundled Docling workflow, conversion script, and requirements file
when those paths do not already exist. Existing files are never overwritten.
The token is stored locally in Chrome extension storage.
Optional API configuration:
- Ghost Content API: base URL, such as
https://example.com/ghost/api/content, plus a browser-safe Content API key - Blogger: optional Google API key; public posts do not require OAuth, but anonymous API calls normally need a key for quota
The primary ChatGPT release uses the hosted MCP service in chatgpt-mcp/.
Its intended endpoint is https://mcp.sourcebraid.com/mcp; deployment, live
authentication tests, and OpenAI review/publication are still required. Follow
the deployment guide and
OpenAI MCP submission kit.
After the service is deployed, users connect through SourceBraid's HTTPS consent page using a fine-grained GitHub token with Contents: Read-only for exactly one private archive repository. Never paste that token into a ChatGPT prompt. The hosted tools provide search, fetch, listing, and status; they cannot capture, edit, delete, migrate metadata, or build a local index. Default-branch search uses GitHub Code Search and verifies matches at a pinned commit. Indexing delays, limits, and any bounded fallback are reported with the results.
The local Codex plugin remains the full archive-management option. It needs Python 3 and local GitHub authentication; a skills-only ZIP does not by itself give ordinary ChatGPT access to those local files or credentials.
For local development, add this checkout as a marketplace and install the plugin:
codex plugin marketplace add .
codex plugin add sourcebraid@sourcebraidAfter the reviewed v1.0.1 tag is published, the tag-bound GitHub installation
is:
codex plugin marketplace add patrickschiller/sourcebraid \
--ref v1.0.1 \
--sparse .agents/plugins \
--sparse codex-plugin/sourcebraid
codex plugin add sourcebraid@sourcebraidStart a new Codex conversation after installation.
Export Plugin Config downloads sourcebraid-config.json. Store it at:
~/.config/sourcebraid/config.json
Or configure the plugin in a terminal:
python3 codex-plugin/sourcebraid/scripts/sourcebraid.py config \
--repo-slug OWNER/REPO --branch main --root-folder web-clipsThe versioned plugin lives under codex-plugin/sourcebraid. It uses a local
SQLite FTS5 index, downloads only files whose Git blob SHAs changed, and
supports search, fetch, listing, and guarded deletion previews:
python3 codex-plugin/sourcebraid/scripts/sourcebraid.py index build
python3 codex-plugin/sourcebraid/scripts/sourcebraid.py index update --max-age 900
python3 codex-plugin/sourcebraid/scripts/sourcebraid.py index verify
python3 codex-plugin/sourcebraid/scripts/sourcebraid.py search "dynamic agents" --tag ai
python3 codex-plugin/sourcebraid/scripts/sourcebraid.py list "dynamic agents" --refresh
python3 codex-plugin/sourcebraid/scripts/sourcebraid.py plan-delete \
--path "web-clips/2026/07/example.md" --jsonThe index is stored separately per repository, branch, and archive root under
~/.cache/sourcebraid/.../search.sqlite3 and is never committed. Search checks
for a changed remote head at most every 15 minutes; when GitHub is unavailable,
the local index remains usable. search --scan is an explicit rg diagnostic
fallback.
The safer cache namespace introduced in this release triggers a fresh initial
build; existing Markdown and Git history are unchanged.
New captures write stable URL-hash shards such as
web-clips/index/47.jsonl. Legacy archives remain readable. Preview and confirm
the one-time migration against an unchanged branch head:
python3 codex-plugin/sourcebraid/scripts/sourcebraid.py index plan-shards --json
python3 codex-plugin/sourcebraid/scripts/sourcebraid.py index migrate-shards \
--expected-head HEAD_SHA --confirm-head HEAD_SHA --jsonBefore deletion, the plugin shows the exact Markdown file, metadata change, and owned assets, then requires explicit confirmation. It writes a normal, non-forced Git commit, so repository history remains recoverable.
The complete repository plugin includes the local stdio MCP server. The optional skills package contains the local Python workflows and is separate from the hosted With MCP submission. Requested sources are processed by the active ChatGPT or Codex environment; hosted requests also pass through SourceBraid and Cloudflare. See the integration guide and privacy notice for authentication, storage, and revocation details.
The native iOS app and Share Extension live under ios/.
After one-time repository and token setup, URLs, selected text, Safari articles
with relevant image assets, PDFs, and other files can be sent to the configured
private archive through the system share sheet.
Android is intentionally not part of the first release. Feedback from the OpenAI community will determine whether an Android share target becomes the next native client and which contributors or testers can help shape it.
SourceBraid is fully open source under the MIT License. See CONTRIBUTING.md for contribution guidance and DCO sign-off, the Code of Conduct for community standards, and the Security Policy for private vulnerability reporting. The public release process documents versioning, validation, and reproducible release artifacts. The privacy notice, terms, and notice document the capture, local-plugin, and hosted-service data flows and licensing boundaries.
The Chrome extension source lives under
chrome-extension/sourcebraid. It has no build
step and bundles no third-party runtime. Docling runs only inside the target
repository's GitHub Action. HTML conversion happens locally in the extension;
API and image requests use either ordinary HTTP or the browser's existing
authenticated session, depending on the source.
