Markdown
Convert a page's main content to clean, well-formatted Markdown. URLpipe keeps the substance — headings, paragraphs, lists, tables, links and fenced code — while dropping navigation, sidebars, cookie banners and other chrome.
This is the go-to endpoint for feeding web pages to an LLM or a RAG pipeline: Markdown is the format models work best with, and because the page is rendered with headless Chrome first, JavaScript-heavy sites produce complete content too.
The same page always produces the same Markdown, and links come back as absolute URLs you can follow without the page they came from.
It uses no AI — the conversion is a walk over the rendered DOM, not a model call — so it costs 1 credit per call, the same as /html. See Credits.
Convert to Markdown
Body parameters
- Name
url- Type
- string
- Required
- Required
- Description
- The absolute URL of the page to process. Rendered with headless Chrome, so JavaScript runs and redirects are followed. It must not include a username or password (
https://user:pass@example.com).
- Name
page_options- Type
- object
- Description
- Wait for the page, and remove ads, cookie banners or your own elements from it before it is read — gone from this result, not merely hidden. See Page options.
- Name
residential- Type
- boolean
- Description
- Fetch the page from a residential exit — an address on a home broadband line rather than one in a datacentre. Reach for it when a site serves you less than it serves a browser, or nothing at all. Defaults to
false. Adds 25 credits per page fetch on top of what the operation costs, and its results are kept separate from the ordinary ones.
- Name
report_to- Type
- string
- Description
- Webhook URL — an
httporhttpsaddress URLpipe POSTs the result to when it's ready. Optional: without it we deliver to your project's default endpoint if it has one, and otherwise send no webhook at all — the result still waits for you at GET /result/:token. A value we cannot deliver to returns422. Ignored on async=truerequest. Deliveries can be signed so your endpoint can verify they came from us.
- Name
sync- Type
- boolean
- Description
- Process the request synchronously, returning the result inline in the response. Defaults to
false(async: return a token now, and either receive the result at a webhook or fetch it with GET /result/:token). See Async & sync modes for the full contract.
- Name
max_age- Type
- string | integer
- Description
- How fresh a cached result must be to be accepted. Either an integer number of seconds (
3600) or a duration string of the form"<number> <unit>"— unitss/min/h/d/w(e.g."2 hours","3 days","30m"). Defaults to7 days, clamped to a max of30 days;0always bypasses the cache. See Caching for all accepted units.
- Name
labels- Type
- object
- Description
- Your own keys to find and account for this request by — a client, a project, a campaign:
{"client": "acme"}. Returned with the result, in the webhook and in theX-Labelsheader, and your dashboard filters history and totals credits by them. Up to 16 keys; string values. See Labels.
Response
Content type text/plain — the body is the page's main content as Markdown.
curl -X POST https://urlpipe.dev/markdown \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com"}'import requests
res = requests.post(
"https://urlpipe.dev/markdown",
headers={"Authorization": "Bearer YOUR_API_KEY"},
json={"url": "https://example.com"},
)const res = await fetch("https://urlpipe.dev/markdown", {
method: "POST",
headers: {
"Authorization": "Bearer YOUR_API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({ url: "https://example.com" }),
})$ch = curl_init("https://urlpipe.dev/markdown");
curl_setopt_array($ch, [
CURLOPT_POST => true,
CURLOPT_RETURNTRANSFER => true,
CURLOPT_HTTPHEADER => [
"Authorization: Bearer YOUR_API_KEY",
"Content-Type: application/json",
],
CURLOPT_POSTFIELDS => json_encode(["url" => "https://example.com"]),
]);
$response = curl_exec($ch);# Example Domain
This domain is for use in illustrative examples in documents. You may
use this domain in literature without prior coordination or asking for
permission.
[More information...](https://www.iana.org/domains/example)Size limit
422 with The page is too big to be processed. — see Errors.Responses
Whatever the status, the response carries metadata headers: the result token, whether it was served from cache and how old that result is, how long we took, what it cost in credits, and the allowance you have left.
200 OK422 Unprocessable Entity429 Too Many Requests504 Gateway Timeout401 UnauthorizedTry it live — no API key needed
Run this endpoint against any URL right in your browser.