Markdown for AI, MCP and Import

This page is for anyone writing Leed Markdown from outside the CMS — an MCP client, an import script, a generator. Three contracts decide whether your first write succeeds: how blocks are addressed, what the page type allows, and what the parser refuses outright. The fourth section is the other direction: the plain-markdown file a reader’s own AI client is served from your published site.

Blocks are addressed by data-id

Every top-level block in a page body carries a data-id. The schema assigns it; you never do. It is stable across edits, it is what an anchored operation targets, and inventing one is the fastest way to have a whole batch of edits rejected.

Read a body and you will see them on the blocks that have them:

## Installing the CLI {data-id="h-install"}

Run the installer and follow the prompts. {data-id="p-install-1"}

New content you send does not need one — ids are assigned on import.

The read-plan-apply loop

Editing an existing page is a three-step loop, and it is not optional: the anchors you edit against have to come from a read you performed moments earlier.

  1. get_page_markdown returns the body as Leed Markdown with every addressable block’s anchor intact. This is the planning input. It reflects the clean accepted-state body — a pending suggested deletion is not shown — and because reads go through the live collaborative document rather than a cached copy, the anchors are current at the moment you read them.
  2. apply_page_markdown_ops applies an ordered list of operations against those anchors.
  3. If any anchor no longer exists because the page changed under you, the whole op-set is rejected — not partially applied. Re-read and re-plan.
OpWhat it doesRequires an anchor?
insertBeforeInserts new Leed Markdown blocks immediately before the anchored blockYes, plus the markdown to insert
insertAfterInserts new blocks immediately after the anchored blockYes, plus the markdown to insert
replaceSwaps the anchored block for new markdownYes, plus the replacement markdown
deleteRemoves the anchored blockYes — and nothing else

A real payload:

{
  "pageId": "c29c7ee8-3d88-4357-a15e-3975ed0a048b",
  "ops": [
    { "type": "replace", "anchorId": "p-install-1",
      "markdown": "Run `leed install` and follow the prompts." },
    { "type": "insertAfter", "anchorId": "p-install-1",
      "markdown": ":::tip Verify the install\nRun `leed --version`.\n:::" }
  ]
}

Whole-body writes

fill_page_markdown sets a draft page’s entire body from markdown, and it is one-shot: it is allowed only while the body is still empty and rejected the moment there is anything there. There is no second fill. The only re-edit path is apply_page_markdown_ops.

That has a real consequence if you are writing at scale. Each body has to be right on the first write, because the fix for a wrong one is not another fill — it is a set of anchored operations that land as suggestions for a human to review. A hundred pages filled carelessly is a hundred-item review queue.

Two tools create pages, and they answer different questions:

ToolUse it whenWhat you get back
create_page_draftYou have the finished body nowA draft page with its content already set, in one call
create_pageOther pages need to link to this one before its body existsA reserved page with an empty body and a pageId

create_page is what makes a cross-linked set possible, because a pageid: link needs its target’s id to exist first. The id is always generated on the server — you cannot propose one — so building a linked set is genuinely two passes: reserve every page, harvest the returned ids, then author the bodies and the menu from that map.

One more thing about create_page: its slug argument is inert. The slug is always derived from the title. If you need a specific URL, set it afterwards with update_page_draft.

Edits land as suggestions

Anchored operations do not overwrite the document. They arrive as tracked additions and deletions attributed to the agent that made them, exactly as a human collaborator’s suggestions would, and a person accepts or rejects them in the editor. That is why concurrent human edits are never clobbered.

It also sets one hard limit: a block with no text cannot be tracked, so an operation whose target or payload is an image-only or iframe-only block is rejected. Suggest a change to the paragraph around it instead.

The page type decides which features are allowed

A page type carries nine editorFormattingOptions. Markdown that uses a feature the page type disables is rejected server-side, on every write path — the same validation the toolbar and the AI writing assistant both consult.

KeyDefaultGatesLabel in the rejection message
codeBlocktrueA fenced code block whose language is not mermaidcode blocks
diagramsfalseA fenced code block whose language is mermaiddiagrams (mermaid)
mathBlockfalseMath fencesmath blocks
alertfalse::: alert containersalerts
iconsfalse{% icon %}icons
tabGroupfalse===tabs-containertab groups
iframefalse{% iframe %}iframes
collapsibleBlockfalse+++ collapsible blockscollapsible blocks
tablefalseTablestables

Headings, paragraphs, lists, task lists, blockquotes, images, links, every inline mark, horizontal rules, hard breaks and the video, audio and form embeds are never gated by these flags.

The crucial scoping rule: only posts-type page types gate anything. On a documentation or an api page type the check returns immediately and every feature is allowed, the toolbar shows every control, and the Settings screen does not even render the toggles. If you are writing into a documentation set, this table describes a mechanism that is switched off for you.

When a gate does fire, the error names the features verbatim:

This page type does not allow: alerts, tables. Allowed formatting features: code blocks.

That message is worth parsing rather than retrying blindly — it tells you both what to remove and what you may use instead. Changing the toggles is a page-type setting, covered on What Each Page Type Lets You Format.

Raw HTML will be rejected

This is the single hardest constraint on writing into Leed, and it has no workaround.

The published site renders arbitrary HTML. The editor’s parser cannot, because the document schema has no node to hold it — so a body containing HTML does not lose the tag, it fails to parse at all:

MarkdownParseError: Could not parse the provided markdown

<br> is the only exception; it becomes a hard line break.

Most hand-written HTML has a Leed Markdown equivalent: alerts, collapsible sections, tabs, and attributes on ordinary blocks between them replace most wrapper markup. What is left usually belongs in a layout rather than in a page body. The full comparison is on Fidelity and Unsupported Syntax.

Escaping when markdown travels as JSON

Two rules, both of which only bite inside a payload and are therefore missed until something renders wrong.

Braces, colons and at-signs come back escaped. The serializer escapes literal {, }, : and @ in body text, so text you sent unescaped is read back with backslashes in front of those characters. That is expected and it does not affect the rendered page. Do not “fix” it on read and send the fixed version back — you will accumulate backslashes.

LaTeX backslashes double inside JSON. A math fence containing \frac{a}{b} has to be written \\frac{a}{b} in a JSON string, like any other backslash. The most common symptom of getting this wrong is a math block that renders as literal text.

There is no third rule about link fragments. A pageid: link may carry a #fragment — [label](pageid:<id>#section-slug) resolves to the target page’s heading — so if your client carries a guard that strips one, remove it. The mechanics are on Links and Internal Links.

The plain-markdown export

Every published page is also served as a plain-markdown file. This is the artifact most external AI clients actually consume, and it is a different thing from the Leed Markdown you write: Leed’s own extensions are stripped, and structure is kept.

flowchart LR
    Y["Page body (Y.Doc)"] --> S["serializeBodyToMarkdown"]
    S --> LM["Leed Markdown"]
    LM --> C["cleanContent<br/>embeds → HTML comments"]
    C --> P["leedMarkdownToPlainMarkdown<br/>strip attributes and fences' attrs"]
    P --> B["buildFlattenedMarkdown<br/>prepend # title"]
    B --> O["&lt;page url&gt;index.md"]
    B -. "api pages only" .-> A["append ## OpenAPI yaml block"]
    A --> O
    O --> D["Docs MCP page content"]

What gets stripped

Leed MarkdownIn the exported .md
{data-id="…"} and any other attribute blockRemoved; a line that is only an attribute block is deleted
Attributes on a code or mermaid fenceRemoved; the fence and its language survive
![Alt =320x200](/p.jpg)![Alt](/p.jpg) — the size is stripped from the alt text
[text](/x){target="_blank"}[text](/x) — the attribute suffix is stripped
A Handlebars commentRemoved entirely
:::note … :::Kept, fence and type word intact
+++ Title … +++Kept, fence and title intact
===tabs-container and @tab LabelKept, attributes stripped from both
Headings, lists, tables, blockquotes, code fencesKept unchanged, contents included

Structure survives; decoration does not. A reader’s client sees your alerts, your tabs and your collapsibles as recognizable fences, which is why an AI answering from these files can still tell a warning from a paragraph.

Embeds become comments

An embed has no plain-markdown equivalent, so each one is replaced by an HTML comment that says what was there:

EmbedIn the exported .md
{% form … %}<!-- embedded Leed form exists in the rendered page -->
{% youtube … %}<!-- embedded youtube video exists in the rendered page -->
{% cloudflare … %}<!-- embedded cloudflare video exists in the rendered page -->
{% audio … %}<!-- embedded audio exists in the rendered page -->
{% iframe … %}<!-- embedded iframe exists in the rendered page -->
{% icon … %}<!-- embedded icon exists in the rendered page -->

A client reading the file therefore knows a video was there and that it cannot show it, rather than silently seeing a gap.

Where it is served

The file lives at the page’s own URL plus index.md. This page is at /docs/leed-markdown/markdown-for-ai-mcp-and-import/index.md; the attributes page is at /docs/leed-markdown/attributes/index.md. The file opens with the page title as an # H1 — which is why a body should not start with one of its own — followed by the stripped body.

An API-type page gets one addition: its OpenAPI operation is appended as a fenced yaml block under a ## OpenAPI heading. Eleventy consumes that operation to render the human page and strips it from the HTML output, so the exported markdown is the only place the structured contract — method, path, parameters, security, responses — survives into hosted markdown. That is deliberate: it gives a client the exact spec alongside the prose describing it.

Who reads it

The Docs MCP hydrates a page’s content from exactly these files. When a reader’s assistant fetches a page from your documentation set, what it receives is the index.md above and nothing else — so anything you need an AI to know has to survive the strip. Docs MCP: AI Access for Your Readers explains how readers connect to it, and Docs MCP Tool Reference documents the tools that serve it.

The syntax guide an AI client can ask for

An MCP client does not have to guess at this format. describe_schema with target: "leed_markdown" returns the complete Leed Markdown syntax guide as text, generated from the same canonical specification these documentation pages are written against. The same guide is embedded in the description of every markdown field on every page tool, so a client that reads tool schemas has it before it makes its first call.

describe_schema also answers page, page_type, form and menu with JSON Schema, which is the fastest way to learn what a create or update call will accept.

The guide and this documentation are kept in step on purpose. If you find them disagreeing, that is a bug in one of them and worth reporting rather than working around.

Importing existing content

An import is a bulk version of everything above. Stated as requirements rather than as a tool walkthrough, content is importable when:

  • It contains no raw HTML except <br>. This is the requirement that fails most imports.
  • It invents no data-ids. Leave them out; they are assigned.
  • Internal links use pageid:, with a #fragment where a specific section is genuinely the destination. That means reserving every page first and rewriting links from the harvested id map, because a pageid: pointing at a page that has not been created resolves to nothing.
  • Images carry an asset URL and a data-assetid obtained beforehand. There is no asset-upload MCP tool — uploading is a separate authenticated HTTP flow — so images have to be in the asset library and their ids in hand before any body is written.
  • The leading H1 is removed. The layout renders the page title; a body starting with # Title publishes a duplicate.
  • The page type’s formatting options are honored — which, on a documentation or api page type, means everything is allowed.

The end-to-end sequence — which call to make when, and what each one returns — is Authoring Pages Over MCP, with full parameter tables in MCP Tools: Pages and Content. The reserve-first pattern is the same one human documentation authors follow, described from their side on Linking Between Docs Pages. And once a body is published, the JSON frontmatter that accompanies it in the site repository is a separate contract with its own required fields: Front Matter Reference.

ESC