Link Health

External links rot. The page you cited last year moves, the vendor renames a product, the domain lapses — and your readers hit the 404 before you do. Leed checks every external link on your site on its own schedule and lists the broken ones where you will see them.

What gets checked, and when

Once a day, Leed gathers the external links from every page in each active workspace and checks them. A workspace is active when it has a real public domain provisioned and has not been deleted — so a workspace that has never been deployed is skipped, which is why a brand-new site reports nothing.

There is no schedule to set and no button to press.

Every absolute http or https link in the latest revision of every page that has not been deleted.

  • Links between your own pages. Internal links resolve by page identity rather than by URL, so they cannot silently break when a page moves — see URL paths and slugs and aliases and redirects. They are never fetched and never appear in the report.
  • Relative links. Only absolute http/https URLs are collected.
  • Excluded hosts. localhost and 127.0.0.1 are never fetched; a link to either is reported with a status of 418 rather than checked.

Leed fetches each link with GET, identifying itself as Leed-LinkChecker/1.0.

Before any fetch, the destination origin’s robots.txt is read and consulted for that user agent. If it disallows the URL, Leed does not fetch it and reports 451 instead. A robots.txt that cannot be fetched at all is treated as permission granted — the assumption is that a site with no robots.txt has not disallowed anything.

Redirects are followed manually, up to 10 of them, and the final destination is what gets reported. Each request times out after 10 seconds. Any 2xx response means the link is healthy and it is not listed. Everything else is reported.

flowchart TD
    A[Daily cron] --> B[For each active workspace]
    B --> C[Collect absolute http/https links<br/>from every page's latest revision]
    C --> D[Group by domain into batches of ~50]
    D --> E{Host excluded?}
    E -- yes --> F[418 — not checked]
    E -- no --> G{robots.txt allows it?}
    G -- no --> H[451 — not verified]
    G -- yes --> I[GET with a 10s timeout]
    I --> J{Redirect?}
    J -- yes, under 10 --> I
    J -- yes, over 10 --> K[310 — too many redirects]
    J -- no --> L{2xx?}
    L -- yes --> M[Healthy — not listed]
    L -- no --> N[Reported with its status]
    F --> O[Write the report]
    H --> O
    K --> O
    N --> O
    O --> P[Broken links table on Know]

Grouping by domain is not cosmetic: each batch runs with its own concurrency limit of six requests, so a site with two thousand links to one busy host is not hammering it, and a slow host cannot hold up the rest of the crawl.

The checker's operating parameters
SettingValue
User agentLeed-LinkChecker/1.0
Request methodGET (never HEAD)
Request timeout10 seconds
Concurrent requests per batch6
Maximum redirects followed10
Batch size~50 URLs, packed by domain and never split mid-domain
Excluded hostslocalhost, 127.0.0.1

GET rather than HEAD is deliberate. A large minority of sites answer HEAD with a 405 or simply do not implement it, which would produce a flood of false failures.

Where results appear

The Broken links table sits at the foot of the Know dashboard, subtitled from the latest link crawl, with a count of how many rows were found. Rows are grouped under the page that contains them, and the page title links straight into the editor.

The Broken links table at the bottom of the Know dashboard, with rows grouped under their pages
ColumnMeaning
PageThe page containing the link, linked to its editor. Shown once per page, above that page’s rows.
URLThe link that failed, opening in a new tab so you can check it yourself.
StatusThe HTTP status, or one of the codes Leed generates — see below.
MessageThe plain-language explanation of that status, e.g. Not Found.

When there is nothing to report the table reads “No broken links detected.” On a workspace that has never been deployed, that is also what an unrun crawl looks like.

Every status you can see

Real HTTP statuses

When the destination answered, its own status is what you see, and it means what it says. A 404 or 410 is a genuinely missing page. A 500 is the destination failing, which may be temporary. Any non-2xx answer is reported.

Statuses Leed generates

The rest are Leed’s own codes, chosen so that “we could not verify this” is never confused with “this is dead”:

StatusWhat Leed means by itIs the link dead?
310More than ten redirects were followed without reaching a final responseProbably — a redirect loop, or a chain that no longer terminates
403The destination refused an automated request. Also what a LinkedIn-style 999 response is rewritten toUsually not. Many large sites block every checker
418The host is on Leed’s exclusion list (localhost, 127.0.0.1) and was never fetchedNot on the public internet — remove it from the page
451The destination’s robots.txt disallows this URL for Leed’s user agent, so no request was madeUnknown. Check it in a browser
502The connection was refused, or the fetch failed outrightLikely — the host is not answering
504The request was aborted after the 10-second timeoutMaybe — a very slow host can produce this while still working
520An error Leed does not have a more specific code for. Also used for a redirect that returned no Location headerInvestigate — this one is genuinely ambiguous
521A socket to the host could not be openedLikely — DNS or the server is down
526The destination’s TLS certificate is expired, self-signed, or does not match its hostnameThe page may load, but browsers will warn readers before it does

Fixing what it finds

  1. Click the page title in the Broken links table to open that page in the editor.
  2. Find the link — searching the page for the distinctive part of the URL is usually fastest.
  3. Update it to the new destination, or remove it if no replacement exists.
  4. Publish the page, as described in publishing changes.

The row disappears on the next daily run. There is no “recheck now” control, and nothing to mark as resolved — the list is rebuilt from scratch each time, so a fixed link simply stops appearing.

The pattern worth watching for is a whole host reporting 403 or 451 at once. That is the destination’s policy, not your content, and no amount of editing on your side will clear it.

ESC