Machine-readable surfaces

Alongside its human pages, the site publishes surfaces meant for tools: crawlers, feed readers, security tooling, and language models. The complete inventory, with live links, is the legal page Public machine-readable files; this page states the rules — which kinds of surfaces exist, why, and how a new one gets added.

Two sitemaps#

An XML sitemap is for search engines. An HTML sitemap is for people.

/sitemap.xml is generated by Hugo and lists canonical URLs for crawlers. The human-readable sitemap is deliberately selective: a structured index of the pages a visitor might actually want, not an inventory of every generated URL. Feeds, language aliases, and individual tag pages do not belong in it.

The HTML sitemap is a normal content page with a custom layout — content files for meaning, templates for structure:

---
title: "Sitemap"
layout: "sitemap"
translationKey: "sitemap"
---

The layout chooses the sections that matter and lists their pages; a page earns its place by being useful to a reader, not by being possible to generate.

LLM context files#

/llms.txt carries concise site context for language models, and /llms-full.txt the expanded version. Both follow the llms.txt proposal, which standardizes a root file that helps LLMs use a website at inference time. Scraping and training permissions stay governed by /robots.txt and X-Robots-Tag headers, as stated in the legal inventory.

Feeds#

Each language publishes three syndication formats: RSS (index.xml), Atom (index.atom), and JSON Feed (feed.json). Feeds are surfaces for readers’ tools, not pages to navigate, so they stay out of the HTML sitemap.

Security and contact files#

/.well-known/security.txt and its root copy carry vulnerability disclosure and security contact metadata; /humans.txt and its compatibility copy credit the people and tools behind the site.

How surfaces are published#

Two generation conventions cover every surface:

  • Content page with an explicit url:: root text files such as security.txt, humans.txt, and llms.txt are normal content files whose front matter claims the root path, per URL strategy
  • Hugo output formats and templates: /sitemap.xml, the feeds, and the search index under /pagefind/ are generated at build time from existing pages, with no content file to edit

Adding a new surface#

  • prefer a standard name and location: established files stay at the site root, newer metadata goes under /.well-known/
  • publish it as a content file with url: when it has editable content; use an output format when it derives from existing pages
  • register it in the legal inventory page, in both languages
  • keep it out of reader-facing indexes when it is not a reader page, using sitemap_exclude and pagefind_exclude