Key takeaways
- The blog is Markdown files and one Python module, and publishing a post means copying one file.
- Posts render on request and are re-parsed only when a file’s modification time changes.
- Google stopped showing FAQ rich results on 7 May 2026, so FAQPage markup earns nothing visible there.
- Escape the closing-tag sequence in inline JSON-LD and never HTML-escape it, because script contents aren’t decoded.
- At 500 posts a warm page takes milliseconds, but the first request after a restart parses for seconds.
The 9io.ai blog has no CMS and no build step. Each post is a Markdown file with YAML front matter, and one Python module of about 600 lines renders it when someone asks for the page. The same module writes the JSON-LD, adds posts to the sitemap, builds the RSS feed and appends a section to llms.txt. Publishing a post means copying one file to the server.
We built it this way because the rest of 9io.ai is already a small FastAPI app serving hand-written HTML, and a second toolchain for a blog this size didn’t pay for itself. The blog added one dependency, Python-Markdown. PyYAML was already in the project. Most of the thinking went into the search plumbing, including two parts that Google changed in the last two years.
This post walks through the request path, the front matter and the checker that enforces it, Markdown rendering, structured data and how we escape it, canonical URLs and the feeds. It ends with timings for 500 synthetic posts and what we’d change at that size.
The request path, from file to page
Four routes serve the blog, in the same FastAPI app as the rest of the site.
| Route | Returns | Cache-Control |
|---|---|---|
/blog |
The index, or a tag-filtered list with ?tag= |
no-cache, must-revalidate |
/blog/feed.xml |
The RSS 2.0 feed | public, max-age=3600 |
/blog/og/{slug}.png |
A post’s share image, if one exists | public, max-age=604800 |
/blog/{slug} |
The post, or a 404 page | no-cache, must-revalidate |
The feed route is registered before the post route because FastAPI matches path operations in the order they are declared (path parameters). A slug that doesn’t match [a-z0-9-]+ gets a 404 straight away, before any lookup.
Every page, feed and sitemap request then calls the same function. It lists posts/*.md, reads each file’s modification time and re-parses a file only when that time has changed. Parsed posts stay in a dict keyed by slug, and entries for deleted files are dropped. A file that fails to parse is logged and skipped, so one bad post can’t take the blog down. Simplified, it looks like this:
def all_posts() -> list[Post]:
today = datetime.now(timezone.utc).date()
posts = []
for path in POSTS_DIR.glob("*.md"):
mtime = path.stat().st_mtime
cached = _cache.get(path.stem)
if cached is None or cached[0] != mtime:
try:
cached = _cache[path.stem] = (mtime, parse(path))
except Exception:
log.exception("could not render %s", path.name)
cached = _cache[path.stem] = (mtime, None)
post = cached[1]
if post is not None and post.published <= today:
posts.append(post)
return sorted(posts, key=lambda p: (p.published, p.slug), reverse=True)
Pages are built from f-strings, with no template engine. HTML goes out with Cache-Control: no-cache, must-revalidate, so an edit shows on the next request, and at our size a warm render costs well under a millisecond (timings below). A storefront with expensive pages is a different case, which we cover in caching server-rendered HTML at the edge.
Scheduled posts and drafts
The date check in that loop is the whole scheduler. It compares each post’s date with today’s date in UTC on every call, after the cache lookup, so a post that was parsed days earlier appears at 00:00 UTC on its date without anyone touching the file. The index, the post page, the sitemap, the feed and llms.txt all read the same list, so the post shows up in all of them together. There is no cron job and no rebuild.
Two side effects come with it. The feed, the sitemap and llms.txt are sent with Cache-Control: public, max-age=3600, so a cache between us and a reader can serve the previous version for up to an hour after midnight. And until its date, a scheduled post returns 404, so other posts shouldn’t link to it early. A file with draft: true never appears anywhere.
Front matter, and the checker that enforces it
Each post starts with a YAML block. The slug comes from the file name.
| Field | Required | Used for |
|---|---|---|
title |
Yes | The page title, the H1, Open Graph tags and the BlogPosting headline |
description |
Yes | Meta description, cards, RSS and llms.txt; 160 characters at most |
date |
Yes | Publish date; a future date hides the post until that day |
updated |
No | The “Updated” line, dateModified and the sitemap’s lastmod |
tags |
No | The first tag is the topic shelf; every tag becomes an article:tag |
author |
No | Defaults to “9io Engineering” |
takeaways |
No | The “Key takeaways” box above the body |
faq |
No | Questions shown after the body and marked up as FAQPage |
image |
No | Share image; defaults to /blog/og/<slug>.png when that file exists |
draft |
No | Keeps the post out of every list and route |
We read it with yaml.safe_load. PyYAML’s documentation warns that yaml.load is as powerful as pickle.load and may call any Python function, while safe_load only builds standard types (PyYAML documentation). Our posts are our own, but nothing in front matter needs more than strings, lists and dates.
Two YAML details matter here. An unquoted 2026-10-08 is a YAML timestamp, so with PyYAML 6.0.3, yaml.safe_load("date: 2026-10-08") returns datetime.date(2026, 10, 8). The module accepts a date, a datetime or a string. An impossible date such as 2026-13-45 raises ValueError while the YAML loads, and the module treats that like any other parse failure. The other detail is the colon. A colon followed by a space inside an unquoted value breaks parsing, so title: Caching: a guide fails with “mapping values are not allowed here”. The template in our style guide puts the title, description, takeaways and FAQ text in quotes, which avoids it.
A checker runs before a post is copied to the server, and a post ships only when it reports no errors. It validates the slug and the front matter, checks the shape of the body (it must open with a paragraph, contain no H1, have at least four sections and cite at least one source) and tries to render the post. The render step matters because the site skips a post that fails to render, logging the error instead of crashing, so a broken file would otherwise just be missing. The checker also blocks client names, hostnames and secrets before a post ships.
Rendering Markdown with Python-Markdown
The body goes through Python-Markdown 3.9 with four extensions.
extraadds abbreviations, attribute lists, definition lists, fenced code, footnotes, tables and Markdown inside HTML (Extra).tocgives every heading an id and exposestoc_tokens, which feed the “On this page” sidebar (Table of Contents).sane_listsmakes list syntax “less surprising”, in the docs’ words (Sane Lists).smartyconverts ASCII dashes, quotes and ellipses to HTML entities (SmartyPants), so authors type straight quotes and readers see curly ones.
After conversion, the module wraps each table in a scrolling container, turns bare URLs in the text into links (footnotes often cite one) and counts words for the reading time at 230 words a minute. Front-matter strings never pass through Markdown, so a small regex gives titles and descriptions the same curly quotes.
Python-Markdown is not a CommonMark implementation, and its documentation says so (Python-Markdown). One difference to know about is indentation. Nested list content needs four spaces, or a tab, per level. Raw HTML passes through untouched. Python-Markdown deprecated its safe_mode option in version 3.0 and recommends running untrusted content through an HTML sanitiser (changelog). Every post here is written by us, so we don’t sanitise. A blog that accepts outside contributions should.
One interaction caught us while drafting this post. smarty runs before toc records heading names, so a heading such as “Why it’s slow” comes back from toc_tokens as Why it’s slow, which is already HTML. Escape it again for the sidebar and readers see ’ in the table of contents. Python-Markdown describes the token’s name as the sanitised label for its own table of contents, so insert it as it is, or unescape it before escaping.
Structured data for posts, breadcrumbs and FAQs
Each post page carries one JSON-LD block whose @graph holds up to three nodes. The index page carries a Blog node listing the newest 50 posts, plus its own breadcrumb.
| Node | Where | What Google does with it, as of 10 September 2026 |
|---|---|---|
BlogPosting |
Every post | Article markup, which Google says helps it show title, image and date information |
BreadcrumbList |
Index and posts | A breadcrumb trail on desktop results only, since January 2025 |
FAQPage |
Posts with an FAQ | Nothing; FAQ rich results stopped appearing on 7 May 2026 |
Blog |
Index | No Google feature we rely on; posts point to it with isPartOf |
Trimmed, the post node looks like this:
{
"@type": "BlogPosting",
"@id": "https://9io.ai/blog/building-the-9io-blog-markdown-structured-data#post",
"headline": "How we built the 9io.ai blog on Markdown files and FastAPI",
"datePublished": "2026-09-10",
"dateModified": "2026-09-10",
"author": {"@type": "Organization", "name": "9io Engineering", "url": "https://9io.ai/"},
"publisher": {"@type": "Organization", "@id": "https://9io.ai/#organization", "name": "9io.ai"},
"isPartOf": {"@id": "https://9io.ai/blog#blog"}
}
Google’s Article documentation lists no required properties. It recommends author, author.name, author.url, dateModified, datePublished, headline and image, with images of at least 50,000 pixels (width times height) in 16x9, 4x3 and 1x1 ratios (Article structured data). We cover most of that. The gap is author.url for named authors. The default author is an Organization with a URL, but a post with a named author gets a Person without one, because we have no author pages yet. The byline shows the same dates as the markup, which Google’s guidance on byline dates asks for (publication dates).
Breadcrumbs changed first. On 22 January 2025 Google limited breadcrumb display to desktop results, because breadcrumbs get truncated on small screens (changelog). Desktop results still show them, and the markup costs three list items.
FAQs changed this year. Google had limited FAQ rich results to well-known government and health sites since 2023 (Changes to HowTo and FAQ rich results). On 8 May 2026 it added a deprecation notice saying the feature would no longer appear from 7 May 2026, and on 15 June 2026 it removed the documentation (changelog). We kept the markup anyway. The questions and answers are part of the page, in <details> elements readers can open, and Google’s general guidelines ask that marked-up content be visible to readers (structured data policies). The code is one dictionary. We expect nothing from Google for it, and we wouldn’t add an FAQ to a post for search. Ours are there for readers who arrive with one specific question.
Google’s guide to its generative AI features adds that structured data isn’t required for them and that there is no special schema.org markup to add (AI features guide).
Escaping JSON-LD inside a script tag
JSON-LD lives in a <script> element, and the HTML spec makes script a raw text element (HTML syntax). Its content can’t contain `
Frequently asked questions
Do you need a CMS for a company engineering blog?
Not when engineers write the posts themselves. Markdown files, a small renderer and a pre-publish checker cover it. A CMS earns its keep when non-engineers publish or you need an editorial workflow.
Is FAQPage structured data still worth adding?
Not for Google. Google stopped showing FAQ rich results on 7 May 2026 and removed the documentation in June 2026. Keep an FAQ only if it helps your readers.
How do you put JSON-LD inside a script tag safely?
Serialise it with a JSON library, then replace every “</” with “<\/” so no string can close the script element. If any value comes from users, replace every “<” with “\u003c” instead.
Does Google use llms.txt?
Not for Search. Google’s documentation says llms.txt files aren’t needed for Google Search and won’t help or harm visibility or rankings. It is a proposal that other tools may read.
How can a blog publish future-dated posts without a cron job?
Filter posts by date on every request. Ours compares each post’s date with today’s date in UTC, so a post dated tomorrow appears on the site, in the sitemap, in RSS and in llms.txt at midnight UTC.
Should blog tag pages be indexed?
Ours aren’t. They list the same posts as the blog index, so they send noindex,follow, which keeps them out of results while crawlers can still follow their links.
Building something like this?
9io is a small team of senior engineers with a fractional CTO, and we work by the hour. Send us a note about your product. The reply comes from the person who'd do the work.