How this blog is published: Markdown, a Python script, S3 and CloudFront

The pipeline behind these posts: a content collection validated at build time, scheduled publishing by date, generated social cards, alias redirects, and a diff-sync publisher that only invalidates the pages that actually changed.

The publishing pipeline as five steps: Markdown files, an Astro build with a Zod schema, a Python sync that diffs by hash, a private S3 bucket, and CloudFront with an edge function.

There is no CMS behind this blog. There is a folder of Markdown files, a build, and a Python script that copies the result into a bucket. I want to write down how the pieces fit together while it is still small enough to describe in one page.

The content is files

Every post is one file in frontend/src/content/blog/, either .md or .mdx. The file name is the slug and therefore the URL, which is the first constraint the whole design has to respect: renaming a file is renaming a published page.

Astro’s content collections turn that folder into typed entries, and the frontmatter goes through a Zod schema at build time. That schema is not decoration. A missing description fails the build rather than shipping a page with no meta description. A heroImage without heroImageAlt fails. An updatedDate before pubDate fails, because a dateModified earlier than datePublished is rejected by structured-data validators and reads as a typo to everyone else. Tags have to already be URL-safe slugs, so nothing has to guess how to spell dev-hub in a path.

The failure mode I care about is the quiet one — a page that ships slightly wrong and nobody notices for a month. Making the build refuse is cheaper than noticing.

Publishing on a date, not on a push

pubDate in the future means the post does not exist yet. It is filtered out of the index, the tag pages, the sitemap, the RSS feed and llms.txt by a single function every consumer reads through, so there is no path by which a scheduled post leaks into one output but not another.

A post with a past pubDate is published and one dated next Tuesday does not exist yet; both pass through one function that feeds the blog index, tag pages, sitemap, RSS feed and llms.txt
Every output reads through the same date filter, so a scheduled post cannot leak into just one of them.

The deploy workflow runs daily. That is the entire scheduling mechanism: a post dated next Tuesday goes live on Tuesday because a build happened on Tuesday, with no commit and nothing running in between.

Each post has a header illustration, and it doubles as the post’s 1200×630 social card. It is laid out from an explicit object tree with satori and rasterised with sharp, using font files committed into the repo, and the image is committed next to the post — a card rendered from whatever fonts the build machine happens to have is a card that changes when the build agent does, and a scraper caches the first version it sees.

For renames there is aliases. Listing an old slug there generates a redirect page at the old URL, and the build throws if an alias collides with a real post or if two posts claim the same one. This post has an alias, which is how I know that path works.

Getting it into the world

The site is a directory of static files, so publishing is a copy. The Python script that does it compares hashes against what is already in the bucket and uploads only what differs, then asks CloudFront to invalidate exactly those paths. Invalidating /* on every deploy would work too, and would throw away a warm cache to republish one typo fix.

Cache headers are set per file class on upload: hashed asset filenames get a long immutable max-age because their names change when their contents do, while HTML gets a short one so a correction is not stuck behind a week of CDN cache.

The bucket itself is private. CloudFront reads it through an origin access control, and a small function at the edge maps /foo/ to /foo/index.html, redirects the extensionless form to the trailing-slash one, and serves 404.html with an actual 404 status. That last detail matters more than it sounds: a soft 404 is a page search engines index as real content.

The static build is compared by hash with the bucket, only changed files are uploaded, and only those paths are invalidated; below, cards for not invalidating /* on every deploy, long immutable cache headers for hashed assets and short ones for HTML, and the edge function mapping /foo/ to /foo/index.html, redirecting /foo to /foo/, and serving 404.html with a 404 status
Publishing is a copy of what changed, with cache lifetimes and URL handling decided per file.

Files in a folder, one build command, one sync. When the blog outgrows that, it will be because something here started costing more than it saves — not before.

Get early access.

Invite-only, released in small waves. Free during early access.

Get early access.

Invite-only, released in small waves. Free during early access.