ALICE · LEARN CHARLIEHUB

Mail club · issue no. 01 posted from the desk where the trip cards get made

Reference · /learn/

How this site builds itself

The trip cards are not written by hand. They are generated, every time, from seven JSON files and a folder of templates. This is what actually happens, with the real files.

The whole idea in one line: content lives in data/*.json, layout lives in templates/*.html.j2, and build.py is the 247 lines of plumbing that puts one through the other. Everything below is a consequence of that split.

~/projects/itinerary-gen 7 data files → 7 pages build.py · 247 lines Jinja2 + stdlib
Marks
  • the data
  • the templates
  • build.py
  • the output
  • determinism
  • the tradeoffs
Colours
  • Paper#F6F1E7the page itself
  • Powder#C3D9E3dots & light fills
  • Dusty#6C8EA0seal, rules, headings
  • Navy ink#2A3E4Eevery word
  • Card#EFE8DAthe things on top
  • Coral#C15C4Fused sparingly

01Part

Why content and markup live apart

One folder holds what the trip is. Another holds what a page looks like. Nothing lives in both.

Every page on the trip side of this site comes from one JSON file in data/. There are seven of them:

~/projects/itinerary-gen/data/content only
hakone.json      # one file per page
kyoto.json
learn.json
osaka.json
singapore.json
tokyo.json
zandvoort.json
friend-recs.md   # not JSON, so build.py ignores it

A JSON file is nothing but nested values — lists, objects, strings, numbers. Here is roughly the shape of one, trimmed hard:

data/tokyo.json — abridged
{
  "template": "tokyo.html.j2",
  "out": "tokyo-field-card.html",
  "tz": "+09:00",
  "areas": [
    {
      "name": "Shibuya",
      "itinerary": { "days": [ /* ... */ ] }
    },
    { "name": "Yanaka",        /* ... */ },
    { "name": "Shimokitazawa", /* ... */ }
  ]
}

Notice there is no <div> anywhere in it. No colours, no fonts, no layout. The JSON does not know it is going to become a web page — it would be equally valid input for a PDF, a text file, or a phone app.

The matching template is the opposite: it is all markup and almost no content.

templates/ — layout only
base.html.j2       # shared scaffolding: nav, weather, currency toggle
_macros.html.j2    # small reusable snippets
tokyo.html.j2      # the Tokyo skin (extends base)
learn.html.j2      # standalone: no nav, no timeline
...
Why bother

Because the two change for completely different reasons. Adding a restaurant to Osaka is a content change. Making every day-heading slightly larger is a layout change. If they lived in the same file, every content edit would risk breaking layout, and every layout edit would mean finding and repeating the same change in six places.

There is a third folder, partials/, which is the escape hatch. Anything genuinely one-off — a hand-drawn diagram, a service worker, a page's own CSS — is a plain file there, pulled into the template whole. That is deliberate: data-drive what repeats, paste in what doesn't. A one-off diagram expressed as JSON would be all overhead and no reuse.

Section Review Card

no. 01 of 06
Concept:
Content and layout never share a file.
Key file:
data/tokyo.json + templates/tokyo.html.j2
One thing I actually understood today:
Rating:

02Part

What Jinja2 actually does — the six-line version

Strip away the file handling and the whole templating engine reduces to one function call that turns data into a string.

This is the entire concept, run in a Python REPL. Nothing is hidden, and there is no file on disk anywhere in it:

python3 — the reductionreal output
>>> from jinja2 import Template
>>> Template('{% for a in areas %}<h3>{{ a.name }}</h3>{% endfor %}').render(
...     areas=[{"name": "Shibuya"},
...            {"name": "Yanaka"},
...            {"name": "Shimokitazawa"}])
'<h3>Shibuya</h3><h3>Yanaka</h3><h3>Shimokitazawa</h3>'

Read what happened. One <h3> was written. Three came out. The list decided how many.

There are only two kinds of marker in that template string, and they are the two you will use ninety per cent of the time:

{{ ... }}

Put a value here. {{ a.name }} is replaced by whatever a.name holds. It produces output.

{% ... %}

Do something here. {% for %}, {% if %}, {% extends %}. It produces no output of its own — it controls whether and how often the surrounding markup is emitted.

The point

build.py is fundamentally that one .render() call. Everything else in its 247 lines is plumbing: finding the files, reading the JSON, deriving a few extra fields, and deciding which directories to write the resulting string into. The templating itself is a single line.

That reframing is worth holding on to, because it tells you where to look when something goes wrong. If the page renders but the content is wrong, the problem is in the JSON. If the shape of the page is wrong, it is in the template. If the file lands in the wrong place or doesn't appear at all, it is in the plumbing — and the plumbing is the only part that can actually crash.

Section Review Card

no. 02 of 06
Concept:
{{ }} puts a value in. {% %} decides what happens.
Key file:
Template(...).render(**data)
One thing I actually understood today:
Rating:

03Part

What build.py does, end to end

Read JSON, render template, write HTML. Everything else is a detail hanging off one of those three.

data/*.json 7 files · content templates/*.j2 layout partials/* verbatim includes build.py render(**data) once per JSON file build/ always site/preview/ always site/ only with --publish ← the live one
Two copies always, a third only when asked. build/ and site/preview/ are safe to overwrite at any time; site/ is what alice.charliehub.net actually serves, so writing to it takes an explicit --publish.

The actual sequence

  1. Find the work

    sorted(DATA.glob("*.json")) — every JSON file in data/, alphabetically. Nothing is registered or listed anywhere; dropping a new file in is all it takes to add a page.

  2. Read one file

    json.loads(path.read_text()) turns the file into an ordinary Python dict. From here on it is just data in memory.

  3. Derive the fields that shouldn't be typed twice

    Three passes add computed values: add_price_html works out the £ equivalent of each price, add_when combines a day's date with a row's "09:30" into a full timestamp, and add_area_dates collects which dates each area covers. All three exist so no date or converted price is ever written out by hand in two places.

  4. Pick the template

    The JSON's "template" key, falling back to <filename>.html.j2. So data/tokyo.json finds templates/tokyo.html.j2 with no configuration at all.

  5. Pull in the verbatim files

    Each entry in "raw_includes" reads a file from partials/ and hands it to the template as a variable. That is how a hand-drawn SVG diagram gets into a page without being mangled into JSON.

  6. Render

    html = template.render(**ctx). The one line from Part 02. Data goes in, a single long string comes out.

  7. Write it, two or three times

    Always build/ and site/preview/; also site/ if --publish was passed. Same string, different directories.

  8. Stamp the sidecar files

    For pages that own a service worker, a SHA-256 hash of the rendered HTML is substituted into it. That is what makes a rebuild invalidate the offline cache — the worker's identity changes only when the page it caches actually changed.

Then it loops back to step 2 for the next file. That is the whole program.

running it
$ cd ~/projects/itinerary-gen
$ . .venv/bin/activate
$ python build.py

built hakone.json -> build/ + preview/
built kyoto.json -> build/ + preview/
built learn.json -> build/ + preview/
built osaka.json -> build/ + preview/
built singapore.json -> build/ + preview/
built tokyo.json -> build/ + preview/
built zandvoort.json -> build/ + preview/

(preview only — run with --publish to update the live site)
Never edit the output

Files in build/, site/preview/ and the published site/*-field-card.html are generated. Editing one works right up until the next build silently overwrites it. If a page is wrong, the fix is in data/ or templates/, always.

Section Review Card

no. 03 of 06
Concept:
Read JSON, render template, write HTML. The rest is detail.
Key file:
~/projects/itinerary-gen/build.py
One thing I actually understood today:
Rating:

04Part

Worked example: 3 of 26

Templating removes repetition it can see. It has no idea what your sentences say.

The test was simple: how much of a page does a single JSON edit actually control? The Zandvoort card mentions Delft. Counting occurrences of the word in the rendered HTML gave 26. Changing the name field in the JSON and rebuilding changed 3 of them.

26 occurrences of “Delft” in the rendered page 3 23 from {{ name }} renamed itself hand-written prose Delftware · Royal Delft · route line · day-label
The green slice is what the template controls. The other 23 are words inside sentences, and no amount of templating reaches them.

The 23 break down into four kinds, and each one is instructive:

Delftware

A different word that merely contains the place name. A rename must not touch this one — the pottery is still called Delftware wherever you are.

Royal Delft

A proper noun. The name of a specific business, which does not change just because the itinerary stops going there.

route line

Prose written into the JSON as a sentence — something like “train to Delft, then 12 minutes on foot”. It is in the data, but as text, not as a field the template substitutes. Templating cannot reach inside a sentence.

day-label

A heading typed out in full, like “Saturday — Delft”, instead of being composed from the name field. This is the one that is arguably a mistake, and the one worth fixing.

The lesson

Templating de-duplicates structural repetition — the same field rendered in many slots — not text you wrote into sentences. A real rename is therefore two jobs: edit the data, then read through the prose. Neither one alone is enough, and the build will not warn you, because 23 hand-written mentions of a place are perfectly valid HTML.

The useful habit that falls out of this: when writing content, notice whether you are typing a value or writing a sentence. A value belongs in its own field, where the template can use it in ten places. A sentence is a sentence, and you own it forever.

Section Review Card

no. 04 of 06
Concept:
Templating removes structural repetition, never prose.
Key file:
data/zandvoort.json
One thing I actually understood today:
Rating:

05Part

Rebuilding changes nothing, and why that matters

Build, hash, build again, hash again. The two hashes matched. That is a property worth protecting.

The test: take the SHA-256 of a rendered page, run build.py again with no changes to any input, and take the hash again. The file came out byte-identical — not merely similar, not the same length, the same bytes.

the shape of the check
$ sha256sum build/zandvoort/index.html
a1b2c3…  build/zandvoort/index.html

$ python build.py           # nothing changed in data/ or templates/
$ sha256sum build/zandvoort/index.html
a1b2c3…  build/zandvoort/index.html   # identical

This sounds like a non-result. It is not, because it is easy to get wrong and plenty of build tools do. A build stops being deterministic the moment it embeds a timestamp, iterates a Python set, reads files in filesystem order, or generates a random ID. Any one of those makes every rebuild produce a different file.

Two things this buys

build/ is disposable

Nothing of value lives in the output directories. They can be deleted entirely and reconstructed exactly from data/ and templates/. That is what makes it safe to treat generated files as worthless — and safe to never back them up.

diff becomes trustworthy

This is the big one. Because a no-op rebuild produces zero differences, any difference between two builds is real. You can change one line of JSON, rebuild, diff the output, and read exactly what your edit did to the page — with no noise to filter out.

Why the second one matters most

Without determinism, a diff of a rebuilt page shows hundreds of changed lines and you cannot tell which are yours. With it, the diff is the review: if you expected to change one heading and the diff shows forty lines, you have learned something before publishing rather than after.

It is worth treating this as a property to defend rather than a happy accident. The temptation to put a Generated at 14:22 line in a page footer is exactly what would destroy it — every rebuild would differ, and every diff would be noise. If a build ever needs a version marker, derive it from the content, the way the service-worker hash already does.

Section Review Card

no. 05 of 06
Concept:
Same inputs, byte-identical output — so every diff is real.
Key file:
build/zandvoort/index.html
One thing I actually understood today:
Rating:

06Part

The glob, and the incremental build I didn't write

Every run rebuilds all seven pages, including the six that didn't change. That is the right call, and it will stay right for a long time.

There is no change detection anywhere in build.py. It globs data/*.json, loops over the results, and rebuilds every single one. Fix a typo in Osaka and Tokyo, Kyoto, Hakone, Singapore, Zandvoort and the learn page are all re-rendered too, byte-for-byte identically to before.

A smarter build would compare modification times and skip the unchanged ones. Here is what that would cost:

What it would needWhy it's harder than it looks
Compare mtimes of input vs output Straightforward, and the only easy part.
Track template dependencies tokyo.html.j2 extends base.html.j2, which imports _macros.html.j2. Editing the base means every page is stale. The build would need to understand Jinja's inheritance graph.
Track partials Each page's raw_includes pulls in several files from partials/. Every one is a dependency, and they're declared in the JSON, not the template.
Handle cross-page state trip_schedule() reads all the JSON files to build the cross-card date routing. Change one card's dates and every other card's output changes. A per-file staleness check would miss this entirely.
Be debuggable when wrong The failure mode of a broken incremental build is a stale page that looks fine. That is a far worse bug than a slow build, and much harder to notice.

Against that: a full rebuild of all seven pages takes well under a second. There is no problem to solve.

The tradeoff, stated plainly

At this scale, rebuild-everything is not a limitation, it is a feature. It removes an entire category of bug — the stale artefact — in exchange for time nobody is waiting on. Dependency tracking is what you add when the build gets slow enough to hurt, and not one moment before.

The signal to revisit it would be concrete: builds slow enough that you hesitate before running one. That is when the tradeoff flips, because a build you avoid running is worse than a build that occasionally serves something stale. Seven files is nowhere near that line.

Footnote

This page and the rest of /learn/ are the exception to everything above: every issue is hand-written HTML in site/learn/, not generated from JSON. One-off documents have no structural repetition to remove, so a build step would be pure overhead — which is the same judgement call as partials/, applied to a whole page. Each issue carries its own stylesheet for the same reason; the one thing they all share, reference.js, is a separate file for exactly the reason Part 01 describes.

Section Review Card

no. 06 of 06
Concept:
Rebuild everything. Dependency tracking is a later problem.
Key file:
sorted(DATA.glob("*.json"))
One thing I actually understood today:
Rating: