ALICE · LEARN CHARLIEHUB

Mail Club · Winter Post the one where four exchanges never agree on when to open

Reference · /learn/

ISSUE 07

Base Layer

Two fetchers, one bug caught before it shipped, and four exchanges that have never agreed on when to open.

First Order ended with the environment set up and CLAUDE.md pinned in the project root. This is the first thing actually pulled in — real exchange rates, real share prices, and the reconciliation problem that shows up the moment you try to line four different calendars up on one chart.

Where this is going

Writing for a scale that doesn’t exist yet, two fetchers built separately on purpose, a cache filename that would have served the wrong data, two decisions about what gets written down and how long it stays true, four exchange calendars that never quite match, and a measurement of exactly how much got filled in — with real numbers, not a guess.

fetch_fx.py fetch_prices.py Frankfurter API yfinance forward-fill
Stickers
  • five portfolios, not one
  • two fetchers, one job each
  • what belongs in a cache filename
  • a log entry for every day
  • four calendars that don’t line up
  • measuring what got filled in
  • data that changes after you fetch it

THE DATA LAYER

01scale

Writing for five portfolios, not one

One driver’s portfolio today. The actual plan is five — worth building for that now, not after.

Before either fetcher existed, the real shape of the thing got settled first: not one portfolio, but up to five — different drivers, different sponsors, up to eight currencies between them. Nothing about that changes what gets built this week. It changes how it gets written.

Both fetchers take a list in, not a fixed set of five names baked into the function. fetch_asset_prices(tickers, ...) runs the same whether it’s handed one ticker or twenty. The difference between “works for George” and “works for the season” is entirely in that one design choice, made before any code existed to retrofit it into.

What actually costs something

A venv and a five-ticker list both looked, at first glance, like things worth keeping small. Neither was — the actual limit turned out to be somewhere else entirely.

The obvious worry was storage — five portfolios’ worth of daily data, cached for a year, sounded like it might add up to something. It doesn’t. Worked out properly: roughly half a megabyte, total, for the full season across every driver. The real risk at that scale isn’t disk space, it’s git history — every re-committed cache file keeps its old version forever, and a folder that gets regenerated and recommitted repeatedly gets heavier over months even while the current files stay tiny. Already solved, and for a reason that had nothing to do with the season expanding: data/ was gitignored from the first commit, so this never becomes a problem no matter how many portfolios get added.

Section Review Card

no. 01 of 07
Concept:
Write for the list you’ll actually have, not the one holding you today.
The check:
Worked out the real scale before worrying about it — half a megabyte, not a concern.
One thing I actually understood today:
Badges:

02fetchers

Two fetchers, one job each

Two separate fetchers, on purpose — a rate and a price are different kinds of thing, and they come from different places.

The first two pieces of code: one for exchange rates, one for share prices. Both built as their own file, each doing exactly one job.

fetch_fx.py — pulls a full date range from the Frankfurter API in a single call, base currency USD, symbols EUR/GBP/CHF. Not a loop, not one request per day — the range endpoint hands back the whole period in one trip:

one call, the whole range
GET https://api.frankfurter.dev/v1/2025-08-22..2026-08-22?base=USD&symbols=EUR,GBP,CHF

fetch_prices.py — loops through five tickers via yfinance, one at a time: ADS.DE, PUM.DE, CFR.SW, PVH, MBG.DE. Written to take any list of tickers, not hardcoded to these five — the plan is more portfolios later, and the fetcher shouldn’t need rewriting when that happens.

Both cache what they pull to a local data/ folder, so a rerun during development reads the saved copy instead of calling the API again every time.

Why USD, when the portfolio is in pounds

fetch_fx.py pulls everything against USD — the base layer everything else gets built on. The portfolio’s actual base currency is GBP, but that conversion is a calculation, not plumbing, so it’s deliberately left for later rather than baked in here.

Section Review Card

no. 02 of 07
Concept:
Two fetchers, one job each — rates and prices never share a file.
The habit:
Cache the raw response locally. Don’t call the API again for data you already have.
One thing I actually understood today:
Badges:

03caching

What belongs in a cache filename

Something that would have worked perfectly the first time and quietly handed back the wrong data the second.

Every cached file needed a filename — something to check against before deciding whether to fetch fresh or read what’s already saved. The first version used only the date range. That would have worked for exactly one scenario and broken silently in every other one.

the catchcaught before it ever ran
The cache filename includes base and symbols, not just the dates —
fx_rates_USD_EUR-GBP-CHF_2025-08-22_2026-08-22.json. Dates-only
would have been a real bug: a USD-based pull and a GBP-based pull
over the same dates are different data, and the second would
silently get served the first one’s cache.

The failure mode, spelled out: pull USD-based rates once, cache them under a dates-only name. Weeks later, pull GBP-based rates for the same date range. Same filename. The GBP request finds a file already sitting there and reads it back — except it’s not GBP data, it’s the old USD data, wearing the right name. Nothing crashes. Nothing warns. The numbers are just quietly wrong, and there’s no obvious moment where that would ever get noticed.

What the fix actually was

Fixed by making the filename carry everything that makes a request unique — base currency and symbols, alongside the dates. The same fix turns up again two sections on, in a different form, when the price fetcher hit its own version of “two different things that could look identical from the outside.”

— noted at the time

Section Review Card

no. 03 of 07
Concept:
A cache key has to encode everything that makes a request unique — not just the obvious part.
The catch:
A silent wrong-data collision, caught before it ever ran once.
One thing I actually understood today:
Badges:

04logging

Why log every day, not just the interesting ones

Every trading day gets recorded, not just the ones that seemed worth keeping at the time.

Two ways to record a year: write down a handful of dates that mattered, or log every day and pull whichever view you want out of it later. The first is smaller. The second is the only one that can answer a question nobody thought to ask yet.

Given the actual goal — daily, weekly, monthly, and yearly views of the same portfolio — there was only one honest answer. A day logged in full can always be zoomed out into a week, a month, a year. A handful of remembered dates can never be zoomed back in — whatever happened between two checkpoints simply wasn’t written down, and there’s no recovering it later. With the storage cost already confirmed trivial, there was no real trade-off to weigh: full daily logs, every series, every day, and let any coarser view get built from that afterward rather than fetched separately.

data/ and output/, and which one is kept

Two folders, two different jobs. data/ holds every cached response — regeneratable any time by calling the API again, never worth keeping a permanent record of. output/ holds the one thing that has to survive: the finished numbers, worth keeping for good.

Section Review Card

no. 04 of 07
Concept:
Log every day. Any coarser view can be built from that later — never the other way round.
The split:
data/ regenerates on demand. output/ is the one folder that’s actually kept.
One thing I actually understood today:
Badges:

05freshness

Data that changes after you fetch it

A reading taken today isn’t guaranteed to say the same thing next month — and the price fetcher needed to know that about itself.

Frankfurter’s rates are permanent the moment they’re published — a Tuesday in March means the same number forever. Share prices aren’t like that. yfinance adjusts history retroactively for stock splits and dividend payouts, so a price cached last week can quietly stop matching a fresh pull for the identical date, through no fault of anything breaking.

Two things came out of taking that seriously rather than shrugging past it:

A fetched_at timestamp, in every cached file, not just the filename. The only way to know how current a reading actually is — and the reason a dates-only cache name would have been the wrong fix here too, same shape as the cache-filename bug two sections back, different cause.

A deliberate choice on what price actually gets recorded. AUTO_ADJUST = True means the cached price already includes dividends paid out, not just the raw share price on the day — a total-return series, not a price-only one. The alternative was simpler to build and quietly wrong for this project’s purpose: a client’s actual return includes what they were paid, not only what the share price did. Set once, as a named constant with a comment explaining the choice, rather than left as an unlabelled default nobody would think to question later.

Last reading taken CFR.SW · fetched_at 2026-08-22T21:17:21+00:00

Nothing decided here changes automatically — force_refresh is a flag you choose, not a clock. A reading stays exactly as fresh as it was on the day it was taken, until someone deliberately asks for a new one.

Section Review Card

no. 05 of 07
Concept:
Not every data source ages the same way — know which kind you’ve got before you cache it.
The choice:
Total return, on purpose, with the reasoning written down next to the constant that decides it.
One thing I actually understood today:
Badges:

06calendars

Four calendars that don't line up

Four exchanges, four sets of opening days, and not one of them agrees with the others.

Once both fetchers were pulling real data, the mismatch became impossible to ignore: every exchange runs on its own calendar. A year of data came back as four different row counts.

ExchangeRow countCalendar
ECB (FX)255Publishes on ECB business days
Xetra253German market holidays (ADS.DE, PUM.DE, MBG.DE)
NYSE251US market holidays (PVH)
SIX250Swiss market holidays (CFR.SW)
Market status what was open, on a real day

Conditions for Friday 2 January 2026

  • ECBFX_EUR / FX_GBP / FX_CHFOpen
  • XetraADS.DE PUM.DE MBG.DEOpen
  • NYSEPVHOpen
  • SIXCFR.SWClosed

Three open, one closed. SIX keeps 2 January; nobody else does. One value forward-filled, on one series.

Two ways to reconcile four calendars into one table: only keep the days every exchange had open — clean, but you lose real days where three were trading and one wasn’t — or take every day any exchange was open, and for the closed ones, carry the last known reading forward. The second was the call: forward-fill, keeping every date, understanding exactly what “closed today” means for the numbers on that date.

Forward-filled — carried over from an earlier day
CFR.SW 2 Jan 2026 172.05 carried forward from 30 Dec

SIX was shut, so nothing was fetched. The value is real — it is just not from this day. was_filled holds True here, which is the only thing that can still tell you so.

Verified properly before trusting it, not just built and assumed: a synthetic test truncated one series’ start date deliberately, to check the one branch real data never triggers — dropping leading dates before a series has any value to carry forward. 101 leading dates dropped, zero gaps remaining afterward. The branch works, on evidence, not on faith.

The caveat worth being able to answer

One caveat, worth being able to answer if asked: forward-filling an FX rate on a day the asset’s market was open but the FX market’s data was stale means that day’s move gets attributed entirely to the asset — none of it to currency. Real, small here, worth knowing exactly why.

Section Review Card

no. 06 of 07
Concept:
Forward-fill over inner-join — keep every date, carry the last known reading forward on days an exchange was closed.
The proof:
The hard branch (leading-date drop) tested on purpose, not left unverified.
One thing I actually understood today:
Badges:

07measuring

Measuring what got filled in

What the combined dataset actually looked like once four calendars were laid on top of each other.

The combined dataset: 258 days — genuinely more than any single exchange’s own count, because each one independently contributes days the others didn’t have. Of 2,064 individual values across all eight series, 43 were forward-filled. 2.1% overall.

What combine_forward_fill() printed, per series.
SeriesFilledOf%
FX_CHF32581.2%
FX_EUR32581.2%
FX_GBP32581.2%
ADS.DE62582.3%
PUM.DE62582.3%
CFR.SW92583.5%
PVH72582.7%
MBG.DE62582.3%

CFR.SW carries the most fill — 3.5%, the highest of any series, and exactly what you’d expect: SIX has the shortest calendar of the four, so it’s closed more often than the others are open. Nothing anywhere near distorting — the whole point of measuring it was having a real number to say that with, instead of just assuming it.

That is where the data work stops. What all of it was actually for — splitting each holding’s return into the part the share did and the part the currency did — is Issue 08, The Scorecard.

Section Review Card

no. 07 of 07
Concept:
Measure what you invented, per series, and print it — don’t hand back a clean frame.
The number:
43 of 2,064 values forward-filled. 2.1%, nothing above 3.5%.
One thing I actually understood today:
Badges:

End of the second build dispatch

Two fetchers, one cache bug caught before it ran, and four exchange calendars reconciled into a single 258-day series. Next time: the actual maths — splitting a year’s return into what the asset did and what the currency did, and the rule for the sliver left over once you’ve done both.