Reference · /learn/
The one where the numbers finally come out, and then have to be proved
Issue 07 stopped the moment the data was clean — four exchange calendars reconciled onto one, 2.1 % of it forward-filled, and not a single return calculated yet. This picks up exactly there and runs to the end: working out what the holdings actually did, turning five of them into one number, writing a rule that decides when to flag a currency, testing it in two different ways, drawing it, and putting it on the internet. Plus the bug that turned up in already-published code because I sat down to write the next plan.
Where this is going
A return split into the bit the share did and the bit the currency did, five holdings weighted into one portfolio figure that has to reconcile to the penny, a hedging threshold that fires on its own with no model anywhere near it, two layers of tests that fail for two different reasons, a chart built so colour is never the only thing carrying the meaning, and a deployment that corrected two things I had written down wrong about my own server — ending with a mistake that only surfaced because I was documenting what came next.
The calculation
01returns
calculate_returns.py
Why a portfolio can go down in a year when most of the shares in it went up.
Every holding here is priced in a currency that isn’t mine. Adidas trades in euros, Richemont in Swiss francs, PVH in dollars. I think in pounds. So when I ask “what did Adidas do for me this year”, two completely separate things have moved: the share price, in euros, and the euro itself, against the pound. Either one can be up while the other is down.
That’s the whole reason this project exists. A client sees one number and asks why it’s negative when the news says their companies did well. The answer is usually that the number is two answers added together, and nobody separated them.
The FX data comes back based on the dollar — how many euros a dollar buys, how many pounds, how many francs. What I need is euros to pounds. That’s just division:
EUR_to_GBP = FX_GBP / FX_EUR
If a dollar buys 0.857 euros and 0.733 pounds, then a euro buys 0.733 / 0.857 = 0.855 pounds. The thing worth noticing is that both numbers come from the same daily snapshot, so they can’t disagree with each other. Pulling the euro rate from one source and the pound rate from another would leave a gap that isn’t a real market movement, just two providers rounding differently.
The obvious split is “asset” and “currency”. But they multiply rather than add — a share that gains 10 % in a currency that gains 10 % doesn’t give you 20 %, it gives you 21 %. That extra 1 % is real money and it belongs to neither leg on its own:
(1 + total) = (1 + local) × (1 + fx) total = local + fx + cross_term
I report that leftover as its own cross-term rather than quietly folding it into the currency figure. Folding it in is easier to explain and slightly dishonest: it makes the currency look like it did more than it did. This year the cross-term stayed under 0.35 % for every holding, because neither leg was large. In a year where a currency moves 20 % it stops being a rounding detail.
Against the previous trading day, 21 August. The share rose half a percent in euros; the euro slipped a little against the pound and took some of that back. In pounds the holding gained 0.38 % — less than the share did, and neither figure is wrong.
These are the numbers baked into the page. With JavaScript on, it fetches today’s instead.
Why the rate is fetched by your browser and the share price isn’t
The panel above gets its exchange rates straight from the FX service, and its share
prices from alice.charliehub.net. That split isn’t a design choice,
it’s a rule the browser enforces. When a page asks a different website for
data, that website has to say it’s allowed. The FX service answers
access-control-allow-origin: *, meaning anyone may ask. Yahoo, where the
prices come from, says nothing — so the browser fetches the data and then throws
it away without letting the page see it.
That rule only binds browsers. A server isn’t a browser, so my own server can
ask Yahoo quite happily, and this page can ask my server, because they’re the
same website. That’s why yfinance works in Python and could never
work here, and why a “CORS error” is never fixed in the page —
it’s fixed by giving the page somewhere of its own to ask.
There is no such thing as the dollar price of a dollar
First run, straight to KeyError: 'FX_USD missing'. PVH is listed in
dollars, and the FX data is based on the dollar — so there is no
“USD” column in it. There never could be. A dollar is worth one dollar by
definition, and nobody publishes that.
The fix is a constant series of 1.0 for whichever currency the FX pull is based on,
behind an explicit fx_base="USD" argument rather than an assumption buried
in the code. It has to be explicit because it must always match what the fetcher was
actually called with — change the base and forget this, and every dollar figure
silently goes wrong instead of crashing.
The function checks its own work: it asserts that total equals
local + fx + cross_term. That assertion can never fail. It’s an
algebraic identity — rearrange the definitions and it’s the same statement on
both sides. It confirms the arithmetic and is completely blind to a wrong input.
So I checked it two more ways that can fail:
fx_return = 0.0 and
total == local, which is exactly right: hold something in your own
currency and you have no currency exposure to it. With EUR, the three euro
holdings collapsed to zero the same way and the Swiss and American ones stayed live.
That proves the function isn’t quietly hard-wired to pounds.The results, once it ran: currency movement was small and remarkably uniform — sterling weakened about 0.99 % against the euro, 0.68 % against the franc and 1.76 % against the dollar. The interesting spread was entirely in the shares themselves: Puma +33.7 %, Mercedes-Benz −13.1 %, a 47-point gap between two holdings in the same country and the same currency.
02portfolio
build_portfolio.py
Five separate answers are not a portfolio. Combining them takes a decision I had no information for.
Section 01 gives a return per holding. To get one number for the whole portfolio I need to know how much of each is held — and I don’t. I have no idea how George Russell’s money is actually allocated, and inventing plausible-looking weights would be claiming knowledge I don’t have.
So: equal weights, 20 % each. It’s simple, it’s transparent, and it’s honest about what is and isn’t known. If someone asks why, the answer is a good one rather than an awkward one.
The portfolio is £20,000,000, and that figure is illustrative — not a claim about anyone’s real finances. I picked a large round number deliberately so the resulting profit and loss is something you can say out loud. “Down 1 % on currency” is easy to wave away. “Currency cost this book £199,847” is not.
| Holding | Ccy | Weight | Local | FX | Cross | Total | P&L |
|---|---|---|---|---|---|---|---|
| ADS.DE | EUR | 20% | −6.87% | −1.06% | 0.07% | −7.86% | −£314,284 |
| PUM.DE | EUR | 20% | 15.14% | −1.06% | −0.16% | 13.92% | +£556,642 |
| CFR.SW | CHF | 20% | 40.22% | −0.73% | −0.29% | 39.20% | +£1,568,001 |
| PVH | USD | 20% | −3.62% | −1.08% | 0.04% | −4.66% | −£186,430 |
| MBG.DE | EUR | 20% | −11.37% | −1.06% | 0.12% | −12.32% | −£492,671 |
| Portfolio | GBP | 100% | 6.70% | −1.00% | −0.04% | 5.66% | +£1,131,258 |
Across holdings the parts add up; inside one holding they multiply
This is the subtlest thing in the whole project and it took a while to see. Within a
single holding, the local and currency legs compound — that’s the
identity from section 01. But the portfolio’s cross-term is the weighted
average of the five individual cross-terms. It is not
(1 + portfolio_local) × (1 + portfolio_fx) − 1.
The reason is a rule that holds everywhere in maths, not just here: the average of a product is not the product of the averages. Average the legs first and you have thrown away which holding each leg belonged to — and the interaction happens inside each holding, not between them. Getting this right is what makes the bottom row of that table reconcile exactly, both as a percentage and in pounds.
add the parts 6.70% + (−1.00%) + (−0.04%) = 5.66%
compound the legs (1 + 6.70%) × (1 − 1.00%) − 1 = 5.63%
At portfolio level the two routes disagree. Only adding the parts ties back to the money, because the cross-term belongs inside each holding and averaging the legs first throws that away. Pick a single holding above and both routes agree exactly.
normalise_weights() takes weights as percentages, as unit counts, or as
market values, and turns any of them into fractions that sum to one. It refuses an
unknown ticker, a missing holding, a negative weight, a NaN, or a set that sums to zero.
Refusing loudly matters more than it sounds: every one of those would otherwise produce a
believable-looking number that is quietly wrong.
One simplification, stated rather than hidden: the weights are fixed for the whole period. Real weights drift as prices move — a holding that gains 40 % becomes a bigger share of the portfolio without anyone buying more of it. Capturing that needs the day-by-day series rather than one figure for the year, which is on the list for version 2.
The judgement
03flags
hedging_flags.py
The one part of this project where I deliberately kept the model out, and wrote down why in the file itself.
A hedging flag says: this holding’s currency exposure has moved past what this client signed up for, and somebody should look at it. It’s a small piece of code and a slightly serious one, because it’s the bit a person might actually act on.
The rule is a threshold. This client’s risk band is “growth”, and the currency limit is −1.5 %. Any holding whose currency leg falls below that gets flagged, with a message naming the holding, the figure, how far past the line it went, and three things that could be done about it.
-0.015, not
-1.5. Passing -1.5 raises an error saying it would mean
−150 %. Percent-versus-decimal is the easiest possible mistake to make here
and the hardest to notice, because both look reasonable. A positive threshold raises
too — this is a downside limit by definition.< instead of <=, and there is a test that sits a
holding precisely on the boundary to prove it.0 of 5 breaching at −1.50%. Closest is PVH at −1.08%, which is 0.42 points clear of the line.
Against the real figures, nothing breaches at −1.5 %. That is the honest result and I kept it. Move the line to −1.0 % and four of the five light up — which is in the test output deliberately, as proof the logic fires, without having to invent a currency crisis to demonstrate it.
Nothing here calls a model, and the file says so
The flag message is built by string formatting. There is no model call, and nothing is imported that could make one. The reason is written into the module’s own docstring so it can’t be quietly forgotten: a flag is a compliance artefact somebody may have to justify a year later. “The system flagged it because the currency leg was 0.3 points past the limit” is a defensible sentence. “The system flagged it and here is roughly what it said at the time” is not.
Language models come in later, for the client-facing commentary, and even there the rule is set in advance: the model only ever sees numbers that have already been calculated, and is forbidden from producing any figure not in what it was handed.
An empty result should be the same shape as a full one
When nothing breaches, the check returns an empty table with the same columns
— not None, and not a special case. Anything reading it can loop over
zero rows without knowing it got zero rows. Returning None for “no
results” pushes an if statement into every single caller, and the day
somebody forgets it is the day the dashboard crashes because the market was calm.
04proof
write_output.py · test_flags.py
Two test layers that fail for two different reasons, and the reason neither one checks a market number.
Everything so far ends in one file. build_summary() runs the whole chain
— returns, then portfolio, then the hedging check — and assembles a single
dictionary: who the client is, the risk band, the base currency, the date range, every
holding with all its components and its profit and loss, the portfolio totals, the
hedging results, and a generated_at timestamp. It goes to
output/portfolio_summary.json.
That file is committed to git, unlike the cached data underneath it. Issue 07
covers why those two folders are treated differently
(what belongs in a cache filename); the short version
is that data/ can be rebuilt by asking the API again, and this can’t
— it’s the finished thing a dashboard reads.
The first layer uses a small table I wrote by hand: four holdings, one breaching badly, one breaching by a hair, one sitting exactly on the limit, and one comfortably safe. Because I chose the inputs, I know the exact output, so the test compares the generated message character for character.
That sounds excessive until you notice what it buys: reword the message and a test goes red. Nobody can casually reword a compliance-facing sentence without being told they just did. Which is exactly what happened later — see section 06.
The second layer runs the real pipeline on cached data, and asserts nothing about the values. It checks the shape: one flag per breaching holding, all six fields present in each, the per-holding booleans agreeing with the flags list, the closest-to-the-line calculation correct in both directions, and a full round-trip through JSON to prove no NumPy types leak into the output.
Never assert a market number in a test
It is very tempting to write assert portfolio_return == 5.66. Don’t.
That test passes today, fails tomorrow when the market moves, and fails every day after
that — and what it actually teaches you is to ignore a red test. A test
suite you have learned to ignore is worse than no test suite, because it looks like
safety.
Instead layer two derives its threshold from the data it just loaded — just above whatever the worst currency figure happens to be — which forces at least one breach no matter what the market did. The test stays meaningful on a calm day and on a chaotic one.
A test that has only ever passed hasn’t been shown to work. So I broke the message on purpose and watched the character-for-character check go red, then put it back. Thirty-four checks across the two layers, and I know at least one of them has teeth because I saw it bite.
The public side
05chart
dashboard.html
Roughly one man in twelve can’t reliably separate red from green, so a chart that says it in colour alone says it to eleven of them.
The dashboard is a single static HTML file. It fetches the JSON from section 04 and draws a header, three summary cards, the holdings table, a bar chart and the hedging status. No backend — the numbers were already calculated, so there is nothing left to compute at the moment somebody looks.
The chart’s job is “above or below zero”, so the sign is what colour should carry. The instinct is green for up and red for down. I used blue and red instead, and the panel below is why.
Blue and red, as shipped, seen with typical colour vision. Try the other three combinations.
What actually happens is more interesting than “green and red look the same”, which is the version everyone repeats. Under deuteranopia both red and green collapse towards the same yellow-olive axis. They don’t become identical — but the only thing left separating them is lightness, which is a property nobody chose on purpose. With the green and red most spreadsheets reach for by default, that lightness gap shrinks from 24.9 to 8.7 and the overall colour difference drops by nearly half. Whether your chart survives becomes luck.
Blue and red keep a genuine difference in hue: the blue goes violet, the red goes olive. Measured, the two are further apart under deuteranopia than they are with typical vision. That’s why blue and red is the standard recommendation for a chart built around a baseline, and it isn’t a matter of taste.
Colour is never the only thing carrying the meaning
Even with a palette that survives, colour is never asked to work alone here. The bar direction already carries the sign. Every bar is labelled with its own figure. The table underneath repeats all the same numbers independently. And the status badges pair an icon with a word — “⚠ Breaching”, “✓ Within limit” — never a bare coloured dot.
Turn the whole chart greyscale and it still reads. That is the actual test, and it’s a much easier standard to hit than picking a clever palette.
Two smaller decisions in the same spirit. There is no legend, because there is only one measure and it’s named in the heading — a legend for a single series is furniture. And dark mode is its own palette, checked separately, not the light one inverted; inverting a palette moves every colour somewhere its contrast was never measured.
The server this was built on has no screen. So the page was verified in pieces: serve it over HTTP and confirm the page, the JSON and the chart library all return 200; then run the page’s own render function directly in Node against the real JSON — 24 checks — and 8 more against a deliberately doctored file where a holding does breach, confirming the flag blocks appear only when they should.
What that cannot check is what it looks like: spacing, wrapping, labels colliding. I wrote that limitation down rather than implying the page was verified, and then looked at it properly in a real browser over an SSH tunnel, where it rendered correctly and matched every number in the pipeline.
06shipping
deploy.sh
Two things I had written down about my own server were wrong, and I only found out by trying to use them.
The dashboard worked on the machine it was built on. Getting it somewhere other people could see it had an obvious-looking answer — open the port — and a better one. Opening a port would mean asking for a change to the router so the whole internet could reach a plain development server with no authentication on it. There was already a perfectly good website running on this box with a certificate and a proper front door, so the dashboard went there instead. No new exposure, no infrastructure change, no conversation needed.
hello-web.service. That unit is dead and disabled — it was replaced
months ago. The real one is guestbook.service, a Python app that answers
its own API routes first and then serves the site files for everything else. I had the
folder right and the service name wrong, which is the kind of error that survives
indefinitely until you try to restart something.build.py on that server and it was
tempting, but it belongs to a completely unrelated project (a travel itinerary
generator). Running it would have done nothing for the dashboard and republished
somebody’s holiday cards. Issue 03 covers how the serving actually works:
how this gets to be on the internet.Two small things came out of doing it. The bare folder address didn’t work at
first, because a directory only resolves if it contains an index.html. And
copying files by hand is how a published site quietly drifts away from the code that
produced it, so it became a script:
$ cp dashboard.html index.html "$SITE/portfolio/" $ cp output/portfolio_summary.json "$SITE/portfolio/output/"
The relative path is preserved on purpose, so the page’s
fetch("output/portfolio_summary.json") resolves in exactly the same way in
both places and no code has to change between the repository and the live site.
Advice that was published, plausible, and impossible
The flag message offered four things a manager could do about a currency breach. One of them was “switch to a hedged share class”. A hedged share class is a currency-protected version of a fund — something the fund provider offers. Every holding here is a direct share in a company. There is no hedged share class of Adidas stock. There cannot be.
It was live, it was public, and it read as advice. I didn’t find it by testing. I found it while writing up the plan for version 2, because describing what the flag did forced me to read the sentence as a reader rather than as its author. The option came out, three valid ones stayed — hedge with a forward, reduce the position, or accept and monitor — and the character-for-character test from section 04 went red exactly as designed, which is the first time that test earned its keep.
I want to be straight about that one, because it is easy to write up as a neat anecdote. It was a real error in shipped, public-facing text, and the thing that caught it was not a clever test or a review. It was writing an explanation carefully enough that the sentence had to make sense.
Where this goes next
Version 1 is finished: fetch, clean, decompose, aggregate, flag, write, draw, deploy.
Version 2 is written down and deliberately not started — a real backend under
/api/portfolio/, somewhere to store state, two roles with proper
authentication, and the big one: a transaction ledger, so the portfolio is derived from
what was actually held on each day rather than from a fixed set of weights. The data
layer for that already exists. Section 01’s function is handed a full daily series
and throws all but the first and last row away.
Three things deliberately do not change: the hedging flags stay deterministic, the decomposition identity stays as it is, and the split between cached data and committed output stays where Issue 07 put it.