Notes

Short ones. A fix, a finding, a link with a reason. Newest first.

Give the model a spec, not a wish

Tried building a small kanban planner with Gemma 4, twice. The first time I described it in a paragraph and got something that looked right and did half of what I meant. The second time I wrote a proper spec first: the entities, what each column can do, drag and drop rules, what the tests must cover. The difference was large. Writing the spec also showed me which parts I had not actually decided yet.

AI training works best in your own repo

Ran a three-day AI training for the engineers. Day one on slides and toy examples got polite interest. Day two, when everyone opened their own repository and tried Claude Code on a real ticket, is when the questions got good. If I run it again, the toy examples go and people bring their own work from the first hour.

A test-writing skill needs a locator rule

The first version of my agent skill for generating Playwright tests produced tests that passed and broke a week later, because it grabbed CSS classes and nth-child selectors. Adding one rule fixed most of it: prefer getByRole with an accessible name, then getByLabel, then getByTestId, and never a class name. A nice side effect is that a page hard to test this way is usually hard to use with a screen reader too.

playwright.dev/docs/locators

Write CLAUDE.md for a colleague who forgets everything

My first CLAUDE.md files were long and full of things any engineer would know. They got better when I imagined a sharp colleague who joins fresh every morning with no memory of yesterday. What do they need that the code does not tell them? How to run it, the commands that are easy to get wrong, the rules we broke once and do not want broken again. Everything else is noise they have to read past.

After a Git migration, keep the old remote read-only

Moving repositories to GitHub Enterprise Cloud went smoothly. What I would repeat: leave the old server up and read-only for a few weeks instead of switching it off. Pipelines, bookmarks and scripts you forgot about fail loudly on a push instead of silently on a missing host, and anyone who needs an old branch can still find it.

Context beats clever prompts

The biggest thing I took from the AI engineering cohort: most bad answers are not prompt problems, they are context problems. The model did not know the naming convention, the folder layout, the one library we are not allowed to use. Writing those down once, in a file the tool reads every time, did more than any amount of rewording the question.

First night with local models

Installed Ollama on the Mac Studio and pulled qwen2.5-coder:14b. Setup took minutes. The surprise was how usable it is for small, closed questions: explain this regex, write the type for this JSON, rename these variables. It falls apart on anything that needs to know the rest of the codebase, which is fair, because it does not. I want to know what a model does with nothing between me and it before I trust a tool built on top of one.

ollama.com

A 401 from Entra ID that was really a version mismatch

An API kept rejecting tokens that looked perfectly valid. The aud claim was the application id GUID, and the API was checking for api://.... The difference comes from the token version: v1 access tokens carry the App ID URI as the audience, v2 tokens carry the client id. Setting accessTokenAcceptedVersion to 2 in the app manifest and accepting the client id as the audience fixed it.

Decode the token before you debug anything else. jwt.ms does it in the browser.

learn.microsoft.com/en-us/entra/identity-platform/access-tokens

With Argo CD, Git is the only way in

Someone fixed a production config with kubectl edit during an incident, which was the right call at the time. Argo CD quietly put it back on the next sync, which was also the right call. With GitOps the change has to land in the repo or it did not happen. We turned on self-heal everywhere after that, and the incident checklist now ends with "open the pull request for whatever you changed by hand."

argo-cd.readthedocs.io/en/stable/user-guide/auto_sync/

Copilot is better at tests than at design

After a few months of the whole team on GitHub Copilot, the pattern is clear to me. It is very good at the boring middle: test cases for a function that already exists, mapping one shape of data to another, the fifth reducer that looks like the other four. It is weak at deciding where a boundary should go. So the advice I give now is to use it freely once the design is settled, and to put it away while you are still deciding what the pieces are.

SLOs made the Monday meeting shorter

Our weekly health review used to be a tour of graphs where every spike needed a story. After we set SLOs in Datadog for the user-facing flows, the question became one line: are we inside the error budget or not? If we are, we move on. If we are not, we look at that flow and nothing else. Same data, far less talking.

sre.google/sre-book/service-level-objectives/

Scan findings are a queue, not a verdict

When Fortify and Black Duck went into every pipeline build, the first reports were long enough that people started ignoring them. What made them useful was a simple rule: the build fails only on new critical findings, and everything else goes into a backlog someone owns. Old findings get worked down on a schedule. A gate that always fails is a gate everyone learns to walk around.

ERR_OSSL_EVP_UNSUPPORTED means it is time for Webpack 5

Moving a project to a newer Node release, the build died with ERR_OSSL_EVP_UNSUPPORTED. Node 17 and later ship with OpenSSL 3, which dropped the MD4 hash that Webpack 4 used for module ids. The quick way out is NODE_OPTIONS=--openssl-legacy-provider, and plenty of answers online stop there.

I used it for one day to unblock a release, then did the real fix: upgrade to Webpack 5. The flag is a note to your future self that something is overdue.

Call createTheme once, outside the component

Another slow leak, this time in a Material UI app. A theme was being built with createTheme() inside a component body, so every render produced a new theme object, every styled child saw a new theme, and style sheets piled up.

Moving createTheme() to module scope, or into a useMemo when it depends on props like the brand, stopped the growth. If a value is expensive and does not change between renders, it should not be created during render.

Split by route, not by component

A Lighthouse pass on a slow page tempted me to wrap everything in React.lazy. That made it worse: dozens of tiny chunks, each its own request, and a flicker of spinners on first load. What worked was splitting at the route level only, plus gzip and long cache headers on the CDN for the chunks that do load. One chunk per page the user can actually land on is a good default. Go smaller only when a profile says so.

Synthetic checks belong in every environment

We had Datadog synthetic tests on production only, which meant they told us about problems at the most expensive moment. I added the same browser tests to dev and test, pointed at each environment with a variable for the base URL. Now a broken config shows up the morning after it is merged instead of the night of the release. The tests cost little to run and they are the closest thing we have to a user clicking through the app every few minutes.

Screen readers read what you wrote, not what you meant

Spent an afternoon with NVDA and VoiceOver on a form that passed every automated check. Two things the scanner never flagged: a clickable div with an aria-label that neither reader announced as a button, and an error message that appeared on screen but was never read out because it was not in a live region.

A real button fixed the first. aria-live="polite" on the error container fixed the second. Automated tools catch maybe the easy half. The other half needs someone to close their eyes and use the page.

Warm up the slot before you swap

Azure App Service slot swaps are supposed to be zero downtime, and they mostly are, but the first requests after a swap were slow because the new instance had never served a real page. App Service can ping a path on the staging slot before it swaps and wait for a good status code:

WEBSITE_SWAP_WARMUP_PING_PATH=/health/ready
WEBSITE_SWAP_WARMUP_PING_STATUSES=200

Point it at an endpoint that touches the database and the caches, not one that just returns 200. Otherwise you are only warming up the health check.

learn.microsoft.com/en-us/azure/app-service/deploy-staging-slots

A MobX store that never let go

Memory kept climbing on a long-lived page. Two heap snapshots a few minutes apart showed components that had unmounted long ago still hanging around. The cause was a singleton store with reaction() calls set up inside components and never disposed, so the store kept a reference to every one of them.

The fix is small: reaction returns a disposer, and it belongs in the effect's cleanup.

useEffect(() => reaction(() => store.filter, load), []);

Returning the disposer from useEffect is enough. The habit I took away: any subscription to something that outlives the component needs an owner for its cleanup.

AZ-305: what actually helped

Passed the Azure Solutions Architect exam this week. The video courses were fine, but the thing that moved my score was reading the Well-Architected Framework pillars and then answering every practice question by naming the pillar it was really asking about. Most questions are a trade-off between cost, reliability and security dressed up as a product choice. Once you see the trade-off, the product name matters less.

learn.microsoft.com/en-us/azure/well-architected/