dhis3/go
PortraitsInfrastructurePhase 1

The One Who Oils the Gauges

Every workshop I have ever walked through has at least one person whose job the visitors cannot immediately name. He does not swing the hammer that lands the cabinet, and he does not sign the drawings that go to the client. What he does, so far as I can tell, is keep the planes sharp, keep the levels true, and keep the one brass gauge on the back wall reading what it ought to read. The cabinet is still a cabinet without him. It is not, however, a cabinet anyone would trust.

In this workshop his name is Olin, and his gauge is a running copy of the reference software our rewrite is trying to equal. For those just joining us, the project is a BSD-licensed Go reimplementation of DHIS2, the open-source health information platform in use in more than seventy countries. You can read about the opening moves here and the first vertical slice here. This note is about something less glamorous and more necessary: the man who makes sure the yardstick is a yardstick.

A detailed cross-section schematic diagram of a workshop steward's station, showing a man in a leather apron tending a brass pressure gauge mounted on a boiler labelled JAVA REFERENCE, with pipes running off to the right carrying labelled valves FIXTURE, SNAPSHOT, RESET, PARITY. 1880s vintage patent illustration style, black ink stippling, aged yellowed parchment background, 19th-century serif typography.

What the gauge measures

The project's correctness strategy is unromantic and unoriginal. Side by side, in two Docker containers on the same workstation, we run the stock Java DHIS2 and our Go rewrite. We fire the same HTTP request at both. We compare the two responses byte for byte, or where the shape demands it, field for field. The Java one is right by definition; the Go one is correct when they match, and wrong when they do not. Martin Fowler calls this parallel change; the engineers call it a differential oracle. The posture is simple enough that a reader can hold it in one hand. The maintenance of the posture is the trade.

That maintenance is Olin's trade. If the Java instance is unhappy. If it has drifted. If its database has been meddled with. If its version does not match the dataset loaded into it. If any one of these things goes quietly wrong, every comparison the specialists run on top of it goes with it, and the project ships a rewrite that is correct against a yardstick that was itself crooked.

He has, accordingly, three standing jobs.

First job: keep it running

On the morning of the eighth of October I asked Olin to bring up the reference instance from zero. By which I mean: no containers, no database files, no configuration, nothing. He was to build the yardstick in front of me so that I could watch him do it and record the time.

He got as far as starting the container. Then the container stopped. Then it started again, and stopped again. The log said, in Tomcat's politely bureaucratic register, that a file called dhis.conf did not exist at /opt/dhis2/.

This was a modest surprise. The container image expects a configuration file; the compose file we had been using passed database credentials as environment variables, in the modern style, assuming the image would synthesise the configuration file from them. The image does not. It never has. The specialist's documentation records this fact in a corner on page four. Our compose file, written by somebody who had read the first three pages, did not.

Olin wrote the configuration file by hand, with the five lines the image actually needs, and committed a template so the next person to burn their reference down to the ground will not spend a morning rediscovering the omission. The container came up. Postgres came up alongside it. The schema was in place in another sixty seconds. The web interface answered a login request thirty seconds after that.

I recorded the wall-clock, which came to eleven minutes from first command to first HTTP 302. Three hundred and twenty-one of those seconds were the Java core cold-starting, which is approximately the time it takes to boil a kettle twice. Three hundred and eighteen more were the database loading the Sierra Leone demo fixture, which is a 90-megabyte compressed SQL dump published by HISP UiO and the closest thing we have to a shared national dataset the whole community can run on a laptop. The database itself is PostgreSQL with the PostGIS extension, because an organisation unit without a geometry is a half-measured organisation unit. The rest was the two of us watching the gauge.

Eleven minutes is not fast. Eleven minutes, however, is the time the whole yardstick takes to come up, including a national demographic dataset, a spatial database, and a login prompt. For a system that runs ministries, I will take that.

Second job: keep it the right instance

A yardstick is not a yardstick in general. It is a yardstick cut to a particular length, kept in a particular place, and labelled so that whoever picks it up knows what it measures. In the shop that means the running copy must declare, at all times, two coordinates: which version of DHIS2 is executing, and which dataset is loaded.

The dataset matters because DHIS2 ships several demos (Sierra Leone, Trainingland, the climate-and-health scenario, and others) and because the specialists write their tests against whatever happens to be in the database. The version matters because the Hibernate migrations will try to lift an older dump to a newer schema at startup, and some of those migrations are not safe on data loaded this way; a well-known failure mode is a duplicate-row surprise in a startup populator, which produces a dark and uninformative stack trace in a log at four in the afternoon on a Friday.

Olin's defence against this is a text file. In the top of the reference-instance directory there sits a CURRENT-FIXTURE.md recording what is loaded and when; and a CURRENT-VERSION.md recording the image tag the container is pinned to; and a rule, written in capital letters at the top of his runbook, which reads: the DHIS2 image version MUST match the Sierra Leone dump version. If they drift, align them before touching anything else. The specialists cite the fixture file by date in their parity reports, so a historical comparison can be re-executed later against the same ground truth.

This sounds like filing. It is filing. It is also the difference between a yardstick and a stick.

Third job: build the handles

A workshop in which the only way to take the measurements is to type thirty-line docker-compose sequences every time is a workshop that will eventually type one of them wrong. Olin's third job is accordingly to wrap each recurring sequence behind a handle. A short command. A verb. Something the specialists can call from a script, or type at a prompt, without consulting the runbook.

He has built a small command-line tool for this purpose. It has verbs like stack status, which asks every declared container whether it is healthy and prints a single-line answer; snapshot, which pg_dumps the current database into a dated gzip in case somebody is about to run something experimental; stack reset, which drops the database, reloads the baseline fixture, and restarts the core. The ergonomic point is to make the dangerous actions easy to invoke correctly and difficult to invoke by accident. The engineering point is quieter: every hand-run shell sequence erodes; every wrapped verb carries its own documentation and its own error messages, and improves the next time somebody uses it.

I asked him what he had learned from building it. He said he had learned how many little decisions a specialist makes implicitly when running one of these sequences by hand, and how many of those decisions are the wrong one. The handles make the right decision the default. That is the ninety percent of the value. The other ten percent is that nobody has to remember the thirty-line sequence anymore.

What the keeper is for

The three jobs together are not interesting on their own. Any of them would be invisible, and together they add up to about the quietest trade in the workshop. The reason to write them down is not that they are heroic. They are not. The reason is that in most digital-health programmes I have observed, the trade does not exist at all.

What tends to happen instead is this. A consultant arrives with a laptop and a plan. The consultant stands up the reference software, loads a dataset, runs the demonstration, writes the report, invoices, and leaves. The reference instance is left to decay in some corner of a server, slowly drifting away from the version the next consultant will be running against. When the second consultant arrives, the yardstick is already crooked, and nobody has the brass to say so, because nobody has the time to prove it. Twelve consultants later, the yardstick measures nothing, and the measurements are entered into a slide deck anyway.

I do not propose that the solution to this old sector problem is a single steward on a single laptop. It is not. It is a cultural adjustment, cumulative and slow, in which engineering teams stop treating the reference instance as a decorative asset and start treating it as a load-bearing one. Which, in this workshop, is Olin's job in a sentence. He is the one who oils the gauges, because the measurement is what the whole shop is for.

If, late one evening, you catch him sharpening the planes, do not interrupt him. The specialists need those planes in the morning.