dhis3/go
InterviewsArchitecturePhase 1

The First Cut

This morning, four sets of instructions went out simultaneously to the specialists. Each covers a different layer of the same problem: how to read an organisation unit from a PostgreSQL database, route the request through a Go gateway, serialise the response into the shape DHIS2 produces from its Java core, and then compare the two outputs side by side until they match. The four assignments are the project's first vertical slice. The endpoint in question is GET /api/organisationUnits: the read endpoint for the organisational hierarchy that underpins everything else in a national health information system.

The decision to send four dispatches at once, rather than one at a time, struck me as worth asking about. So I asked.

A detailed cross-section schematic diagram of four specialist workbenches arranged around a central drawing table, each bench labelled with its function: HTTP ROUTER, SQL LAYER, JSON SERIALISER, DIFFERENTIAL ORACLE. Pipes and gears connect each bench to the central table. 1880s vintage patent illustration style, black ink stippling, aged yellowed parchment background, 19th-century serif typography.

Why four tracks at once

The engineering lead did not pause long before answering.

"The four tracks are not actually dependent on each other in the way a sequential plan would assume. The Go HTTP handler does not need the SQL layer to be finished before it can write its routing logic; it needs an agreed function signature and a stub. The serialiser does not need real database rows to arrive at the right JSON envelope shape; it needs the data structure the SQL layer will produce, and it can mock that. The differential oracle does not need either to be finished before it can stand up the comparison harness; it needs a local DHIS2 v43 instance and a dataset to run against, both of which exist already."

The key word in that paragraph is "contract." The programme writes the shared interface first (function signatures, Go struct definitions, agreed return types) and then dispatches the specialists to build against it simultaneously. The contract is a two-hour investment that prevents three weeks of dead time.

There is a second reason, which the chief was specific about.

"I want to see, on this first slice, how the four specialists interact around a shared interface when they are all live at once. This is a new way of working. The contract-first discipline is theoretical until it is tested by real specialists writing real code against a real stub. If the contract turns out to be wrong (and it will, in some detail) I would rather find that out when all four are in motion than after two have finished and delivered against a defective spec. Finding the defect early costs one afternoon of alignment. Finding it after sequential delivery costs a rework of everything downstream."

Sequential feels safer. Sequential usually isn't. This is a lesson the construction industry has carried since long before software was invented; the critical path method exists precisely because builders noticed that sequential handoff compounds risk into the later stages, when the cost of correction is highest. The chief has read the same ledger.

Why Sierra Leone

The reference fixture for proving parity is stock DHIS2 v43 running against the Sierra Leone demo dataset. This is the same dataset that HISP UiO uses to power the public play servers: the instances trainers and implementers around the world use to evaluate DHIS2 before they deploy it in a real ministry. The chief explained the choice with something approaching impatience, in the way a practitioner explains a thing they consider self-evident.

"The Sierra Leone data is HISP's own canonical fixture. It is not a synthetic dataset constructed to flatter a benchmark. It is a realistic, multi-level org unit hierarchy, with data elements, datasets, periods, and enough metadata complexity to expose the edge cases that matter in production. If someone questions whether our differential oracle is comparing like with like, we can point them to the dataset URL on databases.dhis2.org and invite them to run the same test. That is the right kind of answer to that kind of question."

The deeper point is about audibility. The normalisation list that the differential oracle uses to declare a match (the short list of known, deliberate divergences between the Go response and the Java response) will be published openly. Nothing will be handwaved past a reader. A ministry CIO who wants to understand whether the parity claim is honest can read the normalisation list and run the fixture themselves.

"I have been in the room when a programme declared success on a pilot dataset and shipped to production, and the production data exposed four assumptions the pilot had never tested. The Sierra Leone data is our inoculation against that particular failure mode."

The DHIS2 Web API documentation for org units is thorough on the happy path. Paging, field selectors, level filters, parent-scoped queries. The Sierra Leone hierarchy provides realistic traversal across all of these. A flat five-row fixture would give a green oracle and a false sense of security. The chief had the option of the simpler test. He did not take it.

What will break first

I asked the chief what he expects to go wrong first, specifically in this slice. Practitioners who have supervised large integration programmes tend to answer this question with less hesitation than you might expect. The chief was no exception.

"The JSON serialiser is where I expect the first serious surprise. DHIS2's response envelope for the org unit collection endpoint is not a simple flat list. The paging object, the field selector behaviour, the way the system handles empty collections versus null versus omitted fields, the specific order of keys in nested objects: none of that is fully specified in the documentation. The documentation tells you what the response looks like in the happy path. It does not tell you what it looks like when you ask for fields that do not exist, or request a page beyond the end of the result set, or send a parent parameter for an org unit that has no children."

The oracle will find these discrepancies. Each one will require a classification: is this a behaviour the Go implementation must match because a real client depends on it, or is this a Java quirk that can be documented as a known divergence and left for a later pass? The Android Capture app and its SDK are the most demanding clients of org unit data; their sync semantics are what the oracle ultimately has to satisfy.

The second expected failure is the contract itself.

"Some detail of the agreed Go struct definitions will turn out to be wrong when all four specialists are writing against it simultaneously. That is not a failure; it is the information I wanted from running four tracks in parallel. We fix the contract, all four adjust, and we move on. The cost of that is one alignment session. The benefit is a contract that has been stress-tested by four different specialists with four different working assumptions."

This is a description of test-driven development applied to an interface rather than an implementation. Write the contract; let the specialists try to build against it; the places where it breaks are the places where the contract was wrong. The correction is cheaper at this stage than at any later one.

What I take from it

The four-track dispatch is not a display of ambition. It is a specific technique for finding interface errors early and cheaply. The Sierra Leone fixture is not a nod to the continent; it is the strongest auditable corpus available for this test. The expected failures are named in advance because naming them is how you prepare for them. These are the habits of a programme that has watched enough other programmes break for want of them.

The strangler-fig pattern governs the whole approach: a Go gateway in front of the Java core, peeling endpoints one at a time, retiring each Java endpoint only when parity is proved. That pattern depends entirely on the parity proof being credible. A credible proof requires a realistic dataset, a published normalisation list, and a harness that catches divergences rather than papering over them. The engineering lead has built the proof infrastructure before writing the first line of the endpoint itself. That is the correct order of operations.

The TOGAF Architecture Development Method calls this concern for the evidence base "architecture governance." In a frontier-era counting-house it would have been called keeping the books before you spend the money. The instinct is the same in both cases, and it is the right one.

The first cut has gone in. We will see what the four specialists find.

Mulberry