Engineering note · laytime

September 2026

Golden tests for laytime: why the expected value is never updated

A laytime calculation is a claim for money. Here is how we made ours reproducible, disputable and boring, and why the engine is now open source.

Every voyage charter ends with a small argument about time. The charterer had a number of days to load and discharge the cargo; if they used more, they owe demurrage; if less, the owner owes despatch. On a 10,000 DWT coaster in the Black Sea, a day of demurrage is a few thousand dollars, which is a large share of the voyage result. The calculation is done from a Statement of Facts, a timeline of what happened in port, against the counting rules of the charter party.

In most small owners' offices this is a spreadsheet, rebuilt by hand for each voyage, argued over by email. When we wrote the laytime module for MarOps, a back-office application for exactly those offices, we had three requirements that a spreadsheet cannot meet:

  1. The same input must give the same answer six months later, when the claim is disputed.
  2. Every deduction must be a record with a start and an end, never "minus eight hours".
  3. A change to the code must not be able to quietly change a number that has already been sent to a counterparty.

This post is about the third one, but the first two shape it.

The basis is the evidence

The engine is a single pure function:

result = compute_laytime(basis)

basis is one JSON document holding everything the calculation needs: the charter-party counting rules, the demurrage rate, and one entry per port with NOR, completion time, the allowance and the list of interruptions. It contains no references to anything else, no database ids, no clock reads. result is a JSON document with the used days, demurrage and despatch, in Decimal strings.

MarOps stores the basis in a JSONB column next to the result. When a counterparty disputes a statement, the office does not reconstruct anything; it replays the stored basis and gets the stored result. If the counterparty argues that a Sunday should not have counted, the office changes one flag in a copy of the basis, replays it, and both sides see exactly what that argument is worth. The argument becomes a diff between two JSON documents.

That only works if the function really is pure, and if its behaviour does not drift. Which brings us to the tests.

What a golden test is here

The test file opens with this:

Every expected value was verified by hand against the charter-party rule it exercises. If the engine changes and one of these fails, the engine is wrong; these tests are never "fixed" by updating the expected value.

Each test is one rule, one small statement of facts, and the number a person worked out on paper before the code existed. For example, once on demurrage, always on demurrage:

def test_golden_once_on_demurrage():
    """After demurrage starts neither rain nor Sunday is deducted."""
    rain_after = {"from": "2026-03-04T10:00", "to": "2026-03-04T16:00", "reason": "rain"}
    result = compute_laytime(
        basis(
            {"laytime_basis": "SHEX"},
            [port(completed="2026-03-09T08:00",
                  allowed={"kind": "fixed_days", "days": "1"},
                  interruptions=[rain_after])],
        )
    )
    assert result["ports"][0]["demurrage_start"] == "2026-03-03T08:00"
    assert result["demurrage_days"] == "6.00000"
    assert result["demurrage_amount"] == "48000.00"

Laytime starts Monday 2 March at 08:00, one day is allowed, so demurrage starts Tuesday 3 March at 08:00. Cargo completes the following Monday. Between those two moments there is a six-hour rain stoppage and a full Sunday. Under SHEX both would be deducted before demurrage; after it, neither is. Six full days at 8,000 a day: 48,000. A person can check every line of that with a calendar and a pencil.

The companion test moves the same rain stoppage to before demurrage starts and asserts that it now shifts the demurrage start by six hours and reduces the bill to 46,000. A third test flips the charter-party exception flag and shows the rain being deducted after demurrage. Three tests, one rule and its two edges.

Why "never update the expected value" is the whole point

The usual life of a test suite: a refactor changes an output, the test fails, the developer looks at the new output, decides it looks fine, and updates the assertion. For most code that is reasonable. For this code it is the exact failure we are trying to prevent.

If an engine change turns 48,000 into 46,000 and someone updates the test, then every statement already sent out with 48,000 on it is now contradicted by the office's own software. The counterparty's lawyer will find that.

So the policy is asymmetric. A golden test failing means one of two things:

  • the engine is wrong, and the code is fixed until the test passes, or
  • the rule itself has changed, in which case the old test keeps its number under the old rule (which becomes a flag in the basis), and a new test is added with a new hand-verified number under the new rule.

Either way, an existing basis replays to its existing result. Old statements stay true. The basis carries the rule, so old and new can coexist.

What the rules had to be

Getting to hand-verifiable examples forced us to make each rule explicit as a field in the basis rather than an assumption in the code. The set that survived:

  • SHINC counts everything; SHEX and WWD skip Sundays and a list of holidays. Weather is not inferred; a weather stoppage is an interruption record. So WWD and SHEX are the same to the engine, and the difference is in what the office records.
  • unless_used: excepted days count anyway, and idle time on them is deducted through interruption records like any other.
  • nor_to_start_hours: the gap between NOR acceptance and the start of counting.
  • Allowance as fixed days or as cargo quantity divided by a daily rate.
  • Half or full despatch, on working time saved or all time saved. The second is the subtle one under SHEX: the unused allowance has to be projected across the calendar, Sundays included, to find when it would have run out. There is a golden test where the same two days early are worth four days on working time and five on all time.
  • Reversible laytime: one pooled ledger across ports, so time saved at loading carries to discharge. The test shows the same two port calls giving half a day of demurrage pooled and one and a half days plus a day of despatch when settled separately. Same facts, two numbers, both correct under their own clause.
  • Once on demurrage, always on demurrage, with an explicit opt-out flag for charter parties that say otherwise.

Everything else, notably charter-party text interpretation and the parsing of the Statement of Facts, is deliberately outside the function. The engine takes facts and rules; deciding what the facts are is a human's job.

Two implementation choices that made the tests possible

Whole seconds and Decimal. Time is counted in integer seconds and converted to days only for output, as a five-decimal string rounded half-up. Money is Decimal at two places, from a per-day rate pro rata by seconds. Nothing is a float, so 0.1 plus 0.2 never appears, and a result can be compared with == against a string a person wrote down.

Intervals are merged before counting. Interruptions and excepted days are both intervals on the port-local timeline. They are clipped to the laytime window, merged, and the countable time is what is left. A rain stoppage that overlaps a Sunday is therefore not deducted twice, and overlapping interruption records from a messy SOF do not inflate the deduction. The merge is a few lines and it is where most hand-calculation errors in spreadsheets come from.

Open-sourcing the engine

The engine has no dependencies and reads no state, so we extracted it into a standalone package, laytime, under the MIT licence. The golden tests went with it unchanged.

The reasoning is the same as for the stored basis. A laytime statement is a claim against someone else. If that someone can install the package, feed it the basis printed on the statement, and get the same number, the statement is stronger, not weaker. There is nothing proprietary in how Sundays are skipped; the value of the product is in capturing the facts quickly and keeping them.

pip install laytime
laytime basis.json

What we would tell a small office

Keep the basis. Whatever tool you use, the input that produced a statement should be stored, unmodified, beside the statement. If you can replay it, you can defend it. If your tool cannot show you the exact input it used, it is not a calculator, it is an opinion.


The engine: github.com/oktaybobus/laytime. MarOps case study: zdeck.tech/case-study-marops.html.