A test suite can be green while a financial system remains capable of losing, duplicating or misrepresenting economic state.
This is not an argument against testing. Tests are indispensable. The problem is the conclusion drawn from them. Most tests validate selected examples under controlled conditions. Financial correctness is a property of the whole lifecycle, including retries, concurrency, partial failure, external inconsistency, correction and recovery.
The gap between those two ideas is where many serious defects live.
Tests prove examples; financial systems depend on invariants
A typical test says: given this input and this initial state, the system produces this expected output. That is useful evidence for the path being exercised.
An invariant says: this condition must remain true across every valid path and every relevant failure mode.
Examples of financial invariants include:
- an economic action has one effect even if the initiating request is repeated;
- assets are neither created nor lost across internal representations;
- every posted balance can be reconciled to authoritative movements;
- a transaction cannot be simultaneously final and reversible under conflicting rules;
- a correction preserves audit history rather than rewriting it invisibly;
- custody, ledger and external settlement states cannot drift without detection;
- permissions and approvals in force at execution are provable after the fact.
A suite may contain thousands of tests without expressing these properties explicitly. Coverage is not the same as authority.
A passing test confirms that a chosen example behaved as expected. Financial correctness requires confidence that no allowed execution path can violate the economic truth of the system.
Local correctness can still produce global error
Distributed financial workflows often consist of individually reasonable steps. A service validates a request, another reserves funds, another sends an instruction, and a final component records completion. Every component may pass its own tests.
The failure can exist in the relationship between them:
- the external instruction succeeds, but the response times out;
- a retry creates a second economic action;
- a message is processed before the state on which it depends becomes visible;
- two workers make valid decisions from stale but different snapshots;
- one service treats “submitted” as final while another treats only external settlement as final;
- a compensating action reverses an internal record but not the external effect;
- a manual intervention repairs the customer view without restoring reconciliation.
No single component is obviously broken. The system as a whole has an ambiguous financial state.
This is why correctness cannot be delegated entirely to unit and integration tests around service boundaries. The economic lifecycle needs its own model.
Idempotency is an economic property, not an HTTP feature
Many systems label an endpoint or consumer “idempotent” because a repeated identifier returns the same response or avoids creating a duplicate database row. That may be insufficient.
The relevant question is whether repeated, delayed or reordered execution can produce more than one economic effect across all participating systems.
Consider a transfer request. The local database may reject duplicate rows while an external provider has already accepted two instructions under different retry identifiers. Alternatively, the provider may execute once while the local system retries indefinitely because it cannot establish the result. Technical deduplication does not resolve the business ambiguity.
Economic idempotency requires a stable identity for the intended action, evidence of external outcome and a recovery process that can distinguish “not executed”, “executed once” and “state unknown”.
Concurrency exposes assumptions hidden by sequential tests
Financial workflows frequently read state, make a decision and write a result. Sequential tests can show correct behaviour for each operation while missing the possibility that two operations observe the same precondition.
Examples include:
- two withdrawals both see sufficient available balance;
- two liquidation workers act on the same collateral position;
- an approval and a cancellation cross in flight;
- pricing or fee rules change between authorization and posting;
- a reconciliation job closes a period while late activity is still arriving.
The solution is not simply “more locks”. Correct control may involve transactional boundaries, reservations, version checks, serialisation by economic key, append-only ledgers or explicit state machines. The right mechanism follows the invariant.
Testing should therefore include concurrent schedules and adversarial ordering, not only more input combinations.
Reconciliation is part of correctness, not a reporting feature
When a system interacts with banks, custodians, exchanges or payment providers, it cannot control every participant. External systems may delay, aggregate, reverse or correct activity according to their own semantics.
A reliable architecture assumes that divergence can occur and provides a way to detect and explain it.
Reconciliation should answer:
- Which internal record corresponds to which external movement?
- What is the authoritative amount, currency, asset and effective time?
- Which differences are expected timing gaps, and which represent defects?
- Can every exception be classified and resolved without destroying evidence?
- Does reconciliation cover the entire population or only successfully linked records?
- Can the organisation prove that an empty exception report means agreement rather than missing data?
A report generated from the same mistaken source as the operational ledger is not independent evidence. Good reconciliation creates a separate control path.
Recovery must restore business truth
Operational runbooks often focus on restarting processes, replaying messages or moving a workflow out of an error queue. These actions restore execution, but not necessarily correctness.
A recovery procedure should establish final economic state before applying a change. It should preserve the evidence used, identify whether customer-visible and external records agree, and make the correction itself auditable.
The important distinction is:
- process recovery: the software continues running;
- business recovery: the intended economic obligation is correctly represented and completed.
A system that restarts cleanly after an outage can still carry unresolved financial drift.
Stronger evidence comes from several complementary controls
No single technique proves a complex financial system correct. Confidence comes from layers that challenge different failure classes:
- unit and integration tests for local rules and contracts;
- state-machine tests for allowed lifecycle transitions;
- property-based tests for invariants across broad input spaces;
- concurrency and fault-injection tests for ordering and partial failure;
- deterministic replay for historically observed workflows;
- reconciliation against independent external or accounting evidence;
- runtime assertions and alerts expressed in business terms;
- immutable audit records for approvals, changes and corrections;
- incident review that updates the model, not only the immediate patch.
The value lies in their independence. Several checks derived from the same assumption can all pass together while remaining wrong.
What leadership should ask when the suite is green
A green suite should lead to better questions, not premature certainty:
- Which financial invariants are written down, and where are they enforced?
- Which paths can create an external effect before internal finality is known?
- How are retries identified across service and provider boundaries?
- What happens when events are duplicated, delayed or reordered?
- Can every balance and position be reconstructed from authoritative movements?
- How is unexplained divergence surfaced and owned?
- What evidence proves that recovery restored economic correctness?
- Which critical assumptions are checked independently of the implementation that relies on them?
These questions connect engineering controls to the responsibility carried by the system.
The decision that matters
Passing tests is evidence that the system behaves correctly in the cases the organisation has chosen to specify. Financial correctness requires a stronger claim: that critical economic properties survive the cases nobody intended to run.
The practical goal is not mathematical perfection. It is control. Important invariants should be explicit, failure should become visible before it becomes financial drift, and recovery should restore business truth rather than merely clear technical errors.
A green suite is a good starting point. It is not the final proof.
The appropriate conclusion follows the system, the evidence and the decision — not a pre-packaged diagnosis.
Start with the problem →