Why we updated within a day of the release
Nimiq published v2.2.0 on 17 September 2026. The release is backwards compatible with the 2.x line, which means updating is our choice of moment rather than a network requirement. We chose to do it the next morning, for one reason that matters to our stakers: one of the fixes stops a validator from failing block production when a block contains a transaction that cannot be applied. A validator that can be stalled by a crafted transaction is a validator that can miss its slots, and missing slots is the one thing a staker actually feels. The rest of the release (WebSocket RPC, the macro block separation fix, the testnet history sync fix, a refreshed mainnet checkpoint) is upkeep. That one is uptime.
Before touching anything: the rollback plan
The rule we set: nothing gets stopped until everything needed to undo it exists. Three things were staged:
- The old image, tagged. The running 2.1.0 image was retagged
2.1.0-rollbacklocally, so a rollback is a one line container change. - The data directory, snapshotted. A copy of the node's data volume (529 MB) was made before the stop, so even a bad upgrade that writes new format data cannot force a full resync from scratch.
- The container spec, exported. The full
docker inspectof the working container was saved to disk, so the replacement could be created with exactly the same mounts, ports and restart policy.
This is the part of the procedure that has no upside when it goes right and saves the day when it does not. It costs a few minutes and 529 MB of disk.
The swap, minute by minute
Everything below is from the node's own logs, timestamps included.
| Time (UTC, 18 Sep) | Step | Evidence |
|---|---|---|
| 03:42:47 | Old container stopped, new container from the 2.2.0 image started with the same data volume | docker inspect, container StartedAt |
| 03:42:58 | Consensus established, 51 peers, head at #61,898,524 | node log line "Consensus: established" |
| 03:45:59 | First block produced on 2.2.0: micro block #61,898,709 | node log line "Our turn, producing micro block" |
| 04:27:02 | Restake bot's first cycle after the swap, distributing normally | restake ledger, 49.92 NIM split across stakers |
| 04:28 | Steady state: 92 peers, head advancing, zero penalties | node log + RPC head query |
Total downtime, container stop to consensus: about 11 seconds. The swap was deliberately scheduled in a gap between our production slots; had a slot fallen inside those 11 seconds, the protocol treats a missed slot as an unlucky lottery miss rather than a fault, but we would rather not find out.
The checks that said it worked
A node that has restarted and a node that is producing are different claims. Five checks separated them:
- Consensus established with peers climbing back to the low nineties, matching the pre-update level.
- Blocks produced: the first "Our turn" line at 03:45:59, and a steady stream after. Production on 2.2.0 is the whole point of the swap.
- Zero panics and zero checkpoint mismatches in the swap window. A version jump that disagrees with the chain's state shows up here first.
- RPC sanity: the JSON-RPC server answered a head query inside the new container.
- Downstream services: the restake bot and the uptime monitor both came through the swap without a single missed cycle.
What 2.2.0 actually changes
The release notes, in the order a validator operator cares about:
- Cut unappliable transactions from a block instead of failing production. The uptime fix: a block containing an impossible transaction no longer stalls the producer, the transaction is dropped and the block continues.
- Macro blocks adhere to the 1 second block separation time. Keeps timing consistent at epoch boundaries, the moments our slot math cares most about.
- WebSocket support on the RPC server. Relevant to our tooling: fewer polling round trips for the same data.
- Prevent forged address notifications from stalling the web client. A web client hardening, relevant to anything building on Nimiq's browser stack.
- Refreshed mainnet checkpoint (election block 60,652,800) and the testnet history sync fix: faster starts for new and syncing nodes.
The part we did not use, and why we staged it anyway
The rollback image and the data snapshot are still on disk. That is deliberate: they cost nothing to keep, they turn a bad update from an hours long resync into a one command revert, and deleting them saves nothing that matters. The update worked, so they stay until the next update retires them.
For other validator operators
The short version of the procedure: tag the old image, snapshot the data volume, export the container spec, pull the new image, stop and remove the old container, create the new one with the same spec, then watch for the consensus line before you believe anything. Schedule the stop in a gap between your production slots, and check for your first produced block before you call it done. The whole thing, done slowly, is a five minute job with 11 seconds of downtime.