ESP32 OTA updates for a fleet: A/B partitions, rollback, and what happens when an update fails

By Synacl · 10 min read · Published · Updated

Updating one ESP32 over the air is a weekend tutorial. Updating forty of them inside panels you can't easily reach is a different job, with different questions: what happens to a board that loses Wi-Fi halfway through the download, which version is safe to send, and how you get back to the previous one when a release turns out to be wrong. This guide explains how ESP32 OTA works underneath, what it protects you from and what it doesn't, and how to roll an update across a fleet without a site visit. The last sections show how the Synacl gateway firmware does it, limits included.

How ESP32 OTA works: two slots and a pointer

An ESP32 that can update itself has two application partitions in flash, ota_0 and ota_1, plus a small otadata partition that records which one to boot. The running firmware never overwrites itself. An update runs like this:

  1. The running app downloads the new image into the other slot.
  2. The image is verified: header, checksum and, on normal builds, a SHA-256 digest appended to the image.
  3. Only if that passes is otadata rewritten to point at the new slot.
  4. The chip reboots, and the bootloader starts whatever otadata points at.

Everything outside the two app slots is left alone: the bootloader, the partition table, the NVS partition holding Wi-Fi credentials and settings, and any data partitions. That is why an OTA update keeps a device's configuration and a USB flash with "erase all" ticked doesn't.

The Synacl gateway's partition table is a typical layout for a 4 MB board:

Partition Type Offset Size Holds
nvs data 0x9000 20 KB Wi-Fi, broker and account settings
otadata data 0xE000 8 KB Which app slot to boot
app0 ota_0 0x10000 1.81 MB Firmware, slot A
app1 ota_1 0x1E0000 1.81 MB Firmware, slot B
spiffs data 0x3B0000 320 KB Readings buffered during an outage

The slot size caps the firmware at about 1.8 MB on a 4 MB board (the gateway build is about 1.4 MB today). And there is no third slot: the only earlier version the device itself can fall back to is whatever sits in the other slot, until the next update overwrites it.

What happens when an update fails

Because the pointer only moves after a complete, verified image is in place, most failures cost you a retry and nothing else:

Failure What the ESP32 does What you do
The server answers with an error (404, 500) Writes nothing, logs the HTTP code, carries on Fix the file on the server, send again
Wi-Fi drops mid-download The write ends short and is rejected; the pointer doesn't move Retry once the link is stable
Power is cut mid-download Boots the old slot; the half-written one was never pointed at Retry
The image is corrupt The checksum or digest check fails before the switch Rebuild, check you uploaded the right file
The image is bigger than the slot Refuses to start the update Shrink the build, or repartition (USB only)
The new firmware boots, then misbehaves Depends on rollback, below Often a site visit, if it can't reach the network

In the first five rows the device keeps running what it ran before. The last row strands devices in the field, and it is what people usually mean by rollback.

Rollback: three different things with one name

"Rollback" covers three separate mechanisms, and a fleet needs to know which ones it has.

1. Bootloader rollback on the device

ESP-IDF can boot a new image on probation. With CONFIG_BOOTLOADER_APP_ROLLBACK_ENABLE set, a fresh app starts in the ESP_OTA_IMG_PENDING_VERIFY state and must confirm itself by calling esp_ota_mark_app_valid_cancel_rollback(). If the chip resets first, the bootloader marks the image ESP_OTA_IMG_ABORTED and boots the previous slot. An app that knows it is broken can call esp_ota_mark_app_invalid_rollback_and_reboot(). The ESP-IDF OTA documentation has the full state machine.

The catch is when the confirmation happens. The Arduino core for ESP32 (2.0.x) ships with rollback enabled, but by default it confirms every new image during start-up, before setup() runs. On a stock Arduino or PlatformIO build, a firmware that boots and then can't join Wi-Fi is already marked valid. To get real protection, defer the confirmation until the firmware has proved itself:

#include "esp_ota_ops.h"

// Stop the Arduino core confirming the new image at boot (C linkage required).
extern "C" bool verifyRollbackLater() { return true; }

// Call once the device has proved it works, e.g. connected and read its sensors.
void confirmFirmwareIfPending() {
  esp_ota_img_states_t state;
  if (esp_ota_get_state_partition(esp_ota_get_running_partition(), &state) == ESP_OK &&
      state == ESP_OTA_IMG_PENDING_VERIFY) {
    esp_ota_mark_app_valid_cancel_rollback();
  }
}

Add a deadline: if the device hasn't confirmed within a couple of minutes, call esp_restart() and the bootloader puts the old firmware back. Pick the health check carefully. "Connected to the broker" catches a broken network stack, not a firmware that connects perfectly and reads every sensor wrong.

2. Withdrawing a release across the fleet

The second kind happens on the server: stop offering the bad version and offer the last good one. No device is touched. The point is to stop the damage spreading, so devices that haven't updated yet get the good build. It only works if you kept the previous builds.

3. Downgrading a device

The third kind sends an older build to a device that already took the bad one. Over the air that's just another update with a lower version number, so the device must still be online and able to download. One that can't reach the network needs a cable. ESP-IDF also has the opposite feature, anti-rollback: a secure_version burned into eFuse that makes downgrades impossible. That suits security patches and defeats an escape hatch, so choose before you burn anything.

Rolling an update out across a fleet

Most fleet OTA trouble is process, not technology:

How Synacl does it

The Synacl gateway firmware runs on a classic ESP32 and updates over the air from the web app, on every plan including the free tier. This is what the code does today.

Starting an update. Each card on the Gateways page shows the firmware version the gateway reports. When its channel offers something newer, the card shows an arrow with the new version and a download button; otherwise it says latest. The button needs permission to edit gateways (owners, admins and operators by default). The platform first checks that the build exists, that the gateway isn't already on that version or newer, and that it isn't mid-update, then sends it one message naming the version.

On the gateway. The firmware downloads the image into the idle slot, lets the ESP32 verify it, and only then switches slots and reboots. Wi-Fi, broker and account settings in NVS come through untouched. If anything fails before the switch, the gateway logs the reason (OTA fetch failed: with the HTTP code, or OTA error: with an update error code) under the System category and keeps running the old firmware without rebooting. Those lines stream in live logs while you watch, and show up afterwards in the mobile app's diagnostics snapshot.

In the app. While an update is in flight, the gateway's page shows Firmware update in progress and disables Restart, Resend config and Reset config, since a reboot mid-download cancels the update; the platform refuses those commands for up to five minutes. The confirmation dialog says to allow about a minute. When the gateway reconnects it reports its version again, and that is how you know the update took.

Channels. Each gateway is on Stable or, when beta access is enabled for your account, Beta. A beta gateway is offered the newest pre-release and moves onto the stable release once it ships. Switching back to Stable never downgrades a gateway remotely; it just stops being offered pre-releases.

Rollback, platform side. Synacl keeps at least the last three builds of each channel on the server and can pin a channel to an earlier one when a release is withdrawn. Gateways that haven't updated are then offered the pinned build, and the USB flasher writes it too. A pin never downgrades anyone by itself: a gateway that already took the withdrawn build keeps it and shows no update until a newer release ships. Synacl support can send it a specific older build if it needs to come off sooner.

Recovery by cable. A gateway that can't take an update, because it's offline or runs a build that won't connect, is re-flashed with Flash via USB on its page, in Chrome or Edge. That writes the channel's current build at the bootloader, partition table, OTA-data and app offsets and leaves NVS alone, so the gateway keeps its Wi-Fi and account settings and needs no re-provisioning. It takes the Provision gateways permission (owners and admins). The flashing article walks through it.

What it doesn't do yet

A checklist before you update the fleet

  1. Update the gateway that best tolerates a one-minute gap, and wait for it to report the new version.
  2. Watch its readings and System log for an hour, or one full cycle of whatever it monitors.
  3. Update the rest in small batches, outside the hours when a reboot interrupts a process.
  4. Note which gateways are offline before you start, and come back to them. The uptime monitoring guide covers telling a dead gateway from a quiet one.
  5. Keep a USB cable and a laptop with Chrome on the site-visit list until the last gateway reports in.

The firmware updated here is the same image the ESP32 web flasher writes in the first place. If you haven't put it on a board yet, start with How to flash an ESP32 from your browser; ESP32 Modbus RTU to a cloud dashboard in 15 minutes then gives the freshly updated gateway something to read.

Try it on your own hardware

Synacl is free for five devices — no card, no sales call. Flash a gateway from the browser, add your first device, and see live data in a few minutes.