On 29 July 2026, customers experienced errors and degraded functionality across several Brevo products for approximately an hour and a half. The cause was a configuration change applied during planned preventive work following the previous day's incident. The issue is fully resolved and all services are operating normally.
Duration: approximately 1 hour 27 minutes (06:10–07:37 UTC).
Affected: Automation workflows and exports; Transactional Email; Transactional SMS; Marketing Email; WhatsApp; Segmentation; and some account and user management pages, including contact lists. Outbound Webhooks returned intermittent errors.
During the incident:
Some automation workflows and exports did not run.
Email and SMS sending returned errors for some requests. Messages that were accepted but delayed were processed once service recovered.
Segmentation-based logic did not evaluate correctly for some customers.
No customer data was lost or altered. The affected component is a temporary cache and holds no primary customer data.
Following the incident of 28 July, we began preventive work to strengthen the infrastructure involved. While applying one of these changes, our deployment tooling applied a broader set of changes than we had specified, reverting configuration on part of our infrastructure to an earlier state.
This made a shared caching layer, used by many Brevo services to hold temporary data, briefly unreachable. Services that depend on that cache then returned errors or failed to restart, which is why the impact was visible across several products at once.
We moved the affected cache instances onto healthy infrastructure, provisioned replacement nodes with the correct configuration, removed the unhealthy node so the remaining cache instances could recover, and reconciled our infrastructure configuration so the reverted settings could not return.
Removed the automatic-replacement behaviour that allowed a configuration change to immediately disrupt running infrastructure.
Changed how we verify the scope of a change before applying it. We found that our tooling silently ignored the instruction limiting the change to a single component. Changes are now reviewed against an explicit plan that must match the intended scope.
Reconciled our infrastructure configuration so that it is fully captured and reviewable, with no unrecorded settings that can be lost.
Hardening dependent services so that a brief interruption to the shared cache degrades gracefully instead of causing services to restart.
We are sorry for this disruption, particularly coming so soon after the previous day's incident. Work intended to make the platform more resilient should not have put it at risk, and the changes above are aimed squarely at that.