Creuto is now an OpenAI Select Partner Read More
Envoy Gateway 1.9.1 fixes an OAuth2 padding oracle and reverts a 1.9.0 timeout change. The TLS certificate warning applies to fewer clusters than you think.

Envoy Gateway 1.9.1 shipped on 28 August 2026, and the most expensive line in its release notes is not the security fix. It is a qualifier: the warning that TLS listeners can go active without their certificates is an upgrade note for v1.9.0 users only. If you are on 1.8.x, the path the project tells you to take — straight to 1.9.1, skipping 1.9.0 — is not affected by it at all.
That distinction decides how much work this release is for you. Here is what the release carries, and which group each part lands on.
| Change | Who it affects | What it costs |
|---|---|---|
| AES-256-GCM enabled for OAuth2/OIDC session cookies, legacy AES-256-CBC decryption disabled | Anyone running the OIDC/OAuth2 filter via SecurityPolicy | Every active session is invalidated once |
| SDS/RDS initial fetch timeout reverted from 0s to Envoy's 15-second default | Everyone, but it only bites on the 1.9.0 to 1.9.1 hop | A certificate window on already-running proxies |
| Recommended rolling update to close that window fast | 1.9.0 users | Spare capacity for 2x proxy replicas, and cut long-lived connections |
The release reverts a change made in 1.9.0 that set the initial fetch timeout to zero for SDS (Secret Discovery Service) and RDS. The reason given in the notes is blunt: a cluster waiting on a missing secret or endpoint with no timeout never left warming, which paused CDS updates for the whole proxy and kept its health checks from starting. Both now use Envoy's default 15-second timeout again, as in 1.8.x.
The 1.9.0 notes record the same change from the other side, listed there as a fix for "the initial fetch timing out" on TLS secrets. A fix in August became a revert two weeks later. That is not a criticism of the project — it is a reminder that a point release's risk sits in the sequence of states a cluster passes through, not in the diff.
Alongside it, 1.9.1 drops support for HTTP as an OIDC issuer URL scheme, adds validation for the issuer URL configured in SecurityPolicy, and requires HTTPS for OCI Wasm image pulls — the implicit fallback to plain HTTP is gone, so a Wasm extension backed by a plain-HTTP registry now fails to load unless that registry is explicitly marked insecure.
If you do not configure the OAuth2 or OIDC filter, this part does not apply to you. The vulnerability lives in Envoy's OAuth2 HTTP filter, not in Envoy Gateway's control plane, and it only exists where that filter is in the request path.
The mechanics are worth knowing before you decide how fast to move. Per the Envoy security advisory, the filter's encrypt and decrypt functions used AES-256-CBC with no authentication tag, and the /callback endpoint returned 302 on a successful decrypt and 401 on a padding failure. That difference is the oracle. The advisory's proof of concept recovers the PKCE code verifier in roughly 6,200 requests and about 100 seconds.
Recovering the code verifier is not the attack, though. The attacker still has to have intercepted the victim's CodeVerifier cookie and obtained the victim's authorization code, and has to finish inside the cookie's 600-second lifetime. That is why NVD carries it at CVSS 3.1 6.8, medium, with attack complexity high and user interaction required. Treat it as a real account-takeover path for a gateway fronting an SSO login, and as noise if your gateway does no OIDC.
One trap sits in the fix rather than the flaw. The GCM switch is set in the default Envoy bootstrap, so an EnvoyProxy using spec.bootstrap with type Replace — which is the default when no type is given — does not receive it. If you run a replacement bootstrap, upgrading to 1.9.1 leaves you on the CBC path until you add envoy.reloadable_features.oauth2_use_gcm_encryption: true and envoy.reloadable_features.oauth2_legacy_cbc_decrypt_compat: false to a layered runtime static layer yourself. Check that before you tell anyone the issue is closed.
The release notes answer this one directly, which is unusual and useful. Existing OIDC sessions were encrypted with AES-256-CBC and are no longer accepted, so users with an active session are redirected to re-authenticate once after upgrading.
So yes: if your gateway terminates SSO for an internal estate, upgrading logs everybody out. Not gradually, not on expiry — at the moment the new bootstrap reaches the proxies. There is no compatibility window to configure, because the compatibility window is the vulnerability. Schedule it where a forced re-login is an annoyance rather than an incident, and tell support first.
Logins already in flight are a smaller, separate problem. 1.9.1 also renames the PKCE cookie and rescopes the OIDC flow cookies to the redirect path, and the notes say a login in progress across the rollout may need to be retried once.
Here is the sequence the project documents for 1.9.0 users. Reverting the timeout changes the SDS configuration in every generated listener and cluster. That triggers an open Envoy bug on proxies still running when the upgraded controller starts. About 15 seconds after the new controller pushes configuration, TLS listeners go active without certificates and every new TLS handshake fails.
Clusters using BackendTLSPolicy, Backend TLS or the global rate limit service lose their CA and client certificates the same way. Plain HTTP routing is unaffected. The pods still report Ready — which is the detail that turns this from a blip into an outage, because nothing in your orchestration notices.
The Envoy issue explains why the proxy cannot recover on its own: Envoy keys an SDS provider by the secret name plus the bytes of its ConfigSource, so removing initial_fetch_timeout makes it create a new provider for a secret it already holds, and the new watch folds into the existing subscription without sending a request. After the timeout, the listener is marked ready anyway. The secret is in memory the whole time.
Proxies started fresh on the new version are not affected, and the 1.8.x to 1.9.1 path is not affected. That is the whole argument for skipping 1.9.0.
The project's own mitigation is to shorten the time old proxy pods run beside the new controller: before upgrading, set envoyDeployment.strategy (or envoyDaemonSet.strategy) in the attached EnvoyProxy to {type: RollingUpdate, rollingUpdate: {maxSurge: 100%, maxUnavailable: 0}} so every replacement pod starts at once.
Two costs come attached, and the notes state both. It needs cluster capacity for twice the proxy replicas. And replacing a proxy pod closes every connection it holds: the pod drains gracefully over shutdown.drainTimeout, 60 seconds by default, but anything still open at the end — WebSockets and gRPC streams included — is cut. A gateway is exactly where long-lived connections live, so the mitigation for the certificate window is also the thing that breaks your streaming traffic.
The two mitigations only conflict if you assume pod replacement is the only cure. It is not. The notes say pushing a Secret again, for example by rotating it, restores whatever uses that Secret without a restart. That is the path for a team without spare capacity, or with gRPC streams it will not cut: upgrade, detect, re-push the affected Secrets, and replace pods later on your own schedule. It is slower and more manual, and it fixes one Secret at a time rather than everything at once.
Either way, detection comes first. A proxy in this state logs initial fetch timed out for ...tls.v3.Secret. The notes give two alerts: envoy_sds_init_fetch_timeout > 0, whose counter stays set for the life of the process, and increase(envoy_listener_server_ssl_socket_factory_downstream_context_secrets_not_ready[5m]) > 0, which rises for every handshake rejected because the certificate is not loaded. Wire both before the controller rolls, not after. This is ordinary infrastructure management and monitoring discipline, and it is what separates a 90-second incident from a 40-minute one.
A usable order, then, for a 1.9.0 cluster: set the rolling update strategy and confirm you have the headroom; verify your EnvoyProxy is not using a Replace bootstrap; wire the two alerts; drain or fail over long-lived gRPC and WebSocket traffic if you can; upgrade the controller; watch the 15-second mark; replace pods or re-push Secrets depending on which cost you chose. For a 1.8.x cluster, most of that is unnecessary — go straight to 1.9.1 and plan only for the forced re-authentication.
InfoQ's write-up matches the release notes on the substance, including the 28 August date, but it compresses the certificate warning into a general risk of upgrading rather than the v1.9.0-only note the project wrote. If you read the secondary coverage and not the notes, you would plan a disruptive upgrade you do not need.
The advisory itself carries an internal inconsistency worth knowing if you pin Envoy directly. GitHub's structured advisory data lists the patched Envoy versions as 1.35.13, 1.36.9, 1.37.5 and 1.38.3, while the advisory's own description — mirrored verbatim by NVD — says the flaw exists "prior to 1.35.11, 1.36.7, 1.37.3, and 1.38.1". Both are from the same advisory, published 23 June 2026. If your data plane is version-pinned outside Envoy Gateway, take the higher set.
We do a lot of API development and integration work that terminates at a gateway, and the pattern here is familiar from other releases: the diff is small, the state machine the cluster walks through is not. The same reflex applies to Kubernetes platform work generally — read upgrade notes for who they exempt, not only for what they warn about. We said the same thing about an Artifactory vulnerability chain earlier this year, and about the TLS parameters a domain actually negotiates when we looked at post-quantum key exchange.
If you are on 1.9.0 today, the decision in front of you is not whether to upgrade. It is which cost you would rather pay in the 15-second window: double capacity and severed streams, or a slower manual Secret re-push with the handshakes failing for longer.
Yes. The Envoy Gateway release notes recommend that v1.8.x users upgrade directly to v1.9.1 and skip v1.9.0, because the SDS configuration change that triggers the certificate failure only affects proxies already running v1.9.0 when the upgraded controller starts. The 1.8.x path avoids that window entirely.
No. CVE-2026-47775 is a padding oracle in Envoy's OAuth2 HTTP filter, so it only exists where that filter sits in the request path. An Envoy Gateway deployment with no OAuth2 or OIDC SecurityPolicy configured is not exposed to it, and can treat 1.9.1 as an ordinary maintenance release.
Yes, once. Envoy Gateway 1.9.1 enables AES-256-GCM and stops accepting the legacy AES-256-CBC path, and the release notes state that users with an active OIDC session are redirected to re-authenticate once after upgrading. Sessions issued under the old scheme are not migrated or re-encrypted.
On a v1.9.0 to v1.9.1 upgrade, reverting the SDS initial fetch timeout changes the SDS configuration, which triggers an open Envoy bug. About 15 seconds after the new controller pushes configuration, TLS listeners go active without certificates while the pods still report Ready.
Replacing a proxy pod does. The release notes warn that a replaced pod drains over shutdown.drainTimeout, 60 seconds by default, and anything still open at the end, including WebSockets and gRPC streams, is cut. Re-pushing the affected Secret avoids the restart.
Yes. The Envoy Gateway release notes say pushing a Secret again, for example by rotating it, restores whatever uses that Secret without a restart. It is slower because it works one Secret at a time, but it preserves long-lived connections and needs no spare cluster capacity.
Ready to take the first step towards unlocking opportunities, realizing goals, and embracing innovation? We're here and eager to connect.
11th Floor, O-Hub, Chandaka Industrial Estate, Infocity, Bhubaneswar, Odisha 751024
Level 4, 11 York Street Sydney Startup Hub Sydney, NSW – 2000
30 N. Đinh Nghệ, Phước Mỹ Sơn Trà, Đà Nẵng / Da Nang City – 550000
Level 25, AIDP Business Tower, Dubai Marina, United Arab Emirates
50 Beauchamp Street, Wellington, WGN 5028, New Zealand