Migrating from NGINX to Envoy Gateway Without Downtime
Migrating from NGINX to Envoy Gateway Without Downtime
Table of Contents
- Why Envoy Gateway Instead of Another Ingress Controller?
- How Did We Achieve Zero-Downtime Migration?
- What Gotchas Did We Encounter?
- What Are the Prerequisites?
- How Does the New Architecture Support Advanced Traffic Management?
- The Design Decision: Selector Repointing Over DNS Migration
- Next Steps
In November 2025, Kubernetes SIG Network announced that the community-maintained ingress-nginx project would be retired in March 2026. No further releases, no bugfixes, no security patches. This affects roughly half of all cloud-native environments.
If you're running ingress-nginx (the community project, not F5's commercial NGINX-INGRESS), you now have an infrastructure dependency that is actively accumulating unpatched vulnerabilities. The NGINX "snippets" annotations that once provided flexibility became serious security vulnerabilities. The Ingress API itself is frozen and will not be extended. The entire Kubernetes networking ecosystem is moving to the Gateway API.
This post covers how we migrated Kloudfuse from NGINX Ingress Controller to Envoy Gateway with zero downtime. The key innovation was a service selector repointing technique that preserves the existing load balancer throughout the migration, avoiding DNS changes, IP address shifts, and the traffic interruptions that typically accompany ingress controller swaps. We'll walk through the three-step migration, the technical decisions behind it, and the gotchas we encountered.
Why Envoy Gateway Instead of Another Ingress Controller?
The question isn't just "what replaces NGINX" but "what does the Kubernetes networking ecosystem look like in two years?" The answer is the Kubernetes Gateway API, and Envoy Gateway is the CNCF's own choice for implementing it. When CNCF migrated their internal services cluster, they chose Envoy Gateway. That's a strong signal.
The specific advantages over NGINX Ingress Controller:
- Dynamic configuration via xDS. NGINX relied on config files requiring process reloads. Envoy receives configuration updates at runtime through the xDS protocol. Routes can be swapped without affecting in-flight requests.
- Native rate limiting. Envoy supports both global and local rate limiting simultaneously on the same route. NGINX's rate limiting required non-standard annotations.
- Circuit breaking at the network level. Envoy enforces circuit breaking distribitively, with configurable thresholds. NGINX required per-application configuration.
- Role-based configuration model. The Gateway API separates concerns into distinct resources with separate roles. NGINX's Ingress API relied on a single user role.
- Protocol support. First-class HTTP/2 with transparent proxying, native gRPC support, and protocol-agnostic routing. NGINX Ingress only supported HTTP.
How Did We Achieve Zero-Downtime Migration?
Most migration guides recommend one of two approaches: weighted DNS shifting or parallel load balancers. Both have problems.
Weighted DNS shifting depends on unpredictable DNS TTL propagation. Parallel load balancers mean new IP addresses, which would disrupt many external systems.
The Service Selector Repointing Technique
Our approach preserves the existing NGINX Load Balancer Service throughout the migration. The key is to repoint the service's pod selector from NGINX pods to Envoy pods. The same load balancer keeps the same IP addresses and DNS records.
Here's the three-step migration:
| Step | Configuration Change | Traffic Served By | What Happens |
|---|---|---|---|
| Step 1: Prepare | Enable Envoy + envoyMigration.enabled |
NGINX (unchanged) | Envoy starts alongside NGINX. NGINX LB gets helm.sh/resource-policy: keep annotation. |
| Step 2: Switch | Set envoyMigration.external: true |
Envoy | NGINX LB selector switches to Envoy pods. |
| Step 3: Cleanup | Set ingress-nginx.enabled: false |
Envoy | Remove NGINX controller pods. NGINX LB Service is preserved. |
The helm.sh/resource-policy: keep annotation is critical. It prevents Helm from deleting the service, which is exactly what we're trying to avoid.
Step-by-Step Configuration
Step 1: Prepare. Add to custom_values.yaml:
envoy-gateway:
enabled: true
installGatewayRoutes: true
envoyMigration:
enabled: true
ingress-nginx:
enabled: true
installIngressRules: true
Run helm upgrade. Verify Envoy pods are running alongside NGINX. Traffic still flows through NGINX.
Step 2: Switch. Update custom_values.yaml:
envoy-gateway:
envoyMigration:
external: true # switch external LB to Envoy
internal: true # switch internal LB to Envoy
Run helm upgrade. Verify by checking that HTTPS returns a 401 (Envoy's auth challenge) instead of a 302 (NGINX's redirect behavior).
Step 3: Cleanup. Update custom_values.yaml:
ingress-nginx:
enabled: false
installIngressRules: false
Run helm upgrade. Verify NGINX controller pods are gone and HTTPRoutes show Accepted/Resolved status.
Rollback
If issues arise after Step 2, the rollback is straightforward: remove the resource-policy annotations from the NGINX LB, disable Envoy, re-enable NGINX, and run helm upgrade.
What Gotchas Did We Encounter?
- AWS NLB hairpin problem. The workaround: add a
hostAliasesentry mapping the domain to Envoy's ClusterIP. - CRD scope. In shared clusters, only delete the GatewayClass for your own namespace.
- Orphaned resources. When running
helm delete, these require manual cleanup. - Gateway "Programmed: False." This is expected behavior during migration.
- TLS configuration format change. NGINX's flat TLS configuration migrates to a nested format during the migration period.
What Are the Prerequisites?
The migration requires:
- Kloudfuse 4.0.0 or later
- Kubernetes 1.27+
- Helm 3.x
- cert-manager v1.14+ with Gateway API support enabled.
How Does the New Architecture Support Advanced Traffic Management?
With Envoy Gateway, Kloudfuse deployments gain access to:
- Separate internal and external traffic paths.
- Envoy's native traffic management.
- Protocol and observability improvements.
The Design Decision: Selector Repointing Over DNS Migration
We chose selector repointing to avoid requiring an IP change, which could disrupt customer environments.
Next Steps
The Envoy Gateway setup guide covers installations and migrations from NGINX. If you're still running ingress-nginx, the migration is less disruptive than most teams expect.