Deployment Strategies: Rolling, Blue-Green, Canary
Compare and implement deployment strategies—rolling updates, blue-green deployments, and canary releases—to reduce risk and enable safe production releases.
Rolling updates replace instances in batches, blue-green releases switch traffic between two environments, and canaries expand traffic behind health gates. This guide compares their availability, capacity cost, compatibility requirements, and recovery behavior, with Kubernetes and Argo Rollouts examples. It also covers metric thresholds, feature flags, database changes, and rollback practice so you can choose a strategy and define its safeguards before a production release.
Deployment Strategies: Rolling, Blue-Green, and Canary Releases
Introduction
A release that sends every request to v2 at once can turn a small regression into a full outage. A canary release might send 5% of traffic to v2, compare its error rate and latency with v1, and pause promotion if either crosses a pre-set limit. The first approach is simpler; the second buys time to catch problems, provided traffic volume and telemetry make the comparison meaningful.
This guide compares rolling, blue-green, and canary releases by availability, temporary capacity, traffic control, compatibility between versions, and recovery time. Use those criteria to choose a rollout, then define the health gate and data recovery plan before production traffic moves.
Rolling Update Mechanics
Rolling updates gradually replace old pods with new ones. Kubernetes handles this natively for Deployments — you configure a few parameters and it takes care of the rest.
Basic rolling update configuration:
apiVersion: apps/v1
kind: Deployment
metadata:
name: myapp
spec:
replicas: 6
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 1 # Allow 1 extra pod during update
maxUnavailable: 0 # Never have fewer than desired replicas
selector:
matchLabels:
app: myapp
template:
metadata:
labels:
app: myapp
version: v2
spec:
containers:
- name: myapp
image: myregistry.azurecr.io/myapp:v2.0.0
ports:
- containerPort: 8080
Monitor rolling update progress:
# Watch rollout status
kubectl rollout status deployment/myapp
# View deployment details
kubectl describe deployment myapp
# Check revision history
kubectl rollout history deployment/myapp
Rolling update behavior:
| Parameter | 6 Replicas | Effect |
|---|---|---|
| maxSurge: 1, maxUnavailable: 0 | 7 pods during transition | Maximum availability, slower |
| maxSurge: 2, maxUnavailable: 0 | 8 pods during transition | Faster, more resources |
| maxSurge: 0, maxUnavailable: 1 | 5 pods during transition | Minimum resources, some downtime |
Rollback a rolling update:
# Immediate rollback to previous version
kubectl rollout undo deployment/myapp
# Rollback to specific revision
kubectl rollout undo deployment/myapp --to-revision=3
# Watch rollback
kubectl rollout status deployment/myapp
Blue-Green Deployment Setup
Blue-green deployments run two environments and switch application traffic between them. Keeping the old environment available can make application rollback fast, but a traffic switch cannot reverse database writes or schema changes. Plan those changes separately, usually with backward-compatible expand-contract steps.
Infrastructure setup:
Internet → Load Balancer → Blue (v1) OR Green (v2)
↓ ↓
[Production] [Production]
Kubernetes implementation with two Deployments:
# Blue deployment (current version)
apiVersion: apps/v1
kind: Deployment
metadata:
name: myapp-blue
labels:
app: myapp
slot: blue
spec:
replicas: 6
selector:
matchLabels:
app: myapp
slot: blue
template:
metadata:
labels:
app: myapp
slot: blue
version: v1
spec:
containers:
- name: myapp
image: myregistry.azurecr.io/myapp:v1.0.0
---
# Green deployment (new version)
apiVersion: apps/v1
kind: Deployment
metadata:
name: myapp-green
labels:
app: myapp
slot: green
spec:
replicas: 6
selector:
matchLabels:
app: myapp
slot: green
template:
metadata:
labels:
app: myapp
slot: green
version: v2
spec:
containers:
- name: myapp
image: myregistry.azurecr.io/myapp:v2.0.0
Service switching between slots:
# Initial state: traffic to blue
apiVersion: v1
kind: Service
metadata:
name: myapp
spec:
selector:
app: myapp
slot: blue
ports:
- port: 80
targetPort: 8080
# Switch to green (update selector)
# kubectl patch service myapp -p '{"spec":{"selector":{"slot":"green"}}}'
Blue-green with Argo Rollouts:
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: myapp
spec:
strategy:
blueGreen:
activeService: myapp-blue
previewService: myapp-preview
autoPromotionEnabled: false # Manual promotion
scaleDownDelaySeconds: 600 # Keep old version for 10 min after switch
selector:
matchLabels:
app: myapp
template:
metadata:
labels:
app: myapp
spec:
containers:
- name: myapp
image: myregistry.azurecr.io/myapp:v2.0.0
Canary Deployment with Argo Rollouts
Canary deployments gradually shift traffic to the new version, monitoring metrics to detect issues.
Argo Rollouts canary configuration:
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: myapp
spec:
replicas: 10
strategy:
canary:
steps:
- setWeight: 5 # Start with 5% traffic to new version
- pause: {} # Wait for manual inspection
- setWeight: 20
- pause: { duration: 10m } # Auto-proceed after 10 minutes
- setWeight: 50
- pause: {}
canaryMetadata:
labels:
role: canary
stableMetadata:
labels:
role: stable
trafficRouting:
nginx:
stableIngress: myapp-stable
additionalIngressAnnotations:
canary-by-header: X-Canary
analysis:
templates:
- templateName: success-rate
startingStep: 1
args:
- name: service-name
value: myapp-canary
selector:
matchLabels:
app: myapp
template:
spec:
containers:
- name: myapp
image: myregistry.azurecr.io/myapp:v2.0.0
Analysis template for automated checks:
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
name: success-rate
spec:
args:
- name: service-name
metrics:
- name: success-rate
interval: 2m
successCondition: result[0] >= 0.95
failureLimit: 3
provider:
prometheus:
address: http://prometheus:9090
query: |
sum(rate(http_requests_total{service="{{args.service-name}}",status!~"5.."}[5m]))
/
sum(rate(http_requests_total{service="{{args.service-name}}"}[5m]))
- name: error-rate
interval: 1m
successCondition: result[0] < 0.01
failureLimit: 5
provider:
prometheus:
address: http://prometheus:9090
query: |
sum(rate(http_requests_total{service="{{args.service-name}}",status=~"5.."}[5m]))
Traffic-shifting guardrails
setWeight only changes user traffic when Argo Rollouts is connected to a traffic router such as an ingress controller or service mesh. Without that integration, weights may describe pod proportions rather than a precise request share. Verify the router’s behavior in staging, including how it handles retries, long-lived connections, and requests that bypass the normal ingress.
Treat each weight as a gate, not a timer. At each step, wait for a minimum request count and observation window, then compare candidate and stable versions over the same period. Pause promotion when telemetry is missing, the sample is too small, or the candidate crosses a pre-set service or business metric limit. Keep users on a consistent version during a session when changing versions mid-session could break state; use a stable cohort key when the router supports it.
Feature Flags Integration
Feature flags decouple deployment from release, enabling precise control over who sees new features.
LaunchDarkly in Kubernetes:
# The Secret is provisioned through the deployment secret manager.
apiVersion: apps/v1
kind: Deployment
metadata:
name: myapp
spec:
replicas: 1
selector:
matchLabels:
app: myapp
template:
metadata:
labels:
app: myapp
spec:
containers:
- name: myapp
image: myregistry.azurecr.io/myapp:v2.0.0
env:
- name: LD_CLIENT_KEY
valueFrom:
secretKeyRef:
name: launchdarkly-sdk
key: sdk-key
Progressive percentage rollout with flags:
// Example: gradual rollout of new checkout
const launchDarkly = require("@launchdarkly/node-server-sdk");
const client = launchDarkly.init(process.env.LD_CLIENT_KEY);
async function shouldShowNewCheckout(userId) {
return client.variation("new-checkout-flow", { key: userId }, false);
}
// Route based on flag
app.get("/checkout", async (req, res) => {
const userId = req.user.id;
const useNewCheckout = await shouldShowNewCheckout(userId);
if (useNewCheckout) {
res.redirect("/checkout/new");
} else {
res.redirect("/checkout/legacy");
}
});
Rollback Triggers and Automation
Automated rollback prevents bad releases from affecting users.
Prometheus metrics-triggered rollback:
apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
name: myapp
spec:
strategy:
canary:
analysis:
templates:
- templateName: error-rate-check
# Failed analysis aborts the rollout and returns traffic to the stable version
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
name: error-rate-check
spec:
metrics:
- name: error-rate
interval: 1m
successCondition: result[0] < 0.05
failureCondition: result[0] > 0.05
failureLimit: 2 # Trigger rollback after 2 consecutive failures
provider:
prometheus:
address: http://prometheus:9090
query: |
sum(rate(http_requests_total{service="myapp-canary",status=~"5.."}[5m]))
/
sum(rate(http_requests_total{service="myapp-canary"}[5m]))
GitHub Actions automated rollback:
rollback:
runs-on: ubuntu-latest
if: failure()
steps:
- name: Rollback deployment
run: |
# Rollback in Kubernetes
kubectl rollout undo deployment/myapp -n production
# Or rollback Helm
helm rollback myapp -n production
- name: Notify
uses: slackapi/slack-github-action@v1
with:
channel-id: "deployments"
payload: |
{
"text": "Production deployment failed, rolled back automatically",
"blocks": [
{
"type": "section",
"text": {
"type": "mrkdwn",
"text": ":x: *Production deployment failed*\nRolling back to previous version."
}
}
]
}
Choosing the Right Strategy
| Strategy | Risk | Speed | Cost | Best For |
|---|---|---|---|---|
| Rolling | Low | Medium | Low | Stateless services, Kubernetes native |
| Blue-Green | Very Low | Fast | High (2x resources) | Fast application rollback, isolated validation |
| Canary | Low-Medium | Slow | Medium | New features, A/B testing, gradual rollout |
Decision factors:
- Application state: Stateful apps may have issues with rolling updates
- Traffic sensitivity: User-facing apps benefit from blue-green or canary
- Resource budget: Blue-green requires double the capacity
- Rollback speed: How fast must you recover from a bad deploy
- Testing confidence: Low confidence = canary with analysis
These labels are starting points, not guarantees. A canary has little value when traffic is too low to produce a useful sample, and a blue-green switch can still fail if the router, shared dependencies, or database are not ready. Compare the cost of spare capacity with the time and user impact of recovery. Also check whether the deployment can tolerate old and new versions running together; that compatibility requirement often rules out a strategy before its speed or cost does.
When to Use / When Not to Use
When rolling updates make sense
Rolling updates work best in Kubernetes for stateless services where you can have multiple versions running simultaneously. If your application handles traffic gracefully when some instances run the old version and others run the new version simultaneously, rolling updates are the simplest choice.
Use rolling updates when you need zero-downtime deployments and cannot afford double the infrastructure for blue-green. They are the Kubernetes default for a reason.
The sweet spot for rolling updates is stateless APIs and services where request-level idempotency means mixing versions does not cause problems. If your service reads from one database and writes to one database, rolling updates are safe. Both versions handle the same requests the same way, so traffic splitting between old and new pods does not create data inconsistencies.
What makes rolling updates tricky is state that lives outside the database. In-memory session state, local file writes, cached data on disk — these become problems when a request that started on a v1 pod finishes on a v2 pod. For services with that kind of state, you need either sticky sessions or a different strategy entirely.
Rolling updates also work well when you have many services and cannot afford the operational overhead of blue-green for all of them. Blue-green for one critical service is manageable. Blue-green for twenty microservices is a significant burden. Roll with the default for most services, reserve blue-green for the ones that genuinely need it.
When blue-green makes sense
Blue-green fits services that need a fast traffic switch or a full-environment preview before release. Run smoke tests against green, then move traffic only after it passes. It can also avoid a period where old and new application versions serve requests together, which matters for workflows such as payments.
The switch only reverses application routing. With a shared database, old and new versions still need compatible schemas, and a rollback cannot undo transactions already committed by v2. Plan backward-compatible schema changes, backups, and a tested transaction recovery or repair procedure separately.
The capacity cost is real: keeping two complete environments can nearly double compute during the switch. For GPU-heavy inference or memory-intensive workloads, that cost may outweigh faster application rollback. Compare temporary capacity with the service’s recovery requirements.
When canary makes sense
Canary deployments are best for risky changes where you want real production traffic validation before committing fully. A new algorithm, a major UI redesign, a significant infrastructure change — these are all good canary candidates.
Use canary when you have the metrics infrastructure to validate the change automatically. Without metrics, canary is just slow blue-green.
The defining characteristic of a good canary candidate is uncertainty. You are not sure whether the change works at scale, whether the new algorithm performs better or worse, whether the UI redesign confuses or delights users. Canary lets you find out with real production traffic before rolling out to everyone.
ML model deployments are a canonical canary use case. A model that performed well in staging can degrade in production due to data distribution differences. Sending 5% of traffic to the new model while monitoring error rates and business metrics catches those cases before they affect all users. If the new model performs better, you expand. If it performs worse, you roll back and investigate.
UI redesigns are another strong canary case, but they require care. Canary for a UI change means routing real users to the new interface, which means you need to know what metric you are optimizing for. If the metric is conversion rate, you watch conversion. If it is engagement, you watch session length. Pick the metric before you deploy, not after.
What makes canary a bad fit: low-traffic services where 5% of traffic is not enough to generate meaningful signal, services where you cannot route a small percentage of users without breaking things (some payment flows need atomic consistency), and changes that are clearly safe and just need to ship. Do not use a rocket to kill a mosquito.
Production Failure Scenarios
Common Deployment Failures
| Failure | Impact | First response and recovery |
|---|---|---|
| New pods crash or never become ready | Capacity falls as old pods leave | Pause the rollout; inspect events, logs, probes, and config. Roll back if available capacity or error limits are breached. |
| Blue-green switch sends traffic to the wrong slot | Requests fail or reach mixed versions | Stop further changes, verify Service selectors and endpoint membership, then switch back only after confirming the old slot is healthy. |
| Canary analysis uses stale or unrelated data | Healthy release is blocked, or a bad one passes | Pause promotion and treat missing telemetry as unknown. Check labels, query scope, sample size, and baseline before resuming. |
| PDB blocks necessary eviction | Cluster maintenance stalls | Check allowed disruptions and replica health; adjust the PDB only after confirming the service can tolerate the eviction. |
| New code fails after writing data | Code rollback leaves incompatible records | Stop writes through the new path if possible. Use a forward fix or a tested data repair; restore a backup only with a plan for writes since that backup. |
For example, a canary may look healthy on HTTP errors while a new write path stores values the previous version cannot read. Routing traffic back then restores the old code but leaves incompatible data in place. Include data-shape and business checks in the release gate when a change affects stored records, and decide before rollout whether recovery means disabling the feature, repairing records, or restoring data.
Deployment Rollback Flow
flowchart TD
A[Deploy New Version] --> B{Health Check Pass?}
B -->|No| C[Rollback to Previous]
B -->|Yes| D[Monitor for 10 min]
D --> E{Metrics OK?}
E -->|Yes| F[Deployment Complete]
E -->|No| G[Auto Rollback]
C --> H[Alert Team]
G --> H
H --> I[Investigate Root Cause]
Observability Checklist
Set these checks before a release so each deployment strategy has a clear pause or rollback signal:
- Add a release marker with the service, image digest or version, strategy, region, start time, and each traffic or rollout step. Dashboards and traces should identify which version handled each request.
- Break out readiness and health, error rate, latency percentiles, and saturation by version. Include the resource limits that matter to the service, such as CPU, memory, queue depth, or connection-pool use.
- Compare the candidate with the stable version or a pre-release baseline over the same time window and traffic mix. Set a minimum observation window and sample size before looking at the result.
- Write numeric abort limits before rollout. Cover both service-level limits and regressions against baseline, and state how many failed windows trigger a pause or rollback.
- Treat missing, stale, or delayed telemetry as an unknown result. Pause promotion until the signals return; an empty chart does not show that the release is healthy.
- For rolling updates, check readiness, restarts, capacity, and errors while old and new versions overlap. Pause the next batch when a limit is crossed.
- For blue-green, run the same checks against the preview environment before switching traffic. After the switch, confirm the new version receives expected traffic while the old environment remains available for the rollback window.
- For canaries, evaluate the candidate at each traffic step before increasing its share. Use enough requests and time at each step for the chosen signals to be meaningful.
- After a rollback, confirm traffic returned to the prior version and its health and error signals recovered. Record the release marker, rollback time, and reason for the incident review.
# Check rollout status
kubectl rollout status deployment/myapp --timeout=5m
# Check pod age during rollout
kubectl get pods -l app=myapp -o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.status.phase}{"\t"}{.metadata.creationTimestamp}{"\n"}{end}'
# View rollout history
kubectl rollout history deployment/myapp
Common Pitfalls / Anti-Patterns
Not testing the rollback procedure
A rollback strategy you have never tested is not a rollback strategy. Practice rolling back in staging so you know what happens when you call kubectl rollout undo in production at 2am.
What to practice:
- Roll back to the previous revision and verify the application starts correctly
- Time how long rollback takes so you have realistic expectations during an incident
- Confirm that secrets, ConfigMaps, and volume mounts survive a rollback intact
- Test rolling back a canary deployment that has promoted partially through its steps
- Verify that database connections and in-flight requests are handled gracefully during rollback
Staging validation checklist:
# Deploy a known-good version first
kubectl apply -f deployment-v1.yaml
# Simulate a bad deploy
kubectl apply -f deployment-broken.yaml
# Watch it fail
kubectl rollout status deployment/myapp --timeout=2m
# Roll back immediately
kubectl rollout undo deployment/myapp
# Verify recovery
kubectl get pods -l app=myapp
kubectl logs -l app=myapp --tail=50 | grep ERROR
Without this rehearsal, you discover the gaps in your rollback procedure at 2am when users are affected.
Setting PDB too aggressively
PodDisruptionBudgets limit voluntary disruptions such as node drains. A PDB that requires all 3 pods in a 3-replica deployment to stay available blocks any voluntary eviction; it does not control a Deployment rolling update or prevent involuntary failures.
Common PDB mistakes and their consequences:
| PDB Setting | Replicas | Problem |
|---|---|---|
| minAvailable: 3 | 3 | Node drain blocked forever — no pod can be evicted |
| minAvailable: 2 | 2 | No pod can be voluntarily evicted |
| maxUnavailable: 0 | any | Blocks voluntary pod eviction during node upgrades |
Correct PDB sizing:
# For a 6-replica stateless service
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: myapp-pdb
spec:
# Allow 1 pod to be unavailable during disruption
# This still maintains 5/6 pods = 83% availability
maxUnavailable: 1
selector:
matchLabels:
app: myapp
For stateful services with quorum requirements:
# Kafka with 3 brokers needs at least 2 available for writes
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: kafka-pdb
spec:
minAvailable: 2 # Maintains quorum for writes
selector:
matchLabels:
app: kafka
Think of PDBs as a floor on disruption, not a ceiling on availability. The problem comes when you set that floor too high — the cluster gets stuck and nothing can be evicted when you actually need to drain a node.
Using the same strategy for all services
A simple stateless API and a complex stateful service with database connections need different deployment strategies. Cookie-cutter approaches lead to either over-engineering simple services or under-protecting complex ones.
Strategy selection by service type:
| Service Type | Recommended Strategy | Why |
|---|---|---|
| Stateless REST API | Rolling update | Simple, Kubernetes-native, no traffic complexity |
| Stateful service (DB, cache) | Blue-green or recreate | Data integrity; simultaneous versions risk corruption |
| ML inference service | Canary (manual analysis) | GPU resources expensive; validate accuracy before full rollout |
| Payment service | Blue-green | Instant rollback critical; cannot afford prolonged partial availability |
| Internal tool / admin panel | Recreate | Zero users affected during downtime; fastest deployment |
| User-facing with session state | Canary | Gradual exposure limits blast radius; session affinity matters |
Over-engineering a simple service:
Applying Argo Rollouts with automated metric analysis to a low-traffic internal dashboard wastes operational overhead. A basic rolling update with health checks covers this use case adequately.
Under-protecting a critical service:
Using a rolling update for a payment processing service means partial availability during transition — some transactions go to the old version, others to new. If a schema change is involved, this creates data inconsistency windows. Blue-green eliminates that window entirely.
Decision framework:
Does the service handle user transactions or financial data?
→ Yes: Blue-green
→ No: Continue
Is the service stateful (persists data locally)?
→ Yes: Blue-green or recreate
→ No: Continue
Is the service business-critical with no tolerance for degraded states?
→ Yes: Blue-green with manual promotion gate
→ No: Rolling update
Is this a risky change (new algorithm, major refactor)?
→ Yes: Canary with automated analysis
→ No: Rolling update
Pick the strategy that fits each service’s risk profile. Don’t apply the same approach everywhere.
Ignoring database schema changes
Deployment strategies handle application versions, not schema migrations. If your new version requires a new column that the old version cannot handle, deploying the new version before the migration is a disaster. Treat database migrations as a separate release concern.
The expand-contract pattern for safe schema changes:
This approach separates schema changes from application changes across multiple deployments:
| Phase | Action | Old App | New App |
|---|---|---|---|
| 1 — Expand | Add new column (nullable, with default) | Works | Works |
| 2 — Migrate | Backfill existing rows | Works | Works |
| 3 — Contract | Remove old column/index | Won’t work | Works |
Phase 1 adds the new schema element. Both old and new application versions continue working because the old column still exists. Phase 2 populates the new column with data. Phase 3 removes the old element only after all instances run the new version.
For a large table, make the backfill resumable and bounded: process small batches, track progress, and watch database load and replication lag. If writes continue during the backfill, define how changes made after a row is copied reach the new representation; dual writes, change capture, or a final reconciliation pass are common choices. Deploy code that can read both representations before switching writes, and verify counts or checksums before removing the old one. A migration is not reversible just because the application image is.
What goes wrong without this separation:
- Old version reads missing column: Application queries fail with “column does not exist” errors
- New version writes old column: Data written to deprecated column is lost when that column is dropped
- Index removal during traffic split: Queries that relied on the index become slow or time out on instances still using the old version
Backward-compatible changes that are safe for rolling updates:
- Adding a new table
- Adding a new nullable column with a default value
- Adding a new index (watch for write performance impact)
- Adding a new not-null column with a default
Backward-incompatible changes that require separate handling:
- Removing or renaming a column
- Changing a column type
- Changing a not-null constraint
- Removing an index
For incompatible changes, use the expand-contract pattern or blue-green with a migration freeze window. Deployment strategies handle application versions — they are not a substitute for a data migration plan.
Before each phase, write down the rollback boundary. Adding a nullable field is usually easy to leave in place; dropping or rewriting data may not be reversible. Keep old and new application versions compatible through the overlap window, and delay destructive cleanup until the new code is stable and the recovery window has closed. If a backfill or schema operation cannot meet the required lock and replication limits, split it into smaller work or schedule a maintenance window rather than coupling it to the traffic switch.
Trade-off Summary
| Strategy | Deployment Time | Resource Cost | Rollback Speed | Risk Level |
|---|---|---|---|---|
| Rolling update | Moderate (proportional to batch size) | Low (no extra capacity) | Minutes (reverse batch) | Low-Medium |
| Blue-green | Fast (instant switchover) | 2x (double infrastructure) | Instant (switch traffic back) | Low |
| Canary | Gradual (traffic shifting) | Low-Medium (few extra pods) | Fast (drop traffic to new) | Low |
| Recreate | Fast (no orchestration) | Zero extra | Minutes (redeploy old version) | High |
Security and Compliance Notes
Deployment strategies control how code reaches production. Security concerns here focus on ensuring the deployment process does not become an attack vector, and that the deployed application meets security requirements.
Image Verification Before Deployment
Before deploying, verify that the image comes from a trusted registry and has been scanned for vulnerabilities. Admission controllers like OPA Gatekeeper or Kyverno can enforce this policy at deploy time:
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sRequiredImageTag
metadata:
name: require-tag
spec:
match:
kinds:
- apiGroups: [""]
kinds: ["Pod"]
parameters:
namespaces: ["production"]
exemptImages:
- "myregistry.azurecr.io/base-image:*"
Block images that have not been scanned, images from untrusted registries, or images with known critical vulnerabilities.
Secrets in Deployment Configurations
Deployment manifests should not contain secrets. Use external secrets management (Vault, AWS Secrets Manager, GCP Secret Manager) and the External Secrets Operator to inject secrets at runtime. ConfigMaps are appropriate for non-sensitive configuration; secrets require the additional access controls.
When using Helm, secrets stored in values.yaml are visible in Helm release history. Store secrets in separate files excluded from version control or use external secrets operators.
TLS Configuration for Deployments
Ensure that ingress TLS certificates are valid and use modern TLS versions. The TLS termination point should reject connections using outdated protocols (TLS 1.0, TLS 1.1) and weak cipher suites.
For services that communicate internally, consider mutual TLS (mTLS) via a service mesh. Service mesh certificates rotate automatically and provide authentication between services without application-level certificate management.
Deployment Access Control
Limit who can initiate deployments to production. Use environment protection rules in your CI/CD platform to require manual approval for production deploys. Automated deploys to production bypass human review, which is an important control point.
# GitHub Actions environment protection
environment:
name: production
protection_rules:
- required_reviewers:
- engineering-lead
- wait_timer: 0
Audit logging should capture every production deployment: who initiated it, what image or commit was deployed, and when. This information is essential for incident response and compliance.
Rollback Security
Rollback procedures should be as controlled as forward deployments. The same access controls, approval processes, and audit logging should apply to rollbacks as to forward deployments.
Feature flags provide a rollback mechanism that does not require redeployment. Disabling a flag is faster than rolling back an image and can be done by authorized users without CI/CD access.
Compliance Considerations
Immutable deployments mean that once an image is deployed, it does not change. If a compliance audit requires knowing exactly what code ran at a specific time, immutable image references make that traceability possible.
Deployment approval workflows provide an audit trail of who approved production changes. For regulated industries, this human-in-the-loop control is often a compliance requirement.
Deployment documentation should capture what changed, why, and who approved it. This information supports post-incident review and demonstrates due diligence for compliance.
Quick Recap Checklist
- Confirm there is enough rollout capacity and database changes remain backward-compatible.
- Verify the traffic router, candidate labels, minimum sample size, and observation window before shifting traffic.
- Set health and business signals, plus explicit pause and rollback thresholds, before promotion.
- Practice rollback and check that the previous version still works with current data.
- Keep destructive schema cleanup and old environment removal outside the rollback window.
Interview Questions
unclean.leader.election.enable=false to prevent data loss during broker failures. For Kafka specifically, use Strimzi or Kafka Operator on Kubernetes for managed StatefulSets. Always test the failure scenario in a staging environment first. Incremental rollout with careful monitoring of replication lag is essential.
minAvailable or maxUnavailable to preserve the service’s availability or quorum during maintenance. They do not control a Deployment rolling update; use the Deployment strategy’s `maxUnavailable` and `maxSurge` settings for that.
RollingUpdate with maxSurge: 10-25% and maxUnavailable: 0 so updates happen in controlled batches, staggering deployments across node pools if you have multiple pools, using a wave-based deployment approach where you tag nodes and deploy to wave 1, wait for stability, then proceed. For image pulls specifically, use a local registry mirror or cache (Harbor, Amazon ECR), pre-pull images onto nodes, and set imagePullPolicy: IfNotPresent.
kubectl rollout undo deployment/myapp
kubectl argo rollouts abort myapp
kubectl describe pod shows events and exit codes, kubectl logs shows application output. Common causes: application crashes on startup due to missing environment variables or config maps, health check failing due to incorrect probe configuration, dependency connection failures, or permission issues with service accounts. Fix by checking the deployment spec matches the application's expectations, verify ConfigMaps and Secrets are mounted correctly, adjust readiness/liveness probes if they fail too aggressively, and ensure the container image tag points to the correct version.
terminationGracePeriodSeconds long enough to complete in-flight requests. Use preStop hooks to wait for load balancer to deregister the pod before stopping the container. For databases, use connection pooling with health-check aware connections. Blue-green is often better for stateful connections since you switch entire environments atomically.
# Preserve all six available; allow one extra pod
maxSurge: 1
maxUnavailable: 0
# Use no extra capacity; allow one pod to be unavailable
maxSurge: 0
maxUnavailable: 1
Choose based on spare capacity and the availability the service needs during rollout.
Further Reading
Official Documentation
- Kubernetes Deployment Documentation - Official guide to Deployments and rolling updates
- Kubernetes Pod disruptions - PDB behavior and voluntary disruptions
- Argo Rollouts Documentation - Progressive delivery with Argo Rollouts
- LaunchDarkly Feature Flags Documentation - Feature flag management best practices
Related Guides
- CI/CD Pipeline Design - Pipeline architecture patterns and optimization
- GitOps Implementation - GitOps workflows with ArgoCD and Flux
- Container Registry Setup - Image storage and scanning strategies
- Automated Testing in CI/CD - Testing strategies and quality gates
- Architecture Testing and Fitness Functions - Automate checks that protect release constraints
- Architecture Decision Records - Record the reasoning behind rollout choices
- Distributed Tracing - Compare candidate and stable request behavior
Tools and References
- Argo Rollouts GitHub - Open source progressive delivery controller
- Flagger - Progressive delivery Kubernetes operator
- Weave Flux - GitOps operator for Kubernetes
- Prometheus Operator - Monitoring for Kubernetes deployments
Conclusion
Choose a rollout strategy that matches the service’s state, traffic, and recovery needs. Define the health signals and database recovery plan before deployment, then rehearse the rollback while the stakes are low.
Deployment Checklist
# Before deployment
kubectl rollout history deployment/myapp
kubectl get pdb myapp -o yaml
# During deployment
kubectl rollout status deployment/myapp --timeout=10m
kubectl get pods -l app=myapp --watch
# After deployment
kubectl rollout status deployment/myapp
kubectl logs -l app=myapp --tail=100 | grep ERROR
kubectl get events --sort-by='.lastTimestamp' | grep myapp
Category
Related Posts
Health Checks: Liveness, Readiness, and Service Availability
Master health check implementation for microservices including liveness probes, readiness probes, and graceful degradation patterns.
Zero-Downtime Database Migration Strategies
Learn zero-downtime migration patterns, including expand-contract deployments, backward-compatible changes, rollback planning, and tools for safe releases.
Container Security: Image Scanning and Vulnerability Management
Implement comprehensive container security: from scanning images for vulnerabilities to runtime security monitoring and secrets protection.