
Table of Contents
By Khimananda Oli | Last reviewed: September 2026
Running production on EC2 Spot: Handling Interruptions Safely is the single most effective way to reduce AWS compute spend, but it requires treating capacity as ephemeral rather than guaranteed. Spot Instances offer up to 90% savings over On-Demand pricing, yet they can be reclaimed by AWS with only a two-minute warning when capacity is needed elsewhere. To use them reliably in production, you must architect your application to absorb these interruptions gracefully through automated lifecycle hooks, diversified allocation strategies, and stateless design patterns.
How do you detect EC2 Spot interruptions before termination?
Detection is the foundation of safe Spot usage. AWS provides two distinct signals, and confusing them is a common mistake I see in audits. The Rebalance Recommendation arrives first—typically 5 to 10 minutes before actual reclamation—and indicates that your instance is at elevated risk. This is your window to proactively replace capacity. The Interruption Notice follows later and gives you exactly 2 minutes before forced termination. Relying solely on the interruption notice leaves almost no time for complex cleanup.
Polling instance metadata for signals
Your application or a sidecar process must poll the Instance Metadata Service (IMDSv2) every 5 seconds. Do not use IMDSv1; it lacks the token-based security required for SOC 2 compliance and is disabled on many hardened AMIs. Here is a production-ready polling script:
TOKEN=$(curl -X PUT "http://169.254.169.254/latest/api/token" \
-H "X-aws-ec2-metadata-token-ttl-seconds: 21600")
while true; do
# Check for rebalance recommendation (early signal)
REBALANCE=$(curl -s -o /dev/null -w "%{http_code}" \
-H "X-aws-ec2-metadata-token: $TOKEN" \
http://169.254.169.254/latest/meta-data/events/recommendations/rebalance)
if [ "$REBALANCE" = "200" ]; then
echo "$(date -Iseconds) REBALANCE signal received — initiating proactive drain"
systemctl stop myapp.service
aws autoscaling complete-lifecycle-action \
--auto-scaling-group-name prod-spot-asg \
--lifecycle-hook-name spot-drain-hook \
--instance-id $(curl -s -H "X-aws-ec2-metadata-token: $TOKEN" \
http://169.254.169.254/latest/meta-data/instance-id) \
--lifecycle-action-result CONTINUE
break
fi
# Check for hard interruption notice (2-min warning)
INTERRUPTION=$(curl -s -H "X-aws-ec2-metadata-token: $TOKEN" \
http://169.254.169.254/latest/meta-data/spot/instance-action)
if [ -n "$INTERRUPTION" ] && [ "$INTERRUPTION" != "none" ]; then
echo "$(date -Iseconds) INTERRUPTION notice: $INTERRUPTION"
kill -SIGTERM $(pgrep myapp)
break
fi
sleep 5
done For Kubernetes workloads on EKS, you do not need custom scripts. Install the AWS Node Termination Handler via Helm; it watches both signals and automatically cordons nodes, drains pods respecting PodDisruptionBudgets, and completes ASG lifecycle actions. This is the standard pattern for containerized Spot workloads in 2026.
What is the correct Auto Scaling Group configuration for Spot?
The single biggest reliability lever is the Mixed Instance Policy. Never configure an ASG with a single Spot instance type. When that specific type experiences a capacity shortage, every instance in your fleet gets interrupted simultaneously—a correlated failure that overwhelms your remaining On-Demand base. Instead, define multiple instance types and let AWS allocate from whichever pool has surplus capacity.
Terraform mixed instance policy example
resource "aws_autoscaling_group" "prod_spot" {
name = "prod-web-spot"
desired_capacity = 6
min_size = 3
max_size = 12
vpc_zone_identifier = var.private_subnet_ids
mixed_instances_policy {
instances_distribution {
on_demand_base_capacity = 2
on_demand_percentage_above_base_capacity = 0
spot_allocation_strategy = "price-capacity-optimized"
spot_instance_pools = 0
}
launch_template {
launch_template_specification {
launch_template_id = aws_launch_template.spot_lt.id
version = "$Latest"
}
override {
instance_type = "m7i.large"
weighted_capacity = "1"
}
override {
instance_type = "m7a.large"
weighted_capacity = "1"
}
override {
instance_type = "m6i.large"
weighted_capacity = "1"
}
override {
instance_type = "c7i.large"
weighted_capacity = "1"
}
}
}
lifecycle {
ignore_changes = [desired_capacity]
}
} The price-capacity-optimized allocation strategy is critical. It does not simply pick the cheapest option (which leads to frequent interruptions on low-capacity pools). Instead, it selects from pools that have both sufficient capacity and favorable pricing. In practice, this reduces interruption frequency by 40–60% compared to lowest-price. Always include at least 4–6 instance types spanning multiple families (m, c, r generations) to maximize pool diversity.
How should applications handle SIGTERM for graceful shutdown?
When AWS sends the interruption notice, the operating system delivers SIGTERM to PID 1. Your application must catch this signal and begin an orderly shutdown immediately. Applications that ignore SIGTERM get killed by SIGKILL after the 2-minute window, dropping in-flight requests and corrupting partial writes. This is where most Spot failures actually occur—not in detection, but in inadequate shutdown logic.
Implementing a shutdown handler
For Node.js, Python, Go, or Java applications, register a signal handler at startup. The handler must perform three steps in order: stop accepting new connections, wait for in-flight requests to complete (with a timeout), and flush buffers. Here is a minimal Go example suitable for HTTP services:
ctx, stop := signal.NotifyContext(context.Background(), syscall.SIGTERM)
defer stop()
server := &http.Server{Addr: ":8080", Handler: mux}
go func() {
if err := server.ListenAndServe(); err != nil && err != http.ErrServerClosed {
log.Fatalf("listen: %v", err)
}
}()
<-ctx.Done()
log.Println("SIGTERM received, draining connections...")
shutdownCtx, cancel := context.WithTimeout(context.Background(), 90*time.Second)
defer cancel()
if err := server.Shutdown(shutdownCtx); err != nil {
log.Printf("graceful shutdown failed: %v", err)
} else {
log.Println("all connections drained cleanly")
} Note the 90-second timeout. You have 120 seconds total; reserve 30 seconds for OS-level cleanup and deregistration from the load balancer. If your typical request takes longer than 90 seconds, you must implement checkpointing or move those workloads to SQS-backed processors where visibility timeouts handle retries natively. For guidance on structuring resilient background processing, see the article on event-driven architecture on AWS with EventBridge and SQS.
Which workloads are safe for Spot and which are not?
Not every workload belongs on Spot. Making the wrong choice here causes more outages than any misconfiguration. Use this decision framework based on real production experience across dozens of environments:
| Workload Type | Spot Suitability | Rationale |
|---|---|---|
| Stateless web/API servers behind ALB | Excellent | ALB drains connections automatically; no local state to lose |
| Kubernetes worker nodes (non-critical) | Excellent | Pods reschedule; use PDBs and node disruption budgets |
| Batch processing / ETL jobs | Good | Checkpoint progress to S3/DynamoDB; retry on interruption |
| CI/CD build runners | Good | Jobs are idempotent; failed builds auto-retry |
| Primary databases (RDS, self-managed) | Avoid | Failover during interruption risks data loss and extended downtime |
| Single-instance stateful services | Avoid | No redundancy; interruption equals outage |
| Real-time streaming consumers | Caution | Consumer group rebalancing on interruption causes lag spikes |
A reliable pattern is the base-plus-spot model: run 20–30% of your fleet as On-Demand to guarantee minimum capacity, and scale the remaining 70–80% with Spot. Set on_demand_base_capacity in your ASG to match your observed baseline traffic floor. This ensures you never drop below operational minimum even during severe Spot capacity shortages. Teams adopting this approach alongside proper AWS Auto Scaling typically achieve 60–75% blended savings versus all-On-Demand, with zero customer-facing impact.
How do you monitor Spot reliability and optimize over time?
You cannot improve what you do not measure. Track these metrics continuously using CloudWatch or your existing Prometheus and Grafana monitoring stack:
- Spot interruption rate: Count of interruption notices per hour/day. Sustained rates above 5% per hour indicate poor instance type diversity or a problematic pool.
- Rebalance-to-interruption delta: Time between rebalance signal and actual termination. If consistently under 3 minutes, your early-warning automation may not trigger in time.
- Replacement latency: Time from interruption to healthy replacement registered in the load balancer. Target under 3 minutes for web workloads.
- Blended hourly cost: Weighted average across On-Demand and Spot. Trend this weekly to validate savings targets.
Set up CloudWatch alarms on interruption rate. If a specific instance type exceeds thresholds, remove it from your Mixed Instance Policy overrides. AWS also publishes the Spot Placement Score API, which rates each instance-type/AZ combination from 1–10 based on historical capacity availability. Query this weekly via Lambda and feed results into your Terraform variables to preemptively avoid degrading pools. This proactive approach separates teams that sustain Spot savings from those that abandon it after the first bad week.
Running Production on EC2 Spot: Handling Interruptions Safely
Running production on EC2 Spot: Handling Interruptions Safely is not about eliminating risk—it is about making interruptions a non-event through deliberate engineering. Diversify your instance types with Mixed Instance Policies, implement graceful shutdown handlers that respect the 2-minute deadline, maintain an On-Demand base for guaranteed minimum capacity, and instrument everything. When these controls are in place, Spot becomes a reliable, auditable, and cost-effective foundation for production workloads. If your team needs help designing a Spot-resilient architecture or preparing infrastructure for SOC 2 compliance, reach out to discuss your specific environment.