Running Production on EC2 Spot: Handling Interruptions Safely

Khimananda Oli 8 min read DevOps
Running Production on EC2 Spot: Handling Interruptions Safely

By Khimananda Oli | Last reviewed: September 2026

Running production on EC2 Spot: Handling Interruptions Safely is the single most effective way to reduce AWS compute spend, but it requires treating capacity as ephemeral rather than guaranteed. Spot Instances offer up to 90% savings over On-Demand pricing, yet they can be reclaimed by AWS with only a two-minute warning when capacity is needed elsewhere. To use them reliably in production, you must architect your application to absorb these interruptions gracefully through automated lifecycle hooks, diversified allocation strategies, and stateless design patterns.

Spot InstanceProcessing WorkloadRebalance Signal~5-10 min before reclaim(Early Warning)Interruption Notice2-min SIGTERM countdown(Hard Deadline)Graceful Shutdown• Drain connections• Complete in-flight tasks• Deregister from LBASG Replacement• Launch new instance• Health check passes• Register with LBState PreservedExternal DB / CacheNo local data lossTimeline flows left to right; vertical drop shows action trigger
EC2 Spot interruption lifecycle: from rebalance signal through graceful shutdown to ASG replacement

How do you detect EC2 Spot interruptions before termination?

Detection is the foundation of safe Spot usage. AWS provides two distinct signals, and confusing them is a common mistake I see in audits. The Rebalance Recommendation arrives first—typically 5 to 10 minutes before actual reclamation—and indicates that your instance is at elevated risk. This is your window to proactively replace capacity. The Interruption Notice follows later and gives you exactly 2 minutes before forced termination. Relying solely on the interruption notice leaves almost no time for complex cleanup.

Polling instance metadata for signals

Your application or a sidecar process must poll the Instance Metadata Service (IMDSv2) every 5 seconds. Do not use IMDSv1; it lacks the token-based security required for SOC 2 compliance and is disabled on many hardened AMIs. Here is a production-ready polling script:

TOKEN=$(curl -X PUT "http://169.254.169.254/latest/api/token" \
  -H "X-aws-ec2-metadata-token-ttl-seconds: 21600")

while true; do
  # Check for rebalance recommendation (early signal)
  REBALANCE=$(curl -s -o /dev/null -w "%{http_code}" \
    -H "X-aws-ec2-metadata-token: $TOKEN" \
    http://169.254.169.254/latest/meta-data/events/recommendations/rebalance)

  if [ "$REBALANCE" = "200" ]; then
    echo "$(date -Iseconds) REBALANCE signal received — initiating proactive drain"
    systemctl stop myapp.service
    aws autoscaling complete-lifecycle-action \
      --auto-scaling-group-name prod-spot-asg \
      --lifecycle-hook-name spot-drain-hook \
      --instance-id $(curl -s -H "X-aws-ec2-metadata-token: $TOKEN" \
        http://169.254.169.254/latest/meta-data/instance-id) \
      --lifecycle-action-result CONTINUE
    break
  fi

  # Check for hard interruption notice (2-min warning)
  INTERRUPTION=$(curl -s -H "X-aws-ec2-metadata-token: $TOKEN" \
    http://169.254.169.254/latest/meta-data/spot/instance-action)

  if [ -n "$INTERRUPTION" ] && [ "$INTERRUPTION" != "none" ]; then
    echo "$(date -Iseconds) INTERRUPTION notice: $INTERRUPTION"
    kill -SIGTERM $(pgrep myapp)
    break
  fi

  sleep 5
done

For Kubernetes workloads on EKS, you do not need custom scripts. Install the AWS Node Termination Handler via Helm; it watches both signals and automatically cordons nodes, drains pods respecting PodDisruptionBudgets, and completes ASG lifecycle actions. This is the standard pattern for containerized Spot workloads in 2026.

What is the correct Auto Scaling Group configuration for Spot?

The single biggest reliability lever is the Mixed Instance Policy. Never configure an ASG with a single Spot instance type. When that specific type experiences a capacity shortage, every instance in your fleet gets interrupted simultaneously—a correlated failure that overwhelms your remaining On-Demand base. Instead, define multiple instance types and let AWS allocate from whichever pool has surplus capacity.

Terraform mixed instance policy example

resource "aws_autoscaling_group" "prod_spot" {
  name                = "prod-web-spot"
  desired_capacity    = 6
  min_size            = 3
  max_size            = 12
  vpc_zone_identifier = var.private_subnet_ids

  mixed_instances_policy {
    instances_distribution {
      on_demand_base_capacity                  = 2
      on_demand_percentage_above_base_capacity = 0
      spot_allocation_strategy                 = "price-capacity-optimized"
      spot_instance_pools                      = 0
    }

    launch_template {
      launch_template_specification {
        launch_template_id = aws_launch_template.spot_lt.id
        version            = "$Latest"
      }

      override {
        instance_type     = "m7i.large"
        weighted_capacity = "1"
      }
      override {
        instance_type     = "m7a.large"
        weighted_capacity = "1"
      }
      override {
        instance_type     = "m6i.large"
        weighted_capacity = "1"
      }
      override {
        instance_type     = "c7i.large"
        weighted_capacity = "1"
      }
    }
  }

  lifecycle {
    ignore_changes = [desired_capacity]
  }
}

The price-capacity-optimized allocation strategy is critical. It does not simply pick the cheapest option (which leads to frequent interruptions on low-capacity pools). Instead, it selects from pools that have both sufficient capacity and favorable pricing. In practice, this reduces interruption frequency by 40–60% compared to lowest-price. Always include at least 4–6 instance types spanning multiple families (m, c, r generations) to maximize pool diversity.

Auto Scaling GroupMixed Instance Policym7i.largeIntel · Gen 7m7a.largeAMD · Gen 7m6i.largeIntel · Gen 6c7i.largeCompute · Gen 7us-east-1a2 Spot + 1 ODus-east-1b2 Spot + 1 ODus-east-1c2 Spot + 0 ODDiversification across families, generations, and AZs prevents correlated interruptions
Mixed Instance Policy distributing Spot capacity across four instance types and three Availability Zones

How should applications handle SIGTERM for graceful shutdown?

When AWS sends the interruption notice, the operating system delivers SIGTERM to PID 1. Your application must catch this signal and begin an orderly shutdown immediately. Applications that ignore SIGTERM get killed by SIGKILL after the 2-minute window, dropping in-flight requests and corrupting partial writes. This is where most Spot failures actually occur—not in detection, but in inadequate shutdown logic.

Implementing a shutdown handler

For Node.js, Python, Go, or Java applications, register a signal handler at startup. The handler must perform three steps in order: stop accepting new connections, wait for in-flight requests to complete (with a timeout), and flush buffers. Here is a minimal Go example suitable for HTTP services:

ctx, stop := signal.NotifyContext(context.Background(), syscall.SIGTERM)
defer stop()

server := &http.Server{Addr: ":8080", Handler: mux}

go func() {
    if err := server.ListenAndServe(); err != nil && err != http.ErrServerClosed {
        log.Fatalf("listen: %v", err)
    }
}()

<-ctx.Done()
log.Println("SIGTERM received, draining connections...")

shutdownCtx, cancel := context.WithTimeout(context.Background(), 90*time.Second)
defer cancel()

if err := server.Shutdown(shutdownCtx); err != nil {
    log.Printf("graceful shutdown failed: %v", err)
} else {
    log.Println("all connections drained cleanly")
}

Note the 90-second timeout. You have 120 seconds total; reserve 30 seconds for OS-level cleanup and deregistration from the load balancer. If your typical request takes longer than 90 seconds, you must implement checkpointing or move those workloads to SQS-backed processors where visibility timeouts handle retries natively. For guidance on structuring resilient background processing, see the article on event-driven architecture on AWS with EventBridge and SQS.

Which workloads are safe for Spot and which are not?

Not every workload belongs on Spot. Making the wrong choice here causes more outages than any misconfiguration. Use this decision framework based on real production experience across dozens of environments:

Workload TypeSpot SuitabilityRationale
Stateless web/API servers behind ALBExcellentALB drains connections automatically; no local state to lose
Kubernetes worker nodes (non-critical)ExcellentPods reschedule; use PDBs and node disruption budgets
Batch processing / ETL jobsGoodCheckpoint progress to S3/DynamoDB; retry on interruption
CI/CD build runnersGoodJobs are idempotent; failed builds auto-retry
Primary databases (RDS, self-managed)AvoidFailover during interruption risks data loss and extended downtime
Single-instance stateful servicesAvoidNo redundancy; interruption equals outage
Real-time streaming consumersCautionConsumer group rebalancing on interruption causes lag spikes

A reliable pattern is the base-plus-spot model: run 20–30% of your fleet as On-Demand to guarantee minimum capacity, and scale the remaining 70–80% with Spot. Set on_demand_base_capacity in your ASG to match your observed baseline traffic floor. This ensures you never drop below operational minimum even during severe Spot capacity shortages. Teams adopting this approach alongside proper AWS Auto Scaling typically achieve 60–75% blended savings versus all-On-Demand, with zero customer-facing impact.

$0$2k$4k$6k$8k$7,400All On-DemandSpot $1,200OD Base$1,480Base + SpotSpot $800OD $1,480Optimized Mix↓ 69% savingsMonthly cost for 8-instance web fleet
Monthly cost comparison: all On-Demand vs base-plus-Spot vs optimized mix for an 8-instance production web fleet

How do you monitor Spot reliability and optimize over time?

You cannot improve what you do not measure. Track these metrics continuously using CloudWatch or your existing Prometheus and Grafana monitoring stack:

  • Spot interruption rate: Count of interruption notices per hour/day. Sustained rates above 5% per hour indicate poor instance type diversity or a problematic pool.
  • Rebalance-to-interruption delta: Time between rebalance signal and actual termination. If consistently under 3 minutes, your early-warning automation may not trigger in time.
  • Replacement latency: Time from interruption to healthy replacement registered in the load balancer. Target under 3 minutes for web workloads.
  • Blended hourly cost: Weighted average across On-Demand and Spot. Trend this weekly to validate savings targets.

Set up CloudWatch alarms on interruption rate. If a specific instance type exceeds thresholds, remove it from your Mixed Instance Policy overrides. AWS also publishes the Spot Placement Score API, which rates each instance-type/AZ combination from 1–10 based on historical capacity availability. Query this weekly via Lambda and feed results into your Terraform variables to preemptively avoid degrading pools. This proactive approach separates teams that sustain Spot savings from those that abandon it after the first bad week.

Running Production on EC2 Spot: Handling Interruptions Safely

Running production on EC2 Spot: Handling Interruptions Safely is not about eliminating risk—it is about making interruptions a non-event through deliberate engineering. Diversify your instance types with Mixed Instance Policies, implement graceful shutdown handlers that respect the 2-minute deadline, maintain an On-Demand base for guaranteed minimum capacity, and instrument everything. When these controls are in place, Spot becomes a reliable, auditable, and cost-effective foundation for production workloads. If your team needs help designing a Spot-resilient architecture or preparing infrastructure for SOC 2 compliance, reach out to discuss your specific environment.

Frequently Asked Questions

Only if your application handles interruptions gracefully. Use EFS or S3 for persistent data, implement checkpointing, and ensure tasks are idempotent. Never store local state without replication or backup mechanisms that survive instance termination.

AWS provides a two-minute warning via the instance metadata service before termination. Monitor the termination notice endpoint at 169.254.169.254/latest/meta-data/spot/termination-time to trigger graceful shutdown scripts and drain connections automatically.

Deploy the AWS Node Termination Handler or Karpenter to detect interruption signals. These tools cordon nodes, drain pods gracefully, and reschedule workloads before termination completes, maintaining cluster stability during Spot reclaim events.

No. Spot Instances use identical hardware and networking as On-Demand. Performance differences only occur due to workload interference from other tenants, which affects all instance types equally regardless of purchase option.

Use mixed instances policies in Auto Scaling Groups with a base capacity of On-Demand instances and percentage-based Spot allocation. Define multiple instance types and generations to maximize availability pools and reduce interruption risk significantly.

Root volumes delete by default unless configured otherwise. Attached EBS volumes persist independently if marked for retention. Always set delete-on-termination flags explicitly and use volume snapshots or AMIs for critical data preservation during interruptions.

No. Interruptions are inherent to Spot pricing. Instead, design fault-tolerant architectures using diversified instance types, automated recovery, and health checks to maintain uptime despite reclamation events occurring unpredictably.

Spot placement scores rate capacity pools from one to ten based on historical interruption rates. Higher scores indicate lower reclamation probability. Use this metric when selecting instance types and Availability Zones for production deployments.

CloudWatch Events captures EC2 Spot interruption warnings natively. Combine with Prometheus node exporter metrics and custom Lambda functions polling metadata endpoints to create multi-layered detection systems triggering pre-termination automation workflows reliably.

Yes, for steady baseline traffic. Reserve On-Demand or Savings Plans cover consistent demand at reduced cost. Use Spot only for variable, interruptible, or batch processing workloads where temporary capacity loss is acceptable.

Simulate interruptions using the ec2-spot-labs simulator or manually send SIGTERM signals to containers. Validate that applications shut down cleanly, queues drain properly, and replacement instances join load balancers within acceptable recovery windows.

Grant ec2:DescribeInstances, autoscaling:CompleteLifecycleAction, and ssm:SendCommand permissions minimally. Avoid wildcard actions. Use instance profiles rather than embedded credentials so terminated instances cannot retain access after reclamation.

Yes. Diversifying across instance families, generations, and sizes expands eligible capacity pools. Auto Scaling selects available options dynamically, reducing single-pool dependency and lowering overall interruption probability for production fleets.

Replacement depends on capacity availability and launch template configuration. With diversified instance types and warm pools enabled, replacements typically launch within thirty seconds. Undiversified configurations may experience delays exceeding several minutes during constrained periods.

You pay only for seconds used until termination. Partial hours round up to full seconds billed. Failed tasks requiring restart consume additional compute but incur no penalty fees beyond normal usage charges for replacement capacity.