Skip to content

AWS Batch on ECS Managed Instances for GPU Workloads: When to Choose It Over EKS

On August 25, 2026, AWS Batch added support for Amazon ECS Managed Instances. It works in every region where Batch is available.

Here is why this matters. Until now, GPU batch jobs on AWS meant one of two things. You managed EC2 yourself, with AMIs, NVIDIA drivers, and launch templates. Or you ran a Kubernetes cluster. Fargate never got GPUs. People asked for seven years. AWS closed that issue in December 2025 and pointed everyone to ECS Managed Instances instead.

Now Batch can use it too. You get a job queue, GPUs from the full EC2 catalog, and a bill that goes to zero when the queue is empty. AWS owns the instances, the patching, and the drivers. In this post I'll go over how it works, what it costs, and when EKS is still the right call.

TL;DR

  • Fargate-style operations with EC2 capabilities: GPUs, any instance type, Spot, privileged containers.
  • The management fee is roughly 12% of On-Demand. For GPUs it is lower: about 7.8% on G-series and 4.8% on P-series and Trainium after the July 2026 price cut. EKS Auto Mode charges the same per-instance fees, so price is not the deciding factor.
  • Pick it for single-node GPU jobs: inference backfills, transcription, rendering, fine-tuning on one node.
  • Stay on EKS for distributed training, Capacity Blocks for ML, custom AMIs, or if your GPUs live next to services already on Kubernetes.

What you had before

AWS Batch itself is free. You pay for the compute under it. The problem was that every compute option had a catch:

Compute environment GPUs Who runs the nodes The catch
Fargate ❌ AWS No GPUs, max 32 vCPU per job
EC2 (classic) ✅ You AMIs, drivers, launch templates
EKS ✅ You Bring and operate your own cluster
ECS Managed Instances ✅ AWS New. This post.

What ECS Managed Instances is

It launched in September 2025. Think of it as a middle mode between Fargate and self-managed EC2. AWS provisions, scales, patches, and terminates EC2 instances in your account. You see them on your EC2 bill. You just can't touch them.

The details that matter:

  • Bottlerocket only, no SSH. AWS-owned AMIs with an immutable root filesystem. No custom AMIs. Use ECS Exec if you need a shell.
  • Instances live 14 to 21 days. ECS starts draining every instance at day 14 and kills it by day 21. That is how patching works. Same model as EKS Auto Mode nodes.
  • You pick instance types, or ECS does. Pin explicit types and families, or describe what you need (vCPU, memory, GPU count) and ECS picks the cheapest match.
  • Purchase options: On-Demand, Spot (added December 2025), and Capacity Reservations (February 2026). RIs and Savings Plans apply to the EC2 part automatically. Capacity Blocks for ML are not supported. That is still an open roadmap item.

For GPUs, the managed AMI ships with NVIDIA drivers and CUDA preinstalled. ECS runs DCGM health checks and replaces instances with failing GPUs on its own. GPU metrics show up in CloudWatch Container Insights. The supported list covers basically everything: G4dn, G5, G6, G6e, G6f (fractional L4 slices), P4d, P5, P6-B200, Trn1, Inf2, and more.

How Batch uses it

You create a compute environment with type: ECS_MANAGED_INSTANCES. All infrastructure config lives in one managedInstancesProvider block. Here is a GPU example from the Batch docs:

create-compute-environment.json
{
  "computeEnvironmentName": "gpu-managed-instances-ce",
  "type": "MANAGED",
  "state": "ENABLED",
  "computeResources": {
    "type": "ECS_MANAGED_INSTANCES",
    "maxvCpus": 1000,
    "managedInstancesProvider": {
      "infrastructureRoleArn": "arn:aws:iam::123456789012:role/ecsInfrastructureRole",
      "instanceLaunchTemplate": {
        "ec2InstanceProfileArn": "arn:aws:iam::123456789012:instance-profile/ecsInstanceProfile",
        "networkConfiguration": {
          "subnets": ["subnet-abcde012", "subnet-bcde012a"],
          "securityGroups": ["sg-abcde012"]
        },
        "instanceRequirements": {
          "allowedInstanceTypes": ["g5.xlarge", "g5.2xlarge", "g5.4xlarge"]
        },
        "capacityOptionType": "ON_DEMAND"
      }
    }
  }
}
aws batch create-compute-environment \
  --cli-input-json file://create-compute-environment.json

What you should know about this config:

  • Spot is a flag, not a type. Set capacityOptionType: SPOT. You cannot change it after creation. Switching means a new compute environment.
  • Scale-to-zero is built in. infrastructureOptimization.scaleInAfter controls how long an idle instance lives, from 0 to 3600 seconds. Empty queue means no instances and no bill.
  • A lot of familiar knobs are gone. No allocationStrategy, no minvCpus, no imageId, no launchTemplate, no placementGroup. No warm pool. ECS decides how to allocate Spot.
  • maxvCpus counts vCPUs requested by jobs, not instance vCPUs. ECS bin-packs tasks, so provisioned capacity can exceed what maxvCpus implies. Don't treat it as a hard cost cap.

Job definitions use the ecsProperties format with platformCapabilities: ["MANAGED_INSTANCES"]. Networking is host mode only. You also get everything Fargate blocks: privileged, host devices, tmpfs, ulimits, host path mounts.

register-job-definition.json
{
  "jobDefinitionName": "gpu-inference-job",
  "type": "container",
  "platformCapabilities": ["MANAGED_INSTANCES"],
  "ecsProperties": {
    "taskProperties": [
      {
        "networkMode": "host",
        "executionRoleArn": "arn:aws:iam::123456789012:role/ecsTaskExecutionRole",
        "containers": [
          {
            "name": "inference",
            "image": "123456789012.dkr.ecr.us-east-1.amazonaws.com/batch-inference:latest",
            "command": ["python", "run_batch.py"],
            "resourceRequirements": [
              { "type": "VCPU", "value": "4" },
              { "type": "MEMORY", "value": "16384" },
              { "type": "GPU", "value": "1" }
            ],
            "logConfiguration": { "logDriver": "awslogs" }
          }
        ]
      }
    ]
  }
}

Array jobs, retries, dependencies, and fair-share scheduling all work as usual. Two queue rules: you can't mix platform types in one queue, and On-Demand compute environments must come before Spot ones.

The flow looks like this:

aws batch submit-job
        │
        ▼
   job queue ──► compute environment (ECS_MANAGED_INSTANCES)
                        │
                        ▼
         ECS capacity provider (managedInstancesProvider)
                        │  launch · scale · bin-pack
                        ▼
        ┌───────────────────────────────────────┐
        │  your VPC, your EC2 bill              │
        │                                       │
        │  Bottlerocket + NVIDIA driver + CUDA  │
        │  g5.xlarge    g5.2xlarge    ...       │
        └───────────────────────────────────────┘
          AWS handles patching, the 14-day
          recycle, and GPU health checks

What it costs

You pay the normal EC2 price plus a per-instance management fee. Both are billed per second. Batch itself stays free.

Two things people get wrong about the fee:

  1. There is no flat $0.02/hour fee. That number is just the m5a.xlarge example from the pricing page. The fee varies by instance type. 922 types have one in us-east-1.
  2. The fee is never discounted. Same rate on Spot, Reserved, and Savings Plan capacity. On Spot it becomes a bigger share of your actual spend. If you have heavy RI coverage, the fee is pure overhead.

Real numbers from the Price List API (us-east-1, Aug 27, 2026), after the July 2026 GPU fee cut:

Instance On-Demand Management fee Fee as % of OD
m5.large $0.096/hr $0.0115/hr 12%
g4dn.xlarge $0.526/hr $0.0410/hr ~7.8%
g5.xlarge $1.006/hr $0.0785/hr ~7.8%
g6.xlarge $0.805/hr $0.0628/hr ~7.8%
p4d.24xlarge $21.958/hr $1.0540/hr ~4.8%
p5.48xlarge $55.04/hr $2.6419/hr ~4.8%
trn1.32xlarge $21.50/hr $1.032/hr ~4.8%

Quick example. A nightly backfill on 8 g5.xlarge instances, 6 hours a night, is 1,440 instance hours a month. About $1,449 for EC2 and $113 in fees. For that $113 you delete an AMI pipeline, driver patching, and all the scaling glue. At the top end the fees are real money though. One p5.48xlarge running 24/7 is about $1,900 a month in fees alone.

And here is the punchline for the EKS comparison. EKS Auto Mode charges the same per-instance fees for the same instance types. The EKS control plane adds $73 a month. So the cost difference between the two managed stacks is basically one cluster fee. This decision is about capabilities and operations, not price.

What EKS gives you that this does not

There are three ways to run GPU batch on EKS, and they are very different amounts of work.

DIY with Karpenter. Karpenter for GPU nodes, accelerated AMIs, the NVIDIA device plugin (the AL2023 GPU AMI does not include it), plus a queueing layer like Kueue or Volcano, because plain Kubernetes Jobs don't queue. Add the version treadmill: 14 months of support per Kubernetes minor, then extended support at 6x the price. Very capable. All yours to run.

EKS Auto Mode. The closest sibling. Same locked-down node model, drivers and device plugin baked in, scale-to-zero GPU node pools. And it supports Capacity Blocks for ML, which Managed Instances does not. You still own the cluster and everything above the nodes.

Batch on EKS. Batch as a free layer on your cluster. The headline feature: gang-scheduled multi-node parallel jobs for distributed training. Managed Instances does not support MNP at all. But you still own the cluster lifecycle and AMI updates.

When to choose what

Pick Batch on ECS Managed Instances when:

  • Jobs fit on one node: inference backfills, transcription, rendering, simulations, fine-tuning.
  • You want zero infrastructure ownership. No cluster, no AMIs, no drivers, automatic GPU failure replacement.
  • Cold starts in minutes are fine. Fresh capacity takes 3 to 4 minutes. AWS measured about 13 minutes end to end for a scale-from-zero GPU stack with image pulls.
  • Your team is not already running Kubernetes.

Pick EKS when:

  • You need multi-node distributed training. This is the hard line. No gang scheduling on Managed Instances.
  • You need Capacity Blocks for ML to get H100 or B200 capacity at all.
  • You need custom AMIs, specific driver versions, placement groups, or EFA tuning.
  • You need MIG partitioning or time-slicing beyond the G6f fractional instances.
  • Your GPUs share a platform with services already on Kubernetes.
Requirement Batch on ECS MI EKS
Single-node GPU jobs with a queue ✅ best fit ✅ more moving parts
Distributed training ❌ ✅
Capacity Blocks for ML ❌ ✅
Custom AMIs and drivers ❌ ✅
Scale-to-zero GPUs ✅ built in ✅
No cluster to operate ✅ ❌

Gotchas

  • Always set allowedInstanceTypes for GPU workloads. Unconstrained, ECS optimizes cost across the whole catalog and can pick types you did not expect.
  • The 14-day recycle caps job length. If a job can't checkpoint and resume, don't plan on weeks of runtime.
  • capacityOptionType is permanent per compute environment. Create separate On-Demand and Spot environments up front.
  • Log drivers: only awslogs, splunk, and awsfirelens. Networking is host mode only.
  • Allowed AMIs SCPs need the region-specific ECS AMI accounts allowlisted, or nothing launches.
  • Quotas: 50 compute environments per account, 3 per job queue, array jobs up to 10,000.

My take

For years, "we have some GPU jobs" turned into "we now operate Kubernetes." That was never a good trade if all you needed was a queue, some GPUs, and a bill that stops when the work stops. This closes that gap. It is now the least infrastructure you can own to run GPUs in production on AWS.

The line is capability, not cost. Need distributed training, Capacity Blocks, or custom node software? Use EKS. Already have a healthy EKS platform? Adding a separate silo next to it needs a better reason than novelty. For everything else, I would start here.

I plan to benchmark the cold-start and Spot behavior in a follow-up. If your team is weighing this decision, that's what we do.

Sources