Introduction
PipeCD has served our deployment needs well, particularly for advanced ECS rollouts. As part of a broader internal platform change, we’ve decided to move our continuous delivery to thinner GitHub Actions workflows. The migration covered several deployment targets. In this post, we focus on our ECS Fargate services behind an Application Load Balancer, and PipeCD has been driving the deployments. The moment we removed PipeCD, we had to answer a question: Who drives the ECS deployment now? The answer is the deployment controller: PipeCD had been deploying them through the EXTERNAL controller; after removing PipeCD, GitHub Actions needed the native ECS controller instead. AWS can migrate a service between controllers in place, but our target operating model is Terraform-managed infrastructure. With the Terraform AWS provider version we were using, switching the controller type planned a service recreation.
For live, traffic-serving services, naive recreation can drop requests. This post explains how we weighed the trade-offs and switched controllers with minimal disruption.
Why change the ECS deployment controller?
ECS lets us choose who orchestrates a deployment. For this migration, PipeCD and GitHub Actions sat on opposite sides of that choice.
PipeCD needs the EXTERNAL controller. With EXTERNAL, ECS steps back and hands deployment orchestration to an outside tool. PipeCD uses TaskSet to create parallel revisions, shift target-group traffic, run canary and blue/green rollouts, and roll back by manipulating those primitives directly. For a deeper walkthrough, let’s see ECS External Deployment & TaskSet – Complete Guide.
Our GitHub Actions workflow uses the native ECS controller. Once PipeCD was gone, there was no external orchestrator left, so we let ECS manage deployments natively and chose rolling updates. The important feature for us was the deployment circuit breaker: if a rollout fails its health checks, ECS can roll back to the last healthy task definition.
In practice, our GitHub Actions workflow renders the new image into the task definition, then uses aws-actions/amazon-ecs-deploy-task-definition@v2 to register the task definition and update the ECS service. The workflow can wait for service stability, but it does not own the deployment lifecycle or shift traffic itself.

Fig. GitHub Actions continuous delivery flow
That is the real migration: from EXTERNAL, where PipeCD orchestrates everything, to ECS, where AWS drives the rollout. The two controllers cannot run on the same service at the same time.
The challenges of switching the controller
Two problems sat on top of each other:
First, the services were PipeCD-owned, not Terraform-owned. In our setup, PipeCD created and managed the ECS services directly through the ECS APIs, based on a service definition (servicedef.yaml) stored in a configuration repository. Terraform was not part of that deployment path. In the new setup, Terraform would own the ECS service configuration, while GitHub Actions would deploy new task-definition revisions. Before Terraform could safely manage a service, we needed a complete definition that matched production reality.
Second, our Terraform version treated the switch as a replacement. In the Terraform AWS Provider version we were using (v5.80.0), deployment_controller.type was a ForceNew field, so a Terraform-only switch from EXTERNAL to ECS was a replacement rather than an in-place edit.
So we chose a simple rule: if the production service is healthy, don’t mutate it; build the replacement beside it.
AWS does support in-place controller migration via UpdateService, and newer provider versions later relaxed the replacement behavior (v6.4.0). But both the in-place/API path and the provider-upgrade path still required full Terraform service definitions and state reconciliation for our PipeCD-owned live services. Upgrading also meant crossing a major provider-version boundary. We didn’t want to add that major bump to a production controller cutover.
Instead, we created a new ECS-controller service directly in Terraform, attached it to the same load balancer and service discovery path, validated it under real traffic, and only then removed the old EXTERNAL service.
Implementing the parallel cutover
The rule is to minimize the gap. Run the old and new services side by side behind the same load balancer, and only retire the old one once the new one has proven itself. Concretely, six steps:
- Stop dual-driving. Disable/Stop the app delivery on the PipeCD control plane first. Exactly one system should be in control at a time.
- Define the full service in Terraform. Describe the complete service (load balancer, service discovery, networking, lifecycle) so the recreation is deliberate and reviewable.
- Bootstrap the parallel service. Create a parallel
api-v2service on theECSdeployment controller, attached to the same target group and Cloud Map service as the originalEXTERNAL-controlledapi. Once its tasks pass health checks, the ALB starts routing to them. - Validate under live traffic. Watch
api-v2error rate, latency, and the deployment circuit breaker as it serves real requests. If anything looks wrong, revert it and let the old service carry on. - Tear down the old service. Once
api-v2is trusted, tear down the originalEXTERNALservice. The ALB drains its targets, andapi-v2carries 100%, without a gap, because it was already serving. - Drop the
-v2suffix. Rename back toapiin Terraform. Thanks tocreate_before_destroy, Terraform brings the renamed service up and waits for it to be healthy before destroyingapi-v2, so even this cleanup rename causes no user-visible interruption.

Fig. ECS controller cutover
This pattern assumes the service is safe to run in parallel. For a stretch, because of running two copies in parallel, it is recommended to do the cutover during a low-traffic window. Make sure to have enough Fargate capacity, database connection headroom, no singleton jobs that misbehave when duplicated, and compatible request handling across both versions.
We rehearsed the workflow in dev before prod, including IAM permissions and deployment mechanics, before any of it touched customers.
Key Terraform settings for safe migration
The parallel replacement is only safe because of a few Terraform settings. Here’s the service definition, with the rest of the resource trimmed out:
resource "aws_ecs_service" "api" {
name = "api${var.service_name_suffix}" # "-v2" during the cutover
task_definition = var.task_definition_family
desired_count = var.desired_count
wait_for_steady_state = true # keep sync state until tasks are healthy
deployment_controller {
type = var.deployment_controller_type # "ECS" after the switch (was "EXTERNAL")
}
# Native auto-rollback (replaces PipeCD's), only valid for the ECS controller.
dynamic "deployment_circuit_breaker" {
for_each = var.deployment_controller_type == "ECS" ? [1] : []
content {
enable = true
rollback = true
}
}
...
lifecycle {
create_before_destroy = true # new healthy before old dies
ignore_changes = [task_definition, desired_count] # allowed changes not re-create resources
}
}
Four settings did most of the work.
create_before_destroy = truekeeps the old service alive while Terraform creates the replacement.wait_for_steady_state = truemakes Terraform wait until the replacement service is healthy before it considers creation complete.ignore_changes = [task_definition, desired_count]keeps Terraform from fighting the systems that own those fields after the migration: GitHub Actions owns task-definition updates, and autoscaling owns desired count.deployment_circuit_breaker { rollback = true }replaces PipeCD’s rollback orchestration. If an ECS deployment fails health checks, ECS rolls back to the last successful deployment.
One operational detail: IAM permissions usually need iteration. A role that can run terraform plan may still fail on apply, especially once Terraform starts creating or updating ECS services and load balancer attachments. Start minimal, expand as errors surface, and verify the result with aws ecs describe-services: deploymentController.type should be ECS, and the service should be steady.
Summary
The controller switch looked like a risky service recreation, but the rule stayed simple: don’t mutate a healthy production service if you can build the replacement beside it.
Our repeatable recipe was:
- Make sure one system drives the service.
- Own the full service in Terraform.
- Create a parallel service on the new controller.
- Let it serve real traffic, then remove the old service and rename it back to its original under
create_before_destroy.
The pattern is not specific to deployment controllers. The -v2 + create_before_destroy approach works well for many live resources that Terraform treats as replacements, as long as the workload can safely run in parallel. If you’re staring down a similar EXTERNAL → ECS cutover, we hope this provides a safer path than editing the live service.
