From 47845c76d966266e044b6201766017027a8da9c8 Mon Sep 17 00:00:00 2001 From: Guy Saar Date: Tue, 1 Sep 2026 10:16:02 -0500 Subject: [PATCH 1/6] DND-1573: v6 brownfield migration support (removed.tf + provider requirements) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Basis for the v6.0.1-migration tag — the one-apply stepping stone the remaining 11 brownfield envs need to reach v6.0.0. v5.6.2-migration cannot be reused. It predates DND-875, so it lacks rds_auto_minor_version_upgrade (which porsche, si, waystar and zoox all pass) plus rds_parameter_group_family, rds_use_proxy_endpoint and rds_proxy_ack_no_iam_auth. Those envs fail on an unsupported argument before the removed{} blocks ever run. Cut from v6.0.0 rather than v5.7.0 — both carry the full DND-875 surface, but starting at v6 means Stage 1 lands on the final surface and Stage 2 is only "drop the providers map", with no second behaviour change (the DND-1522 toggle removal and redis_vpn re-index happen once, in Stage 1). - modules/comet_eks/removed.tf: 13 removed{} blocks, all destroy = false. Every address verified present on both v2.1.2 and v1.20.11. 6 shared: blueprints{alb,cert_manager,external_dns}, storage_class.gp3, namespace.monitoring, secret.monitoring 1 zoox: blueprints.aws_cloudwatch_metrics 6 v1.20.x: helm_release{karpenter_stsaas,external_secrets, external_secrets_crds}, storage_class.comet_generic, annotations{admin,app}_ns_node_selector - kubernetes + helm requirements restored in both versions.tf files (requirements only — no provider config blocks; the wrapper passes its own). - MIGRATION.md: restored and rewritten for the v6 sequence. Coverage checked against real state: waystar 6/6, zoox 5/5, si 6 of 11, porsche 4 of 11. The remainder on si and porsche sit at the STATE ROOT (agentro_extras, agentro_view, agentro_portforward, and on porsche also namespace/secret.monitoring) — not module-addressed, so they are removed{} in the wrapper. Documented in MIGRATION.md and in the file header. terraform validate passes. Co-Authored-By: Claude Opus 5 (1M context) --- .gitignore | 1 + MIGRATION.md | 95 ++++++++++++++++++++++ modules/comet_eks/removed.tf | 149 ++++++++++++++++++++++++++++++++++ modules/comet_eks/versions.tf | 14 ++++ versions.tf | 14 ++++ 5 files changed, 273 insertions(+) create mode 100644 MIGRATION.md create mode 100644 modules/comet_eks/removed.tf diff --git a/.gitignore b/.gitignore index 0059dc7..d1c794f 100644 --- a/.gitignore +++ b/.gitignore @@ -37,3 +37,4 @@ override.tf.json # Ignore CLI configuration files .terraformrc terraform.rc +*.local.md diff --git a/MIGRATION.md b/MIGRATION.md new file mode 100644 index 0000000..95920e7 --- /dev/null +++ b/MIGRATION.md @@ -0,0 +1,95 @@ +# Brownfield migration to v6.0.0 (DND-1573 / DND-1257) + +For clusters still on a v1.20.x or v2.1.x module version. Greenfield clusters and +anything already on v5.x/v6.x do not need this — bump straight to `v6.0.0`. + +## Why a temporary tag + +The v5 "Infra-Only + ArgoCD" refactor deleted the module's in-cluster and +eks-blueprints-addons resources; GitOps and native EKS add-ons own them now. +A cluster upgrading from v1/v2 still carries those objects in state, and because +the module also dropped its kubernetes/helm provider *config blocks*, `plan` fails +with `Provider configuration not present` for every one of them. + +`v6.0.1-migration` is `v6.0.0` plus: + +- `modules/comet_eks/removed.tf` — 13 `removed{}` blocks, all `destroy = false` +- `kubernetes` + `helm` back in both `versions.tf` files (requirements only, no + provider config blocks) + +It is consumed for exactly ONE apply per cluster, then the wrapper moves to the +permanent `v6.0.0`. + +Do not use `v5.6.2-migration` for this. It predates DND-875 and lacks +`rds_auto_minor_version_upgrade`, `rds_parameter_group_family`, +`rds_use_proxy_endpoint` and `rds_proxy_ack_no_iam_auth` — envs that pass any of +them fail on an unsupported argument before the removed blocks run. + +## Stage 1 — drop the orphans + +Wrapper: + +```hcl +module "comet" { + source = "github.com/comet-ml/terraform-aws-comet-stsaas?ref=v6.0.1-migration" + + providers = { + aws = aws + kubernetes = kubernetes + helm = helm + } + ... +} +``` + +The `providers` map is required: the orphans bind to +`module.comet.provider["...{aws,kubernetes,helm}"]`, and this populates that with +the wrapper's real assume_role/region. An empty `provider "aws" {}` inside the +module would hijack credentials instead (AccessDenied). + +Also in the same PR: + +- **Root-level orphans** — porsche and si carry `kubernetes_cluster_role.agentro_extras`, + `kubernetes_cluster_role_binding.{agentro_extras,agentro_view}` and + `kubernetes_role{,_binding}.agentro_portforward` at the STATE ROOT. They are not + module-addressed, so add `removed{}` blocks for them in the WRAPPER. +- **`mysql_vpn` import** — v6/DND-1522 creates the VPN→MySQL rule unconditionally, + but every env already has it out of state from the DND-752 era. Without an + `import{}` the apply fails `InvalidPermission.Duplicate`. See the rule-id table in + the rollout plan. +- **Drop deleted vars** — `enable_argocd_management_eks_access`, + `enable_vpn_eks_api_access`, `enable_ci_runners_eks_api_access`, + `enable_vpn_redis_access`, `enable_monitoring_setup`, `monitoring_namespace`, + `eks_create_comet_generic_storage_class`, `enable_redis_insights_ns`. + +Apply. The objects leave state; live infrastructure is untouched. + +## Stage 1.5 — GitOps adoption + +comet-infra (ArgoCD) / native EKS add-ons must own the ALB controller, cert-manager, +external-dns, the gp3 StorageClass and the monitoring namespace/secret. Coordinate +this with Stage 1 — it is the real risk in the sequence, not the terraform. + +## Stage 2 — land on the permanent tag + +```hcl +source = "github.com/comet-ml/terraform-aws-comet-stsaas?ref=v6.0.0" +``` + +Drop the `providers` map and `helm` from the wrapper's `required_providers`. **Keep +`kubernetes`** if the wrapper's own modules use it — `agentro_k8s_rbac` and +`automation_smoke_rbac` own kubernetes resources outside `module.comet`, and dropping +the provider strands them. + +Delete the wrapper's one-shot `imports.tf` and any Stage 1 `removed{}` blocks once +applied. + +Expected plan: no infrastructure change. Stage 1 already landed on the v6 surface. + +## Order + +Canary **waystar** — its orphan set is exactly the six blocks proven against bayer, +with no root-level extras and no `aws_cloudwatch_metrics`. Then zoox, si, porsche. +Group C (v1.20.x: circuit, circuit-dev, eonnext, fetch, mercedesamgf1, netflix) +after Group B is complete; those additionally hit the Karpenter-Helm → EKS Auto Mode +migration, which is its own piece of work. diff --git a/modules/comet_eks/removed.tf b/modules/comet_eks/removed.tf new file mode 100644 index 0000000..fc0d29a --- /dev/null +++ b/modules/comet_eks/removed.tf @@ -0,0 +1,149 @@ +# Brownfield state migration (v1.x/v2.x → v6.x) — DND-1573 / DND-1257. +# +# ⚠️ THIS FILE MAKES THIS A TEMPORARY "MIGRATION" MODULE VERSION. It is meant to be +# consumed for exactly ONE apply per brownfield cluster, then the wrapper bumps to +# the permanent v6.0.0 tag WITHOUT this file. See MIGRATION.md at the repo root. +# +# Supersedes v5.6.2-migration, which cannot be used by the remaining brownfield envs: +# it predates DND-875, so it lacks rds_auto_minor_version_upgrade (which porsche, si, +# waystar and zoox all pass) plus rds_parameter_group_family, rds_use_proxy_endpoint +# and rds_proxy_ack_no_iam_auth. Those envs would fail on an unsupported argument +# before the removed{} blocks below ever ran. +# +# Cut from v6.0.0 rather than v5.7.0 so Stage 1 lands directly on the v6 surface — +# Stage 2 is then only "drop the providers map", with no second behaviour change. +# +# The v5 "Infra-Only + ArgoCD" refactor DELETED these in-cluster / eks-blueprints-addons +# resources from the module (they are now owned by GitOps / native EKS add-ons). Clusters +# upgrading from a v1/v2 module version still carry them in state, and because the module +# root's kubernetes/helm/aws provider *config blocks* were also removed (child-module +# design), `terraform plan` fails with "Provider configuration not present" for all of +# them until the objects leave state. +# +# HOW THIS WORKS (proven against bayer, DND-1573 / #2205): +# - These `removed` blocks drop the objects from state WITHOUT destroying the live +# resources (`lifecycle { destroy = false }`) — the live workloads are adopted by +# comet-infra (ArgoCD) / native EKS add-ons as part of the coordinated cutover. +# - The orphans bind in state to `module.comet.provider["...{aws,kubernetes,helm}"]`. +# To satisfy that binding the module declares the kubernetes+helm *requirements* +# (versions.tf) — but declares NO provider config blocks. Instead the WRAPPER passes +# its fully-configured providers in: +# module "comet" { +# providers = { aws = aws, kubernetes = kubernetes, helm = helm } +# } +# This populates module.comet.provider[...] with the wrapper's real assume_role/region +# (an empty `provider "aws" {}` here would instead HIJACK credentials → AccessDenied). +# - Greenfield v6 clusters never had these resources, so the blocks are a harmless no-op. +# +# NOT COVERED HERE — root-level orphans. porsche and si additionally carry +# kubernetes_cluster_role.agentro_extras, kubernetes_cluster_role_binding.agentro_extras, +# kubernetes_cluster_role_binding.agentro_view, kubernetes_role.agentro_portforward and +# kubernetes_role_binding.agentro_portforward at the STATE ROOT, not under module.comet. +# They are not module-addressed, so they must be removed{} in the WRAPPER, not here. +# +# NOTE: do NOT `removed` the whole `module.eks_blueprints_addons` — its `aws_eks_addon.this[*]` +# children (coredns/kube-proxy/vpc-cni/metrics-server/ebs-csi) are relocated via the `moved` +# blocks in main.tf. Only the Helm-based sub-modules below were deleted. + +# --- eks-blueprints-addons Helm sub-modules --- +# Removing the sub-module call drops all of its resources (aws_iam_policy/role/attachment + +# helm_release). ALB controller → comet-infra (ArgoCD); cert-manager + external-dns → native +# EKS managed add-ons (see main.tf module.eks.addons). +removed { + from = module.eks_blueprints_addons.module.aws_load_balancer_controller + lifecycle { + destroy = false + } +} + +removed { + from = module.eks_blueprints_addons.module.cert_manager + lifecycle { + destroy = false + } +} + +removed { + from = module.eks_blueprints_addons.module.external_dns + lifecycle { + destroy = false + } +} + +# zoox only — enabled there via eks_aws_cloudwatch_metrics, which v6 removed. Harmless +# no-op on every other cluster. +removed { + from = module.eks_blueprints_addons.module.aws_cloudwatch_metrics + lifecycle { + destroy = false + } +} + +# --- In-cluster resources now owned by comet-infra (ArgoCD) --- +# gp3 StorageClass + monitoring namespace/secret moved to the comet-infra Helm chart. +removed { + from = kubernetes_storage_class.gp3 + lifecycle { + destroy = false + } +} + +removed { + from = kubernetes_namespace.monitoring + lifecycle { + destroy = false + } +} + +removed { + from = kubernetes_secret.monitoring + lifecycle { + destroy = false + } +} + +# --- v1.20.x-only in-cluster resources (Group C) --- +# Present on circuit, circuit-dev, eonnext, fetch, mercedesamgf1 and netflix; a no-op on +# the v2.1.x envs. Karpenter-via-Helm is superseded by EKS Auto Mode, external-secrets and +# the comet-generic StorageClass by comet-infra GitOps. +removed { + from = helm_release.karpenter_stsaas + lifecycle { + destroy = false + } +} + +removed { + from = helm_release.external_secrets + lifecycle { + destroy = false + } +} + +removed { + from = helm_release.external_secrets_crds + lifecycle { + destroy = false + } +} + +removed { + from = kubernetes_storage_class.comet_generic + lifecycle { + destroy = false + } +} + +removed { + from = kubernetes_annotations.admin_ns_node_selector + lifecycle { + destroy = false + } +} + +removed { + from = kubernetes_annotations.app_ns_node_selector + lifecycle { + destroy = false + } +} diff --git a/modules/comet_eks/versions.tf b/modules/comet_eks/versions.tf index 3060ce7..f80d535 100644 --- a/modules/comet_eks/versions.tf +++ b/modules/comet_eks/versions.tf @@ -11,5 +11,19 @@ terraform { source = "hashicorp/time" version = ">= 0.9" } + # kubernetes + helm are declared ONLY so brownfield clusters (upgrading from a + # v1/v2 module version) can associate their leftover in-cluster / helm_release + # resources with a provider long enough for the `removed` blocks in removed.tf to + # drop them from state (destroy = false). v6 no longer CREATES any kubernetes/helm + # resource — greenfield clusters never instantiate these providers. Removed again + # in the permanent v6.0.0 tag (DND-1573 / DND-1257 brownfield migration). + kubernetes = { + source = "hashicorp/kubernetes" + version = ">= 2.0" + } + helm = { + source = "hashicorp/helm" + version = ">= 2.0" + } } } diff --git a/versions.tf b/versions.tf index cd3ae62..60cf9df 100644 --- a/versions.tf +++ b/versions.tf @@ -13,5 +13,19 @@ terraform { source = "hashicorp/random" version = ">= 3.0" } + # kubernetes + helm are declared ONLY so brownfield clusters (upgrading from a + # v1/v2 module version) can associate their leftover in-cluster / helm_release + # resources with a provider long enough for the `removed` blocks in removed.tf to + # drop them from state (destroy = false). v6 no longer CREATES any kubernetes/helm + # resource — greenfield clusters never instantiate these providers. Removed again + # in the permanent v6.0.0 tag (DND-1573 / DND-1257 brownfield migration). + kubernetes = { + source = "hashicorp/kubernetes" + version = ">= 2.0" + } + helm = { + source = "hashicorp/helm" + version = ">= 2.0" + } } } From 66d964df456f362d0995f14c6592c74661e7ada3 Mon Sep 17 00:00:00 2001 From: Guy Saar Date: Tue, 1 Sep 2026 11:22:12 -0500 Subject: [PATCH 2/6] fix: omit compute_config when Auto Mode is off (unblocks 10 of 13 envs) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit waystar's v6 apply (#2213) failed on: Error: updating EKS Cluster (waystar-use1) Auto Mode settings: InvalidRequestException: Cannot modify EKS Auto Mode configuration. Auto Mode is not enabled on this cluster. comet_eks sent compute_config unconditionally, so a cluster with enable_auto_mode = false got an explicit { enabled = false }. EKS rejects that on a cluster that never had Auto Mode — it is not a no-op, it is an error, and it fails the whole UpdateClusterConfig. Only bayer and stsaasuat set eks_enable_auto_mode = true. The other ten envs leave it unset, so every one of them would have hit this on its v6 bump. It did not surface on bayer (#2205) precisely because bayer has Auto Mode on. Fix: pass null when disabled. The upstream module guards with `for_each = var.compute_config != null ? [var.compute_config] : []`, so null omits the block rather than sending a disable. Trade-off, kept deliberate: the previous comment argued the block should always be sent so a cluster with Auto Mode on could be turned back off. That still works, but now needs an explicit one-off (set enabled = false for that apply, or disable out of band). Turning Auto Mode off is the rarer operation, and unlike this failure it is not silent — you are doing it on purpose. validate passes. Co-Authored-By: Claude Opus 5 (1M context) --- modules/comet_eks/main.tf | 23 +++++++++++++++-------- 1 file changed, 15 insertions(+), 8 deletions(-) diff --git a/modules/comet_eks/main.tf b/modules/comet_eks/main.tf index 9a85785..808dae4 100644 --- a/modules/comet_eks/main.tf +++ b/modules/comet_eks/main.tf @@ -244,14 +244,21 @@ module "eks" { # EKS Auto Mode. When enabled, the control plane provisions nodes via the # built-in node pools and the upstream module auto-creates/wires the Auto Mode - # node IAM role (so node_role_arn is intentionally omitted). The block is - # always sent — enabled = false explicitly disables Auto Mode so a cluster that - # previously had it on can be turned back off (a bare null would omit the - # argument and leave the last-applied config in place). - compute_config = { - enabled = var.enable_auto_mode - node_pools = var.enable_auto_mode ? var.auto_mode_node_pools : [] - } + # node IAM role (so node_role_arn is intentionally omitted). + # + # null when disabled, NOT { enabled = false }. Sending an explicit disable to a + # cluster that never had Auto Mode makes EKS reject the whole UpdateClusterConfig + # with "Cannot modify EKS Auto Mode configuration. Auto Mode is not enabled on + # this cluster." That failed waystar's v6 apply (#2213) and would break every env + # where eks_enable_auto_mode is unset — 10 of 13. + # + # Turning Auto Mode back OFF on a cluster that has it therefore needs a + # deliberate one-off: set enabled = false here for that apply, or disable it out + # of band. That is the rarer operation, and unlike this failure it is not silent. + compute_config = var.enable_auto_mode ? { + enabled = true + node_pools = var.auto_mode_node_pools + } : null # Bake the Karpenter discovery tag directly into the node SG so it is never # dropped when Terraform modifies the security group during subsequent applies. From ed517f2073738b4b31cadadedfa797aa111f0523 Mon Sep 17 00:00:00 2001 From: Guy Saar Date: Tue, 1 Sep 2026 11:25:00 -0500 Subject: [PATCH 3/6] docs: point MIGRATION.md at v6.0.1-migration-2 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit v6.0.1-migration is abandoned — it shipped compute_config unconditionally and EKS rejects that on clusters that never had Auto Mode (comet-devops#2213). Tags are immutable, so the fix ships as a new tag rather than moving the old one. Co-Authored-By: Claude Opus 5 (1M context) --- MIGRATION.md | 8 ++++++-- 1 file changed, 6 insertions(+), 2 deletions(-) diff --git a/MIGRATION.md b/MIGRATION.md index 95920e7..fb096ac 100644 --- a/MIGRATION.md +++ b/MIGRATION.md @@ -11,7 +11,7 @@ A cluster upgrading from v1/v2 still carries those objects in state, and because the module also dropped its kubernetes/helm provider *config blocks*, `plan` fails with `Provider configuration not present` for every one of them. -`v6.0.1-migration` is `v6.0.0` plus: +`v6.0.1-migration-2` is `v6.0.0` plus: - `modules/comet_eks/removed.tf` — 13 `removed{}` blocks, all `destroy = false` - `kubernetes` + `helm` back in both `versions.tf` files (requirements only, no @@ -20,6 +20,10 @@ with `Provider configuration not present` for every one of them. It is consumed for exactly ONE apply per cluster, then the wrapper moves to the permanent `v6.0.0`. +`v6.0.1-migration` (without the `-2`) is abandoned — it shipped `compute_config` +unconditionally, which EKS rejects on any cluster that never had Auto Mode. Do not +use it. + Do not use `v5.6.2-migration` for this. It predates DND-875 and lacks `rds_auto_minor_version_upgrade`, `rds_parameter_group_family`, `rds_use_proxy_endpoint` and `rds_proxy_ack_no_iam_auth` — envs that pass any of @@ -31,7 +35,7 @@ Wrapper: ```hcl module "comet" { - source = "github.com/comet-ml/terraform-aws-comet-stsaas?ref=v6.0.1-migration" + source = "github.com/comet-ml/terraform-aws-comet-stsaas?ref=v6.0.1-migration-2" providers = { aws = aws From f663c8831f75c24054e07bb1239dd41bc4f411bb Mon Sep 17 00:00:00 2001 From: Guy Saar Date: Tue, 1 Sep 2026 11:29:22 -0500 Subject: [PATCH 4/6] docs: v6.0.1-migration-3 is the usable migration tag MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit v6.0.1-migration-2 is deleted — it was cut after v6.0.1-migration had already been force-moved onto the same commit, so the two were identical and the 'superseded' note was wrong. v6.0.1-migration is now restored to its original commit (47845c7) so the compare against -3 shows the real one-file fix. Co-Authored-By: Claude Opus 5 (1M context) --- MIGRATION.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/MIGRATION.md b/MIGRATION.md index fb096ac..68e4b54 100644 --- a/MIGRATION.md +++ b/MIGRATION.md @@ -11,7 +11,7 @@ A cluster upgrading from v1/v2 still carries those objects in state, and because the module also dropped its kubernetes/helm provider *config blocks*, `plan` fails with `Provider configuration not present` for every one of them. -`v6.0.1-migration-2` is `v6.0.0` plus: +`v6.0.1-migration-3` is `v6.0.0` plus: - `modules/comet_eks/removed.tf` — 13 `removed{}` blocks, all `destroy = false` - `kubernetes` + `helm` back in both `versions.tf` files (requirements only, no @@ -35,7 +35,7 @@ Wrapper: ```hcl module "comet" { - source = "github.com/comet-ml/terraform-aws-comet-stsaas?ref=v6.0.1-migration-2" + source = "github.com/comet-ml/terraform-aws-comet-stsaas?ref=v6.0.1-migration-3" providers = { aws = aws From eceea3d17c06f3738f118e48c62a8649b1b68ac1 Mon Sep 17 00:00:00 2001 From: Guy Saar Date: Tue, 1 Sep 2026 11:29:34 -0500 Subject: [PATCH 5/6] docs: fix stale -2 reference in the abandoned-tag note Co-Authored-By: Claude Opus 5 (1M context) --- MIGRATION.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/MIGRATION.md b/MIGRATION.md index 68e4b54..4dc3740 100644 --- a/MIGRATION.md +++ b/MIGRATION.md @@ -20,7 +20,7 @@ with `Provider configuration not present` for every one of them. It is consumed for exactly ONE apply per cluster, then the wrapper moves to the permanent `v6.0.0`. -`v6.0.1-migration` (without the `-2`) is abandoned — it shipped `compute_config` +`v6.0.1-migration` (the unsuffixed tag) is abandoned — it shipped `compute_config` unconditionally, which EKS rejects on any cluster that never had Auto Mode. Do not use it. From 65220e6335b222cff3c49020a3900154b7494f7c Mon Sep 17 00:00:00 2001 From: Guy Saar Date: Tue, 1 Sep 2026 13:34:54 -0500 Subject: [PATCH 6/6] fix: OVERWRITE resolve_conflicts on cert-manager/external-dns (forward-port 4b6b30e) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit waystar's apply cleared the Auto Mode error and then failed on the next one: Error: waiting for EKS Add-On (waystar-use1:cert-manager) create: CREATE_FAILED ... ConfigurationConflict: Conflicts found when trying to apply. Will not continue due to resolve conflicts mode. Same for external-dns. On a brownfield cluster the native EKS add-on refuses to take over the objects the eks_blueprints_addons Helm release still owns, because module.eks defaults resolve_conflicts_on_create = NONE. Alex fixed exactly this for bayer in 4b6b30e, but that commit only ever existed on the v5.6.2-migration side branch — it was never merged to main, so v6.0.0 lost it. v6 had 1 occurrence of resolve_conflicts_on_create (the ebs_csi addon); v5.6.2-migration had 3. This restores the other two verbatim, comments included. The native add-on adopts the Helm-owned objects in place, relabelling them managed-by=EKS; pods keep running. Harmless on greenfield. Operational note carried over from 4b6b30e: resolve_conflicts does NOT cover EKS Pod Identity associations. A pre-existing external-dns:external-dns association must be deleted before apply or it fails with ResourceInUseException. Ships as v6.0.1-migration-4; -3 is superseded. Co-Authored-By: Claude Opus 5 (1M context) --- MIGRATION.md | 4 ++-- modules/comet_eks/main.tf | 15 +++++++++++++++ 2 files changed, 17 insertions(+), 2 deletions(-) diff --git a/MIGRATION.md b/MIGRATION.md index 4dc3740..c1b0b2e 100644 --- a/MIGRATION.md +++ b/MIGRATION.md @@ -11,7 +11,7 @@ A cluster upgrading from v1/v2 still carries those objects in state, and because the module also dropped its kubernetes/helm provider *config blocks*, `plan` fails with `Provider configuration not present` for every one of them. -`v6.0.1-migration-3` is `v6.0.0` plus: +`v6.0.1-migration-4` is `v6.0.0` plus: - `modules/comet_eks/removed.tf` — 13 `removed{}` blocks, all `destroy = false` - `kubernetes` + `helm` back in both `versions.tf` files (requirements only, no @@ -35,7 +35,7 @@ Wrapper: ```hcl module "comet" { - source = "github.com/comet-ml/terraform-aws-comet-stsaas?ref=v6.0.1-migration-3" + source = "github.com/comet-ml/terraform-aws-comet-stsaas?ref=v6.0.1-migration-4" providers = { aws = aws diff --git a/modules/comet_eks/main.tf b/modules/comet_eks/main.tf index 808dae4..5f63cc9 100644 --- a/modules/comet_eks/main.tf +++ b/modules/comet_eks/main.tf @@ -319,6 +319,14 @@ module "eks" { } : {}, { configuration_values = local.coredns_config + # OVERWRITE so a brownfield cluster (upgrading from the eks_blueprints_addons + # Helm cert-manager) lets the native add-on ADOPT the existing Helm-owned + # objects (SAs/CRDs/Deployments/webhooks) instead of failing with + # "ConfigurationConflict … resolve conflicts mode" (default NONE). The add-on + # relabels them managed-by=EKS in place — cert-manager keeps running. Harmless + # on greenfield (nothing to conflict). DND-1573. + resolve_conflicts_on_create = "OVERWRITE" + resolve_conflicts_on_update = "OVERWRITE" } ) } : {}, @@ -335,6 +343,13 @@ module "eks" { role_arn = aws_iam_role.external_dns[0].arn service_account = "external-dns" }] + # OVERWRITE so a brownfield cluster adopts any external-dns k8s objects left by + # the old eks_blueprints_addons Helm release rather than conflict-failing. NOTE: + # a pre-existing Pod Identity association for external-dns:external-dns must be + # deleted first (ResourceInUseException otherwise) — resolve_conflicts does not + # cover Pod Identity associations. DND-1573. + resolve_conflicts_on_create = "OVERWRITE" + resolve_conflicts_on_update = "OVERWRITE" }, var.eks_external_dns_addon_version != null ? { addon_version = var.eks_external_dns_addon_version