Skip to content

feat: allow EKS Auto Mode nodes to reach ElastiCache/RDS/RDS-proxy - #63

Merged
obezpalko merged 1 commit into
mainfrom
feat/auto-mode-datalayer-sg-access
Aug 11, 2026
Merged

obezpalko merged 1 commit into
mainfrom
feat/auto-mode-datalayer-sg-access

Conversation

@obezpalko

@obezpalko obezpalko commented Aug 11, 2026 •

Copy link
Copy Markdown

User description

Problem

EKS Auto Mode nodes attach the cluster primary security group (module.eks.cluster_primary_security_group_id, e.g. sg-04eea458… / eks-cluster-sg-<cluster>), while managed node groups use the module's node shared SG (node_security_group_id, e.g. sg-00c20cef…). The data-layer SG ingress rules (ElastiCache 6379, RDS 3306) referenced only the managed node SG — so any pod that lands on an Auto Mode node is blocked from Redis/MySQL.

The module already knows these SGs are distinct (comet_eks/main.tf adds node↔cluster coexistence rules gated on enable_auto_mode) but never extended that to the data layer.

Observed on stsaasuat: after adding the arm64 multiarch Auto Mode nodepool, the mysql-db-migration sync hook landed there, its waitForResources looped redis timeout forever → never completed → the whole comet-ml ArgoCD app wedged "waiting for completion of hook … mysql-db-migration". (S3 is fine — Gateway VPC endpoint, route-based.)

Change

  • comet_eks: expose a new output cluster_primary_security_group_id (the SG Auto Mode nodes attach; previously unexposed).
  • comet_elasticache / comet_rds: add *_auto_mode_allow_from_sg variable (string, default null) + a second ingress rule (Redis 6379 / MySQL 3306) created only when the var is set. Mirrors the existing single-SG rule.
  • rds_proxy: allowed_sg_ids is already list/for_each — root now passes both the managed node SG and (when Auto Mode is on) the cluster primary SG.
  • root main.tf: wire all three, guarded by var.enable_eks && var.eks_enable_auto_mode.

Safety

Default null / Auto-Mode-off → zero diff for every existing non-Auto-Mode cluster. On an Auto Mode cluster, terraform plan shows exactly: +1 ElastiCache ingress rule, +1 RDS ingress rule, and the RDS-proxy allow-list gaining the cluster SG — 0 destroy.

Suggested release: v5.5.0. stsaasuat bumps ?ref to it + applies → Auto Mode pods reach Redis/MySQL, the wedged migration hook completes, and the sync (incl. the dply-utils 2.8.0 fix) unblocks.

🤖 Generated with Claude Code


Generated description

Below is a concise technical summary of the changes proposed in this PR:
Enable EKS Auto Mode nodes to access ElastiCache, RDS, and the RDS proxy by exposing the cluster primary security group and wiring it into data-layer access rules. Add conditional Redis/MySQL ingress rules and extend the proxy allow-list while preserving zero changes when Auto Mode is disabled.

TopicDetails
Expose cluster SG Expose cluster_primary_security_group_id from the EKS module so the root configuration can conditionally authorize Auto Mode nodes.
Modified files (1)
  • modules/comet_eks/outputs.tf
Latest Contributors(2)
UserCommitDate
alexb@comet.comfeat: allow EKS Auto M...August 11, 2026
jms200feat(eks): add metrics...April 13, 2026
Auto Mode data access Grant EKS Auto Mode nodes access to Redis, MySQL, and the RDS proxy using the cluster primary security group, while retaining the managed node group permissions.
Modified files (5)
  • main.tf
  • modules/comet_elasticache/main.tf
  • modules/comet_elasticache/variables.tf
  • modules/comet_rds/main.tf
  • modules/comet_rds/variables.tf
Latest Contributors(2)
UserCommitDate
alexb@comet.comfeat: allow EKS Auto M...August 11, 2026
CRThazeMerge pull request #52...July 29, 2026
Review this PR on Baz | Customize your next review

EKS Auto Mode nodes attach the cluster PRIMARY security group
(module.eks.cluster_primary_security_group_id), while managed node groups use the
module's node shared SG (node_security_group_id). The data-layer SG ingress rules only
referenced the managed node SG, so any pod on an Auto Mode node was blocked from Redis
(6379) and MySQL (3306). The module already handles this SG split for node<->node
coexistence (comet_eks/main.tf) but never extended it to the data layer.

Symptom: on a cluster with an Auto Mode nodepool (e.g. stsaasuat's arm64 `multiarch`
pool), a pod that lands there can't reach ElastiCache — e.g. the mysql-db-migration
sync hook's waitForResources loops "redis timeout" forever and wedges the ArgoCD sync.

- comet_eks: expose `cluster_primary_security_group_id` output (the SG Auto Mode nodes use).
- comet_elasticache / comet_rds: add `*_auto_mode_allow_from_sg` var (string, default null)
  + a second ingress rule (Redis 6379 / MySQL 3306) created only when it's set. Mirrors the
  module's existing coexistence-rule pattern.
- rds_proxy: allowed_sg_ids already list/for_each — root now passes both the managed node SG
  and (when Auto Mode is enabled) the cluster primary SG.
- root main.tf: wire all three, guarded by `var.enable_eks && var.eks_enable_auto_mode`.

Default null / Auto-Mode-off => ZERO diff for existing clusters. Suggested release: v5.5.0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@obezpalko
obezpalko requested a review from a team as a code owner August 11, 2026 15:56
@obezpalko obezpalko self-assigned this Aug 11, 2026
@obezpalko
obezpalko merged commit 6e716a0 into main Aug 11, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant