Skip to content

Commit a906fb3

Browse files
obezpalkoclaude
andauthored
feat(eks): land Chunks C+D on main — module is Kubernetes-API-free (v5.0.0) (#56)
* feat(eks): remove all kubernetes_* resources; module is now k8s-API-free in effect (v5.0.0) Deletes the last in-cluster resources this module created, moving each to its GitOps owner: - storage_class.gp3 / .comet_generic -> comet-infra umbrella (ArgoCD) - namespace.monitoring -> comet-infra umbrella (ArgoCD) - secret.monitoring -> External Secrets Operator (owned) - annotations.app/admin_ns_node_selector -> DROPPED (obsolete under Auto Mode; NodePools/NodeClasses schedule now) - namespace.redis_insights -> agentro-role/rbac local module Also removes time_sleep.wait_for_alb_webhook (only the storage classes / monitoring ns depended on it) and every now-orphaned variable at both module and root level: enable_monitoring_setup, manage_monitoring_secret, monitoring_namespace, create_comet_generic_storage_class, storage_class_reclaim_policy (+ eks_* root aliases), enable_namespace_nodegroup_pinning, app_namespace, admin_pinned_namespaces, enable_redis_insights_ns, and the comet_eks pass-throughs for all of them. Kept: grafana_admin_user / grafana_admin_password root vars — they feed comet_secretsmanager (the AWS Secrets Manager secret ESO reads), not the EKS module. No kubernetes_/helm_/kubectl_ resource remains in the module. The kubernetes provider is now declared-but-unused; it is dropped in the next chunk (D) together with time_sleep.wait_for_cluster_access. terraform init + validate pass. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(eks): drop kubernetes provider — module is Kubernetes-API-free (v5.0.0) With every kubernetes_* resource gone (Chunk C), the kubernetes provider was declared-but-unused. Remove it from required_providers (root + comet_eks) and delete the exec-auth provider "kubernetes" config block in providers.tf (the chicken-and-egg source: it wired the provider to the not-yet-reachable cluster endpoint). Also delete time_sleep.wait_for_cluster_access — the last time_sleep, now with zero dependents. Result: terraform providers lists only aws / tls / time / random (+ cloudinit / null, transitive node-bootstrap deps of terraform-aws-modules/eks). No kubernetes/helm/kubectl anywhere. A step-1 apply on a private-endpoint EKS cluster the runner cannot yet reach now makes ZERO Kubernetes API-server connections; all in-cluster state is owned by ArgoCD (comet-infra) + ESO + the agentro RBAC module. terraform init + validate pass. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * remove terraform.tfvars --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
1 parent 61dea00 commit a906fb3

9 files changed

Lines changed: 45 additions & 386 deletions

File tree

‎README.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -14,6 +14,7 @@ Terraform module for deploying infrastructure components to run CometML.
1414
- Clone the repository to your local machine: `git clone https://github.com/comet-ml/terraform_aws_comet.git`
1515
- Move into the deployment directory: `cd terraform-aws-comet`
1616
- Initialize the directory: `terraform init`
17+
- Copy the example variables file: `cp terraform.tfvars.example terraform.tfvars` (`terraform.tfvars` is gitignored — never commit it)
1718
- Within terraform.tfvars, set your module toggles to enable the desired infrastructure components and set any related inputs
1819
- Provision the resources: `terraform apply`
1920

‎main.tf‎

Lines changed: 0 additions & 18 deletions
Original file line numberDiff line numberDiff line change
@@ -324,10 +324,6 @@ module "comet_eks" {
324324
external_secrets_iam_role_name_override = var.external_secrets_iam_role_name_override
325325
secretsmanager_environment = var.secretsmanager_environment
326326

327-
# Storage class configuration
328-
storage_class_reclaim_policy = var.eks_storage_class_reclaim_policy
329-
create_comet_generic_storage_class = var.eks_create_comet_generic_storage_class
330-
331327
# Loki IRSA for S3 access
332328
enable_loki = var.enable_loki_bucket
333329
loki_s3_bucket_arn = var.enable_s3 && var.enable_loki_bucket ? module.comet_s3[0].comet_loki_bucket_arn : null
@@ -340,13 +336,6 @@ module "comet_eks" {
340336
enable_cloudwatch_exporter = var.enable_cloudwatch_exporter
341337
cloudwatch_exporter_iam_role_name_override = var.cloudwatch_exporter_iam_role_name_override
342338

343-
# Monitoring namespace and Grafana credentials
344-
enable_monitoring_setup = var.enable_monitoring_setup
345-
manage_monitoring_secret = var.manage_monitoring_secret
346-
monitoring_namespace = var.monitoring_namespace
347-
grafana_admin_user = var.grafana_admin_user
348-
grafana_admin_password = var.grafana_admin_password
349-
350339
# EKS Auto Mode (mutually exclusive with Karpenter)
351340
enable_auto_mode = var.eks_enable_auto_mode
352341
auto_mode_node_pools = var.eks_auto_mode_node_pools
@@ -373,13 +362,6 @@ module "comet_eks" {
373362
enable_ci_runners_eks_api_access = var.enable_ci_runners_eks_api_access
374363
ci_runners_cidr = var.ci_runners_cidr
375364

376-
# Namespace nodegroup pinning
377-
enable_namespace_nodegroup_pinning = var.enable_namespace_nodegroup_pinning
378-
app_namespace = var.app_namespace
379-
admin_pinned_namespaces = var.admin_pinned_namespaces
380-
381-
# Redis Insights namespace + agentro port-forward RBAC
382-
enable_redis_insights_ns = var.enable_redis_insights_ns
383365
}
384366

385367
module "comet_elasticache" {

‎modules/comet_eks/main.tf‎

Lines changed: 15 additions & 180 deletions
Original file line numberDiff line numberDiff line change
@@ -583,13 +583,6 @@ module "irsa-ebs-csi" {
583583
}
584584
}
585585

586-
resource "time_sleep" "wait_for_cluster_access" {
587-
count = var.eks_enable_cluster_creator_admin_permissions ? 1 : 0
588-
589-
depends_on = [module.eks]
590-
create_duration = "60s"
591-
}
592-
593586
# State migration: native addons moved out of the eks_blueprints_addons module.
594587
# coredns/kube-proxy/metrics-server now live on module.eks.addons (aws_eks_addon.this),
595588
# vpc-cni is provisioned before_compute (aws_eks_addon.before_compute), and
@@ -730,93 +723,11 @@ resource "aws_iam_role_policy" "external_dns" {
730723
policy = data.aws_iam_policy_document.external_dns[0].json
731724
}
732725

733-
# Best-effort settle delay before this module creates Services/objects that a
734-
# webhook might mutate.
735-
#
736-
# The AWS Load Balancer Controller is now deployed out-of-band by ArgoCD, NOT by
737-
# this module, so its webhook readiness is OUTSIDE Terraform's dependency graph —
738-
# there is no in-graph resource to wait on. We therefore cannot truly gate on the
739-
# ALB webhook here; this is a fixed post-cluster-access delay only. Real ordering
740-
# for anything that depends on the ALB webhook must be enforced ArgoCD-side (sync
741-
# waves / health checks), not here. depends_on is on wait_for_cluster_access
742-
# (which this module DOES own), not on eks_blueprints_addons (which no longer
743-
# installs the controller).
744-
resource "time_sleep" "wait_for_alb_webhook" {
745-
count = var.eks_aws_load_balancer_controller ? 1 : 0
746-
747-
depends_on = [time_sleep.wait_for_cluster_access]
748-
create_duration = "60s"
749-
}
750-
751-
locals {
752-
# Build tag specifications for EBS CSI driver
753-
# Each tag needs to be a separate tagSpecification_N parameter with format "key=value"
754-
# Note: common_tags passed from root module already includes Terraform=true and Environment tags
755-
common_tags_list = [for k, v in var.common_tags : "${k}=${v}"]
756-
757-
# Base tags for gp3 storage class (only StorageClass identifier, other tags come from common_tags)
758-
gp3_base_tags = ["StorageClass=gp3"]
759-
gp3_all_tags = concat(local.gp3_base_tags, local.common_tags_list)
760-
gp3_tag_params = { for idx, tag in local.gp3_all_tags : "tagSpecification_${idx + 1}" => tag }
761-
762-
# Base tags for comet-generic storage class (only StorageClass identifier, other tags come from common_tags)
763-
comet_generic_base_tags = ["StorageClass=comet-generic"]
764-
comet_generic_all_tags = concat(local.comet_generic_base_tags, local.common_tags_list)
765-
comet_generic_tag_params = { for idx, tag in local.comet_generic_all_tags : "tagSpecification_${idx + 1}" => tag }
766-
}
767-
768-
resource "kubernetes_storage_class" "gp3" {
769-
depends_on = [time_sleep.wait_for_cluster_access]
770-
771-
metadata {
772-
name = "gp3"
773-
labels = var.common_tags
774-
annotations = {
775-
"storageclass.kubernetes.io/is-default-class" = "true"
776-
}
777-
}
778-
779-
storage_provisioner = "ebs.csi.aws.com"
780-
781-
parameters = merge(
782-
{
783-
type = "gp3"
784-
# Optionally, set iops and throughput:
785-
# iops = "3000"
786-
# throughput = "125"
787-
},
788-
local.gp3_tag_params
789-
)
790-
791-
reclaim_policy = var.storage_class_reclaim_policy
792-
volume_binding_mode = "WaitForFirstConsumer"
793-
allow_volume_expansion = true
794-
}
795-
796-
resource "kubernetes_storage_class" "comet_generic" {
797-
# Some deployments have comet-generic created by the comet-ml Helm chart
798-
# (Helm/ArgoCD-owned). Set create_comet_generic_storage_class=false there so
799-
# this module does not fight the chart over ownership of the same SC.
800-
count = var.create_comet_generic_storage_class ? 1 : 0
801-
802-
depends_on = [time_sleep.wait_for_cluster_access]
803-
804-
metadata {
805-
name = "comet-generic"
806-
labels = var.common_tags
807-
}
808-
809-
storage_provisioner = "ebs.csi.aws.com"
810-
811-
parameters = merge(
812-
{ type = "gp3" },
813-
local.comet_generic_tag_params
814-
)
815-
816-
reclaim_policy = var.storage_class_reclaim_policy
817-
volume_binding_mode = "WaitForFirstConsumer"
818-
allow_volume_expansion = true
819-
}
726+
# StorageClasses (gp3 default + comet-generic) moved to the comet-infra umbrella
727+
# chart (ArgoCD-owned) — see comet-devops-helm/charts/comet-infra. The former
728+
# wait_for_alb_webhook settle delay went with them; ordering for ALB-webhook
729+
# dependents is enforced ArgoCD-side (sync waves), not in Terraform. This module
730+
# no longer touches the Kubernetes API for storage classes.
820731

821732
#########################################
822733
#### Cluster Autoscaler IRSA Role ####
@@ -1146,41 +1057,10 @@ module "cloudwatch_exporter_irsa_role" {
11461057
#########################################
11471058
#### Monitoring Namespace and Secrets ####
11481059
#########################################
1149-
resource "kubernetes_namespace" "monitoring" {
1150-
count = var.enable_monitoring_setup ? 1 : 0
1151-
1152-
metadata {
1153-
name = var.monitoring_namespace
1154-
}
1155-
1156-
depends_on = [
1157-
module.eks,
1158-
time_sleep.wait_for_alb_webhook
1159-
]
1160-
}
1161-
1162-
resource "kubernetes_secret" "monitoring" {
1163-
# Set manage_monitoring_secret = false where the monitoring Secret is owned by
1164-
# External Secrets Operator (ExternalSecret with creationPolicy: Owner). Letting
1165-
# Terraform also manage it causes a reconcile fight: TF strips ESO's labels and
1166-
# replaces the whole data map (dropping ESO-only keys) on every apply.
1167-
count = var.enable_monitoring_setup && var.manage_monitoring_secret ? 1 : 0
1168-
1169-
metadata {
1170-
name = "monitoring"
1171-
namespace = kubernetes_namespace.monitoring[0].metadata[0].name
1172-
}
1173-
1174-
data = {
1175-
grafana-admin-user = var.grafana_admin_user
1176-
grafana-admin-password = var.grafana_admin_password
1177-
}
1178-
1179-
type = "Opaque"
1180-
immutable = false
1181-
1182-
depends_on = [kubernetes_namespace.monitoring]
1183-
}
1060+
# The monitoring namespace moved to the comet-infra umbrella chart (ArgoCD-owned).
1061+
# The monitoring Secret is owned by External Secrets Operator (ExternalSecret with
1062+
# creationPolicy: Owner) — Terraform no longer creates it. Both are out of this
1063+
# module so it never touches the Kubernetes API for monitoring bootstrap.
11841064

11851065
#########################################
11861066
#### Karpenter Prerequisites ####
@@ -1626,58 +1506,13 @@ resource "aws_vpc_security_group_ingress_rule" "eks_api" {
16261506
#########################################
16271507
#### Namespace nodegroup pinning ####
16281508
#########################################
1629-
# Annotates namespaces with scheduler.alpha.kubernetes.io/node-selector to route
1630-
# every Pod admitted into the namespace onto a specific node group. Skips
1631-
# kube-system + monitoring (they host DaemonSets and must schedule everywhere).
1632-
# Patches in place — does NOT create the namespace; create via Helm first.
1633-
1634-
resource "kubernetes_annotations" "app_ns_node_selector" {
1635-
count = var.enable_namespace_nodegroup_pinning ? 1 : 0
1636-
1637-
depends_on = [time_sleep.wait_for_cluster_access]
1638-
1639-
api_version = "v1"
1640-
kind = "Namespace"
1641-
metadata {
1642-
name = coalesce(var.app_namespace, var.environment)
1643-
}
1644-
annotations = {
1645-
"scheduler.alpha.kubernetes.io/node-selector" = "nodegroup_name=comet"
1646-
}
1647-
force = true
1648-
}
1649-
1650-
resource "kubernetes_annotations" "admin_ns_node_selector" {
1651-
for_each = var.enable_namespace_nodegroup_pinning ? toset(var.admin_pinned_namespaces) : []
1652-
1653-
depends_on = [time_sleep.wait_for_cluster_access]
1654-
1655-
api_version = "v1"
1656-
kind = "Namespace"
1657-
metadata {
1658-
name = each.value
1659-
}
1660-
annotations = {
1661-
"scheduler.alpha.kubernetes.io/node-selector" = "nodegroup_name=admin"
1662-
}
1663-
force = true
1664-
}
1509+
# DROPPED. The scheduler.alpha.kubernetes.io/node-selector annotations pinned
1510+
# app/admin namespaces onto legacy managed node groups (nodegroup_name=…). Under
1511+
# EKS Auto Mode, scheduling is handled by NodePools/NodeClasses (comet-infra), so
1512+
# this in-cluster patching is obsolete and has been removed.
16651513

16661514
#########################################
16671515
#### Redis Insights namespace ####
16681516
#########################################
1669-
# Operational debug surface — provides a namespace for the redis-insights helm
1670-
# chart (installed by FRED-helm-apply) pinned to the admin NG.
1671-
1672-
resource "kubernetes_namespace" "redis_insights" {
1673-
count = var.enable_redis_insights_ns ? 1 : 0
1674-
1675-
depends_on = [time_sleep.wait_for_cluster_access]
1676-
1677-
metadata {
1678-
name = "redis-insights"
1679-
annotations = {
1680-
"scheduler.alpha.kubernetes.io/node-selector" = "nodegroup_name=admin"
1681-
}
1682-
}
1683-
}
1517+
# Moved to the agentro-role/rbac local module (comet-devops), which owns agentro's
1518+
# in-cluster objects in one place. Not created by this module anymore.

0 commit comments

Comments
 (0)