Software Engineer 3, Platform
Role details
Job location
Tech stack
Job description
As a Software Engineer 3 on the Engineering Experience (EngExp) platform team, you help run and evolve the platform that powers CJ's production systems across multiple AWS regions. "Platform" here is broad - it is the Kubernetes clusters, but also the observability stack every squad depends on, the CI/CD and artifact infrastructure their builds run through, the AWS networking that connects them, the secrets and access systems that gate them, and the cost visibility that keeps them accountable. EngExp owns all of it. This is not just an infrastructure role - your value is in engineering judgment: how you evaluate systems, detect risk, and make decisions under uncertainty. You'll own meaningful pieces of these systems independently and drive changes from design through production. We want real depth in the systems below, not just familiarity with the tool names.
Responsibilities
The Systems You Work On:
EngExp owns the systems below. You'll own pieces of them independently and be a credible reviewer of changes to them:
-
Observability & monitoring - Prometheus, Alertmanager, Grafana, and OpenTelemetry across production regions. This is not dashboard-building: you'll own cardinality budgets and recording-rule design, keep a production Prometheus healthy as it outgrows a single shard (federation / sharding / long-term store), and understand Alertmanager HA and the blast radius of alert-routing config. Deep Prometheus and Alertmanager knowledge is a core requirement.
-
Kubernetes & cloud infrastructure - multi-region EKS clusters: upgrades, node group and Karpenter management, controller lifecycle, and add-on / configuration management. Spot failure modes before they happen (subnet IP exhaustion, API server latency, ArgoCD reconciliation lag, Prometheus cardinality, Karpenter consolidation disruption).
-
AWS networking - VPC and subnet design, CIDR management, VPC peering, Route53, security groups, and NAT gateway topology across accounts and regions, plus the 24/7 networking alarms for prod networking between clusters and squad resources.
-
CI/CD & artifact management - GitLab administration (runner fleet, cache, access - not just pipeline authoring), GitOps delivery through ArgoCD, and the Nexus artifact repository including its storage lifecycle as it grows.
-
Access & identity - Vault secrets management, IAM roles and service accounts for apps in clusters, cluster permission management for audit compliance, and AI model access management. Turn recurring access requests into self-service workflows that are hard to misuse.
Requirements
-
AWS networking (VPC, VPC peering, Transit Gateway, Route53, NAT Gateway, security groups, subnet/CIDR design across accounts and regions)
-
Terraform, AWS (IAM, EKS, S3, EBS)
-
ArgoCD, GitLab CI/CD, Nexus (artifact registry), Docker, container image build pipelines
-
Vault, OpenCost
-
Kubernetes controllers/operators (reconciliation patterns, restart safety) - Go experience is a plus, not required
Engineering Practices We Employ:
-
Agile software development
-
Infrastructure as Code (IaC)
-
Pair programming
-
Test-Driven Development (TDD)
-
Continuous Delivery, What We Look For:
-
3+ years of experience in software or infrastructure engineering
-
Bachelor's degree or equivalent experience
-
Hands-on production experience with Kubernetes and at least one major cloud (AWS preferred)
-
Real operational depth in at least one system we own beyond the cluster - most valuably the observability stack (Prometheus/Alertmanager at scale), but AWS networking, Vault, or artifact/CI infrastructure also count. We are filtering for people who have run these systems, not just used them.
-
Comfortable owning infrastructure-as-code (Terraform) and CI/CD pipelines
-
Can reason about tradeoffs and communicate the pros and cons of multiple approaches
-
Effective communication; thrives in a collaborative, pair-friendly team culture
Nice to Have:
-
AWS networking depth (Transit Gateway, multi-account topology)
-
Prometheus long-term storage / sharding (Thanos, Cortex, Mimir, or equivalent)
-
Kubernetes controllers/operators - Go experience is a plus, not required
Benefits & conditions
Apart from offering competitive salaries, 401K matching, wellness programs, and comprehensive medical, dental, and vision coverage, we provide:
-
Flexible time off without the hassle of accrual
-
A generous number of paid holidays
-
Company-sponsored team-building events
-
An Employee Referral Program
-
Annual recognition awards
-
Hybrid work arrangements for optimal work-life balance
-
Parental bonding leave
-
Backup care options for children and elders
-
An employee discount program
-
International SOS program for global support
-
Business Resource Groups, where employees connect over shared interests to cultivate an engaging, inclusive environment
*and those are just a few of our great perks! Come join us and see what makes our company a great place to work., Compensation Range: USD $8,721.00 - USD $131,230.00/Annually. This is the pay range the Company believes it will pay for this position at the time of this posting. Consistent with applicable law, compensation will be determined based on the skills, qualifications, and experience of the applicant along with the requirements of the position, and the Company reserves the right to modify this pay range at any time. Temporary roles may be eligible to participate in our freelancer/temporary employee medical plan through a third-party benefits administration system once certain criteria have been met. Temporary roles may also qualify for participation in our 401(k) plan after eligibility criteria have been met. For regular roles, the Company will offer medical coverage, dental, vision, disability, 401k, and paid time off. The Company anticipates the application deadline for this job posting will be 8/28/2026.