> Markdown version of [/jobs/ext/1452010-software-engineer-3-platform](https://www.wearedevelopers.com/jobs/ext/1452010-software-engineer-3-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer 3, Platform - **Company:** Publicis Groupe - **Location:** Agoura Hills, CA, United States - **Experience:** Experienced - **Salary:** $104,652.0 - **Contract:** Temporary contract - **Skills:** Application Programming Interfaces (APIs), Agile Methodology, Artificial Intelligence, Amazon Web Services, Amazon S3, ARM Architecture, Cloud Computing, Configuration Management, Continuous Delivery, Continuous Integration, Shard (Database Architecture), Monitoring of Systems, Identity and Access Management, Subnetting, Key Management, Cisco Nexus Switches, Octopus Deploy, Pair Programming, Peering, Prometheus, Test-Driven Development (TDD), Grafana, Infrastructure as Code (IaC), Amazon Virtual Private Cloud (VPC), Gitlab, Gitlab-ci, Kubernetes, Route53, Terraform, Docker - **Published:** July 26, 2026 - **Apply:** https://www.juju.com/job/00000000gji7su ## About the Role + AWS networking (VPC, VPC peering, Transit Gateway, Route53, NAT Gateway, security groups, subnet/CIDR design across accounts and regions) + Terraform, AWS (IAM, EKS, S3, EBS) + ArgoCD, GitLab CI/CD, Nexus (artifact registry), Docker, container image build pipelines + Vault, OpenCost + Kubernetes controllers/operators (reconciliation patterns, restart safety) - Go experience is a plus, not required Engineering Practices We Employ: + Agile software development + Infrastructure as Code (IaC) + Pair programming + Test-Driven Development (TDD) + Continuous Delivery, What We Look For: + 3+ years of experience in software or infrastructure engineering + Bachelor's degree or equivalent experience + Hands-on production experience with Kubernetes and at least one major cloud (AWS preferred) + Real operational depth in at least one system we own beyond the cluster - most valuably the observability stack (Prometheus/Alertmanager at scale), but AWS networking, Vault, or artifact/CI infrastructure also count. We are filtering for people who have run these systems, not just used them. + Comfortable owning infrastructure-as-code (Terraform) and CI/CD pipelines + Can reason about tradeoffs and communicate the pros and cons of multiple approaches + Effective communication; thrives in a collaborative, pair-friendly team culture Nice to Have: + AWS networking depth (Transit Gateway, multi-account topology) + Prometheus long-term storage / sharding (Thanos, Cortex, Mimir, or equivalent) + Kubernetes controllers/operators - Go experience is a plus, not required ## Description As a Software Engineer 3 on the Engineering Experience (EngExp) platform team, you help run and evolve the platform that powers CJ's production systems across multiple AWS regions. "Platform" here is broad - it is the Kubernetes clusters, but also the observability stack every squad depends on, the CI/CD and artifact infrastructure their builds run through, the AWS networking that connects them, the secrets and access systems that gate them, and the cost visibility that keeps them accountable. EngExp owns all of it. This is not just an infrastructure role - your value is in engineering judgment: how you evaluate systems, detect risk, and make decisions under uncertainty. You'll own meaningful pieces of these systems independently and drive changes from design through production. We want real depth in the systems below, not just familiarity with the tool names. Responsibilities The Systems You Work On: EngExp owns the systems below. You'll own pieces of them independently and be a credible reviewer of changes to them: + **Observability & monitoring -** Prometheus, Alertmanager, Grafana, and OpenTelemetry across production regions. This is not dashboard-building: you'll own cardinality budgets and recording-rule design, keep a production Prometheus healthy as it outgrows a single shard (federation / sharding / long-term store), and understand Alertmanager HA and the blast radius of alert-routing config. Deep Prometheus and Alertmanager knowledge is a core requirement. + **Kubernetes & cloud infrastructure -** multi-region EKS clusters: upgrades, node group and Karpenter management, controller lifecycle, and add-on / configuration management. Spot failure modes before they happen (subnet IP exhaustion, API server latency, ArgoCD reconciliation lag, Prometheus cardinality, Karpenter consolidation disruption). + **AWS networking -** VPC and subnet design, CIDR management, VPC peering, Route53, security groups, and NAT gateway topology across accounts and regions, plus the 24/7 networking alarms for prod networking between clusters and squad resources. + **CI/CD & artifact management -** GitLab administration (runner fleet, cache, access - not just pipeline authoring), GitOps delivery through ArgoCD, and the Nexus artifact repository including its storage lifecycle as it grows. + **Access & identity -** Vault secrets management, IAM roles and service accounts for apps in clusters, cluster permission management for audit compliance, and AI model access management. Turn recurring access requests into self-service workflows that are hard to misuse. ## Related Videos - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [WeAreDevelopers LIVE - Modern DevOps for IoT Devices and More](https://www.wearedevelopers.com/videos/1805-wearedevelopers-live-modern-devops-for-iot-devices-and-more) - [Transforming Education: A Journey from interactive Markdown to Remote-Labs](https://www.wearedevelopers.com/videos/941-transforming-education-a-journey-from-interactive-markdown-to-remote-labs) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Enabling automated 1-click customer deployments with built-in quality and security](https://www.wearedevelopers.com/videos/83-enabling-automated-1-click-customer-deployments-with-built-in-quality-and-security) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read) - [Dev Digest 131 - AI'm not sure about OSS](https://www.wearedevelopers.com/magazine/472-dev-digest-131-ai-m-not-sure-about-oss)