> Markdown version of [/jobs/ext/1447358-senior-software-engineer-platform](https://www.wearedevelopers.com/jobs/ext/1447358-senior-software-engineer-platform). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Senior Software Engineer, Platform - **Company:** Publicis Groupe - **Location:** Agoura Hills, CA, United States - **Experience:** Expert - **Salary:** $110,580.0 - **Contract:** Temporary contract - **Skills:** Application Programming Interfaces (APIs), Agile Methodology, Artificial Intelligence, Amazon Web Services, Amazon Cloudfront, Amazon S3, ARM Architecture, Cloud Computing, Configuration Management, Continuous Delivery, Continuous Integration, Domain Name System (DNS), Monitoring of Systems, Identity and Access Management, Subnetting, Key Management, Cisco Nexus Switches, Octopus Deploy, Pair Programming, Peering, Role-Based Access Control, Prometheus, Policy as Code, Test-Driven Development (TDD), Grafana, Infrastructure as Code (IaC), Amazon Virtual Private Cloud (VPC), Gitlab, Gitlab-ci, Kubernetes, Route53, Terraform, Webhooks, Docker - **Published:** July 26, 2026 - **Apply:** https://www.juju.com/job/00000000gji818 ## About the Role + Terraform, AWS (IAM, EKS, S3, EBS) + ArgoCD, GitLab CI/CD, Nexus (artifact registry), Docker, container image build pipelines + Vault, OpenCost + Gateway API / Kgateway + Kubernetes controllers/operators (reconciliation patterns, restart safety) - Go experience is a plus, not required Engineering Practices We Employ: + Agile software development + Infrastructure as Code (IaC) + Pair programming + Test-Driven Development (TDD) + Continuous Delivery Qualifications What We Look For: + 6+ years of experience in software and/or infrastructure engineering + Bachelor's degree or equivalent experience + Deep, hands-on production experience operating Kubernetes and AWS at scale, across multiple accounts and regions + Real operational depth in at least one system we own beyond the cluster - most importantly the observability stack (Prometheus/Alertmanager at scale), but AWS networking, Vault, or artifact/CI infrastructure also count. We are filtering for people who have run these systems, not just used them. + Strong AWS networking judgment (VPC, peering, Transit Gateway, subnet/CIDR design) + A track record as a critical reviewer - spotting subtle infrastructure issues and long-term risks before they ship + Experience leading technical work and mentoring engineers; can manage, clarify, and plan around uncertainty + Effective communication and the ability to influence design in a product-focused way Nice to Have: + Prometheus long-term storage / sharding (Thanos, Cortex, Mimir, or equivalent) run in production + Experience owning a container image / base image pipeline + Policy-as-code (Kyverno / OPA) and admission webhook design ## Description As a Senior Software Engineer on the Engineering Experience (EngExp) platform team, you help run and evolve the platform that powers CJ's production systems across multiple AWS regions. "Platform" here is broad - it is the Kubernetes clusters, but also the observability stack every squad depends on, the CI/CD and artifact infrastructure their builds run through, the AWS networking that connects them, the secrets and access systems that gate them, and the cost visibility that keeps them accountable. EngExp owns all of it, and this role touches most of it. This is not just an infrastructure role - your value is in engineering judgment. You are a leader on the team: you own the reliability and operability of the platform, act as a critical reviewer of systems and changes, set standards, and mentor other engineers. We especially want someone with real depth in the systems below - we have made platform decisions we later had to reverse because the team lacked deep expertise in a component we owned, and that depth is exactly what this role brings. Responsibilities The Systems You Work On: You'll own their reliability, be the critical reviewer of changes to them, and raise the team's depth in them: + **Observability & monitoring -** Prometheus, Alertmanager, Grafana, and OpenTelemetry across every production region. This is not dashboard-building: you own cardinality budgets and recording-rule design, keep a production Prometheus healthy as it outgrows a single shard (federation / sharding / long-term store strategy), and own Alertmanager HA and the blast radius of alert-routing config. Deep Prometheus and Alertmanager expertise is a core requirement. + **Kubernetes & cloud infrastructure -** multi-region EKS clusters: upgrades, node group and Karpenter management, controller lifecycle, and add-on / configuration management. Identify failure modes before they happen (subnet IP exhaustion, API server latency, ArgoCD reconciliation lag, Prometheus cardinality, Karpenter consolidation disruption). + **AWS networking -** VPC design, subnet allocation and CIDR management, VPC peering, Transit Gateway, security groups, Route53, and NAT gateway topology across multiple accounts and regions, plus 24/7 networking alarms for all prod networking between clusters and squad resources. + **CI/CD & artifact management -** GitLab administration (runner fleet, AMI updates, cache, access - not just pipeline authoring), GitOps delivery through ArgoCD, and the Nexus artifact repository including its storage lifecycle as it grows. + **Access & identity -** Vault secrets management, IAM roles and service accounts for apps in clusters, cluster permission management for audit compliance, and AI model access management (including cost alerts and reporting). + **Cost observability -** OpenCost, EBS orphan cleanup, cost anomaly investigation, and rightsizing attribution across teams, so waste is attributable rather than shared overhead. + **Internal tools & delivery -** the container image build pipeline and base image standards, code audit tooling, HedgeDoc, and the UI CDN (S3 + CloudFront), plus adopted applications with no other owner. What You'll Do: + Own the reliability and operability of the systems above - focused on what is happening and why + Establish and enforce platform standards: RBAC, admission webhooks, resource limits, LimitRanges, policy-as-code + Manage infrastructure-as-code with Terraform across AWS accounts + Act as a high-quality reviewer of infrastructure changes - Terraform, Kubernetes configs, CI/CD pipelines, observability config - catching subtle issues and long-term risks before they ship + Turn recurring requests (ingress, DNS, service accounts) into self-service workflows that are hard to misuse + Drive resolution of platform incidents with a focus on learning and lasting system improvement + Evaluate new patterns (Gateway API / Kgateway, claim-based self-service) on tradeoffs, not hype + Mentor less-senior engineers and raise the team's depth in the components we own ## Related Videos - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [WeAreDevelopers LIVE - Modern DevOps for IoT Devices and More](https://www.wearedevelopers.com/videos/1805-wearedevelopers-live-modern-devops-for-iot-devices-and-more) - [Building a Cloud Platform Where Everything is Just Another Kubernetes Resource](https://www.wearedevelopers.com/videos/100137-building-a-cloud-platform-where-everything-is-just-another-kubernetes-resource) - [Enabling automated 1-click customer deployments with built-in quality and security](https://www.wearedevelopers.com/videos/83-enabling-automated-1-click-customer-deployments-with-built-in-quality-and-security) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) - [Our GitOps approach for deploying an Identity Provider and an API Gateway in a SaaS company](https://www.wearedevelopers.com/videos/776-our-gitops-approach-for-deploying-an-identity-provider-and-an-api-gateway-in-a-saas-company) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [The Best X (Twitter) Accounts for Developers](https://www.wearedevelopers.com/magazine/294-the-best-x-twitter-accounts-for-developers) - [The Best Software Developer Blogs to Read](https://www.wearedevelopers.com/magazine/156-the-best-software-developer-blogs-to-read) - [Dev Digest 131 - AI'm not sure about OSS](https://www.wearedevelopers.com/magazine/472-dev-digest-131-ai-m-not-sure-about-oss)