> Markdown version of [/jobs/ext/663610-platform-ops-team-within-cloudops](https://www.wearedevelopers.com/jobs/ext/663610-platform-ops-team-within-cloudops). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Platform Ops team within CloudOps - **Company:** DigiCert, Inc. - **Location:** United States - **Experience:** Expert - **Salary:** $332,800.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Computing Platforms, Microsoft Azure, Bash Shell, Cloud Computing, Distributed Systems, Domain Name System (DNS), Github, Python (Programming Language), Networking Basics, Open Source Technology, Public Key Infrastructure, Role-Based Access Control, Reliability Engineering, Prometheus, Zero Trust Network Access, Software Engineering, Software Technical Review, Google Cloud, Load Balancing, Cloud Platform System, Istio, Grafana, Multi-Cloud, Kubernetes, Cloud Migration, Terraform, Splunk, Serverless Computing - **Published:** June 27, 2026 - **Apply:** https://www.dice.com/job-detail/f4a510c6-f1b6-4a79-bea5-38035ac03d13 ## About the Role * 5+ years of experience in SRE, platform engineering, or infrastructure engineering roles * Deep proficiency in at least one major cloud provider (AWS, Google Cloud Platform, or Azure) with working knowledge of multi-cloud environments * Strong software engineering skills in Python, Go, or Bash; comfortable writing production-grade automation and tooling * Hands-on Kubernetes experience: cluster operations, workload management, networking (CNI/service mesh), and security (RBAC, pod security) * Infrastructure-as-code expertise with Terraform or equivalent; experience with GitOps workflows * Proven experience designing and operating observability systems and responding to production incidents at scale * Strong understanding of networking fundamentals: DNS, TLS/PKI, load balancing, and zero-trust networking concepts Nice to have * Experience in PKI, certificate lifecycle management, or security-adjacent infrastructure * Familiarity with compliance frameworks such as SOC 2, FedRAMP, or ISO 27001 in cloud environments * Prior experience driving cloud migration or modernization programs at scale * Contributions to open-source infrastructure or platform projects * AWS/Google Cloud Platform/Azure professional-level certifications (e.g., AWS Solutions Architect Professional, CKA/CKS) What success looks like In your first 90 days, you'll have a deep understanding of our platform's reliability posture, contributed to at least one automation or modernization initiative, and be a trusted voice in incident response. Within a year, you'll have measurably reduced toil, improved SLO attainment across key services, and delivered at least one major platform capability that enables product teams to move faster. ## Description The Platform Ops team within CloudOps is responsible for the reliability, scalability, and modernization of DigiCert's cloud infrastructure. As a Principle SRE, you will own the intersection of software engineering and operations-driving automation-first practices, reducing toil, and accelerating our cloud transformation across AWS, Azure, and Google Cloud Platform environments. You will be a technical force multiplier: raising reliability standards across the organization, defining SLOs that matter, and building the internal platforms and tooling that enable product teams to ship with confidence. What you will do Reliability Engineering * Define, implement, and own SLIs, SLOs, and error budgets for critical platform services * Lead blameless post-mortems and drive systemic reliability improvements across the platform * Design and implement observability pipelines (metrics, logs, traces) using tools such as Splunk, Prometheus, Grafana, or OpenTelemetry * Participate in on-call rotation and serve as an incident commander for P0/P1 events Cloud Modernization * Architect and execute migration strategies from legacy infrastructure to cloud-native patterns (containers, serverless, managed services) * Champion adoption of Kubernetes, service mesh, and managed cloud services (EKS, GKE, AKS) * Evaluate and introduce emerging cloud technologies that improve availability, cost efficiency, and developer experience * Partner with architecture and security teams to embed reliability and compliance into platform design Automation & Platform Development * Build and maintain infrastructure-as-code using Terraform across multi-cloud environments * Develop internal tooling, self-service platforms, and golden-path templates that reduce operational burden for development teams * Automate operational workflows including provisioning, scaling, patching, and secret rotation * Contribute to and maintain CI/CD pipelines (GitHub Actions) to enable safe, frequent deployments Engineering Leadership * Mentor mid-level engineers on SRE principles, distributed systems, and infrastructure best practices * Collaborate cross-functionally with product, security, and compliance teams to deliver on platform roadmap commitments * Document architectural decisions, runbooks, and platform standards; raise the engineering bar through code and design reviews ## Related Videos - [Platform Engineering vs. DevOps Why not both?](https://www.wearedevelopers.com/videos/885-platform-engineering-vs-devops-why-not-both) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [From Pipelines to Platforms: How DevOps Automation Becomes a Force Multiplier at Scale](https://www.wearedevelopers.com/videos/1946-from-pipelines-to-platforms-how-devops-automation-becomes-a-force-multiplier-at-scale) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) ## Related Articles - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Learning Kubernetes made easy with KubeCampus](https://www.wearedevelopers.com/magazine/348-learning-kubernetes-made-easy-with-kubecampus) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Events like RSAC Get You CISOs. Developers Decide What Actually Gets Deployed.](https://www.wearedevelopers.com/magazine/693-events-like-rsac-get-you-cisos-developers-decide-what-actually-gets-deployed) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers)