> Markdown version of [/jobs/ext/2710333-cloud-ops-engineer](https://www.wearedevelopers.com/jobs/ext/2710333-cloud-ops-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Cloud Ops Engineer - **Company:** Hypori, Inc. - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $180,000.0 - $195,000.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Systems Engineering, Software as a Service, Cloud Computing, Cloud Computing Security, Cloud Engineering, Continuous Integration, Software Design Patterns, DevOps, Disaster Recovery, Github, Identity and Access Management, Uptime, Reliability Engineering, Prometheus, Datadog, Data Logging, Delivery Pipeline, Grafana, AWS Lambda, Amazon Virtual Private Cloud (VPC), Cloudformation, SC Clearance, Git Flow, Kubernetes, Information Technology, Deployment Automation, Route53, Cloudwatch, Terraform, Splunk, Docker, Jenkins - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/senior-engineer-cloud-operations-hypori-8730024 ## About the Role * Bachelor's in Computer Science, Engineering, or equivalent hands-on experience in cloud infrastructure, SRE, or systems engineering. * 8+ years in Cloud Ops, SRE, DevOps, or Infrastructure Engineering, with a demonstrated track record of progression into senior/architectural ownership. * Deep AWS expertise (EC2, VPC, IAM, S3, EBS, EFS, Route 53, CloudWatch); AWS GovCloud experience strongly preferred. * Proven experience designing multi-account AWS environments and landing zone frameworks (e.g., AWS LZA), including isolation, governance, and guardrail design. * Expert-level Infrastructure as Code (Terraform and/or CloudFormation), with a track record of building reusable modules and org-wide standards, not just consuming them. * Production-grade Kubernetes and Docker experience, with the judgment to make orchestration and architectural trade-off decisions, not just operate what exists. * CI/CD architecture experience (GitHub Actions, Jenkins), with fluency in modern deployment strategies (blue/green, canary, rolling) as design patterns, not checkboxes. * Track record operating high-availability SaaS platforms under strict security, reliability, and uptime SLAs. * Advanced observability and incident management experience (Datadog, Splunk, Prometheus, Grafana, ELK), including designing monitoring/alerting strategy for 24/7 NOC environments. * Demonstrated incident command experience - leading high-severity incidents from detection through resolution and systemic RCA. * Experience operating in regulated environments (FedRAMP, DoD IL4/IL5, SOC 2 Type 2), with working knowledge of NIST 800-53. * Strong grounding in cloud security, compliance, and resilient architecture design. * Demonstrated ability to mentor senior and mid-level engineers and set technical direction beyond your own individual work. * Fluency in Git-based workflows and modern development/deployment practices. * Preferred certifications: AWS Solutions Architect Associate (or higher), AWS CloudFormation, AWS Lambda, AWS CodePipeline. * U.S. citizenship required. * Active Secret clearance required; Top Secret preferred. * Common Access Card (CAC) eligible under DoD 8140. Personal Qualifications * A strategic, hands-on operator who can independently own ambiguous, high-stakes infrastructure problems and turn them into scalable, durable solutions. * Sharp architectural judgment - able to weigh trade-offs, make sound calls under uncertainty, and defend those decisions to both engineers and leadership. * Calm, decisive, and clear-headed under pressure; a natural incident commander others look to during high-severity events. * A strong communicator who moves easily between engineering, security, and executive conversations, and who elevates the people around them. * Deep sense of ownership and accountability - treats reliability, automation, and operational excellence as personal standards, not job requirements. * Genuinely energized by hard infrastructure problems and staying ahead of the cloud landscape. * Willing to serve as escalation resource in a 24/7 on-call rotation. ## Description You'll independently own ambiguous, high-impact infrastructure problems, guide technical direction for the team, and acting as a senior escalation point when production systems are on the line. This is a builder-designer hybrid role - someone equally comfortable writing Terraform and designing infrastructure modernization solutions. You will drive the strategy and execution of Infrastructure as Code, observability, CI/CD, and operational frameworks, while mentoring engineers and raising the technical bar across the org. This role sits at the intersection of DevOps, SRE, and cloud architecture, with a strong emphasis on engineering rigor, operational maturity, and continuous improvement in a mission-critical, regulated environment. Responsibilities * Own the architecture and operation of secure, scalable, highly available AWS and AWS GovCloud infrastructure supporting the Hypori SaaS platform end to end. * Set the technical direction for Infrastructure as Code (Terraform, OpenTofu, CloudFormation), establishing org-wide standards, reusable modules, and governance patterns other engineers build on. * Drive engineering tasks for platform reliability, scalability, and capacity; helping define SLA/SLO targets. * Lead the end-to-end observability strategy - monitoring, logging, alerting, telemetry - across Datadog, Splunk, Prometheus, Grafana, and/or ELK. * Serve as incident commander for high-severity production events; drive root cause analysis and lead systemic, multi-phase remediation - not just fixes, but prevention. * Design and implement CI/CD platforms and deployment automation (GitHub Actions, Jenkins, AWS Code Pipeline), championing progressive delivery patterns as the org standard. * Lead the operation and evolution of containerized workloads on Kubernetes and Docker, making the orchestration and scaling trade-off calls for the platform. * Proactively identify architectural scalability and performance bottlenecks before they become incidents, and design the systems that eliminate them. * Partner directly with Security and Compliance leadership to implement secure-by-design infrastructure supporting FedRAMP, DoD IL5, and SOC 2 Type 2 requirements. * Participate in owning disaster recovery and business continuity - design, test, and continuously validate RTO/RPO targets. * Identify and develop solutions for cloud cost optimization as a discipline, not a cleanup task - identifying inefficiencies at the design level and implementing measurable, sustained cost controls. * Mentor engineers, lead design reviews, and shape the technical standards the broader team is measured against. * Lead cross-team resolution of systemic operational issues, delivering road mapped improvements with measurable, reported impact. * Represent Hypori in technical conversations with external cloud and infrastructure vendors, holding them to performance, reliability, and security commitments. * Participate in a 24/7 on-call rotation ## Related Videos - [DevOps at Netflix](https://www.wearedevelopers.com/videos/270-devops-at-netflix) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [We adopted DevOps and are Cloud-native, Now What?](https://www.wearedevelopers.com/videos/485-we-adopted-devops-and-are-cloud-native-now-what) - [Leading with Reliability: Applying SRE Principles to Build Stronger Engineering Organizations](https://www.wearedevelopers.com/videos/100185-leading-with-reliability-applying-sre-principles-to-build-stronger-engineering-organizations) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers)