Cloud Ops Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+23 more
Job description
You’ll independently own ambiguous, high-impact infrastructure problems, guide technical direction for the team, and acting as a senior escalation point when production systems are on the line. This is a builder-designer hybrid role - someone equally comfortable writing Terraform and designing infrastructure modernization solutions.
You will drive the strategy and execution of Infrastructure as Code, observability, CI/CD, and operational frameworks, while mentoring engineers and raising the technical bar across the org. This role sits at the intersection of DevOps, SRE, and cloud architecture, with a strong emphasis on engineering rigor, operational maturity, and continuous improvement in a mission-critical, regulated environment.
Responsibilities
-
Own the architecture and operation of secure, scalable, highly available AWS and AWS GovCloud infrastructure supporting the Hypori SaaS platform end to end.
-
Set the technical direction for Infrastructure as Code (Terraform, OpenTofu, CloudFormation), establishing org-wide standards, reusable modules, and governance patterns other engineers build on.
-
Drive engineering tasks for platform reliability, scalability, and capacity; helping define SLA/SLO targets.
-
Lead the end-to-end observability strategy - monitoring, logging, alerting, telemetry - across Datadog, Splunk, Prometheus, Grafana, and/or ELK.
-
Serve as incident commander for high-severity production events; drive root cause analysis and lead systemic, multi-phase remediation - not just fixes, but prevention.
-
Design and implement CI/CD platforms and deployment automation (GitHub Actions, Jenkins, AWS Code Pipeline), championing progressive delivery patterns as the org standard.
-
Lead the operation and evolution of containerized workloads on Kubernetes and Docker, making the orchestration and scaling trade-off calls for the platform.
-
Proactively identify architectural scalability and performance bottlenecks before they become incidents, and design the systems that eliminate them.
-
Partner directly with Security and Compliance leadership to implement secure-by-design infrastructure supporting FedRAMP, DoD IL5, and SOC 2 Type 2 requirements.
-
Participate in owning disaster recovery and business continuity - design, test, and continuously validate RTO/RPO targets.
-
Identify and develop solutions for cloud cost optimization as a discipline, not a cleanup task - identifying inefficiencies at the design level and implementing measurable, sustained cost controls.
-
Mentor engineers, lead design reviews, and shape the technical standards the broader team is measured against.
-
Lead cross-team resolution of systemic operational issues, delivering road mapped improvements with measurable, reported impact.
-
Represent Hypori in technical conversations with external cloud and infrastructure vendors, holding them to performance, reliability, and security commitments.
-
Participate in a 24/7 on-call rotation
Requirements
-
Bachelor’s in Computer Science, Engineering, or equivalent hands-on experience in cloud infrastructure, SRE, or systems engineering.
-
8+ years in Cloud Ops, SRE, DevOps, or Infrastructure Engineering, with a demonstrated track record of progression into senior/architectural ownership.
-
Deep AWS expertise (EC2, VPC, IAM, S3, EBS, EFS, Route 53, CloudWatch); AWS GovCloud experience strongly preferred.
-
Proven experience designing multi-account AWS environments and landing zone frameworks (e.g., AWS LZA), including isolation, governance, and guardrail design.
-
Expert-level Infrastructure as Code (Terraform and/or CloudFormation), with a track record of building reusable modules and org-wide standards, not just consuming them.
-
Production-grade Kubernetes and Docker experience, with the judgment to make orchestration and architectural trade-off decisions, not just operate what exists.
-
CI/CD architecture experience (GitHub Actions, Jenkins), with fluency in modern deployment strategies (blue/green, canary, rolling) as design patterns, not checkboxes.
-
Track record operating high-availability SaaS platforms under strict security, reliability, and uptime SLAs.
-
Advanced observability and incident management experience (Datadog, Splunk, Prometheus, Grafana, ELK), including designing monitoring/alerting strategy for 24/7 NOC environments.
-
Demonstrated incident command experience - leading high-severity incidents from detection through resolution and systemic RCA.
-
Experience operating in regulated environments (FedRAMP, DoD IL4/IL5, SOC 2 Type 2), with working knowledge of NIST 800-53.
-
Strong grounding in cloud security, compliance, and resilient architecture design.
-
Demonstrated ability to mentor senior and mid-level engineers and set technical direction beyond your own individual work.
-
Fluency in Git-based workflows and modern development/deployment practices.
-
Preferred certifications: AWS Solutions Architect Associate (or higher), AWS CloudFormation, AWS Lambda, AWS CodePipeline.
-
U.S. citizenship required.
-
Active Secret clearance required; Top Secret preferred.
-
Common Access Card (CAC) eligible under DoD 8140.
Personal Qualifications
-
A strategic, hands-on operator who can independently own ambiguous, high-stakes infrastructure problems and turn them into scalable, durable solutions.
-
Sharp architectural judgment - able to weigh trade-offs, make sound calls under uncertainty, and defend those decisions to both engineers and leadership.
-
Calm, decisive, and clear-headed under pressure; a natural incident commander others look to during high-severity events.
-
A strong communicator who moves easily between engineering, security, and executive conversations, and who elevates the people around them.
-
Deep sense of ownership and accountability - treats reliability, automation, and operational excellence as personal standards, not job requirements.
-
Genuinely energized by hard infrastructure problems and staying ahead of the cloud landscape.
-
Willing to serve as escalation resource in a 24/7 on-call rotation.
About the company
Hypori Inc. provides a generous benefits package for full-time employees that includes medical, dental, and vision insurance, parental leave, and life and disability packages. We also invest in our employees’ futures by providing a 401(k) plan with employer-matching contributions that vest starting from your first day of employment. In addition to the base compensation, Hypori also offers a performance bonus, which is primarily contingent upon company-wide performance. We are dedicated to investing in the tools and skills required to be strong, collaborative colleagues and people managers to help build and retain a strong workforce.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
What Are The Top Skills Required For Azure Developers?
Is Software Engineering Over-Saturated?
Fully Remote Software Engineer Jobs
DevOps Engineer Salary [2023]