Senior Manager, Cloud & Infrastructure Engineering and Operations
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+21 more
Job description
Infrastructure Engineering & Automation
- Design, implement, and operate secure, resilient infrastructure across AWS and Azure.
- Personally develop and maintain infrastructure as code using Terraform or equivalent tools.
- Establish reusable infrastructure modules, version-controlled configurations, automated validation, and controlled deployment pipelines.
- Automate provisioning, configuration, maintenance, and operational tasks.
- Review infrastructure code and technical designs for security, reliability, maintainability, and cost.
Cloud Operations & Reliability
- Lead day-to-day infrastructure operations, including monitoring, alerting, patching, capacity planning, and lifecycle management.
- Troubleshoot complex issues across cloud services, networking, identity, operating systems, compute, and storage.
- Participate directly in major incident response and drive root-cause analysis and preventive improvements.
- Establish and test backup, restoration, and disaster recovery procedures.
- Maintain practical runbooks, architecture documentation, and operational standards.
- Identify opportunities to improve performance and reduce unnecessary cloud spending.
Enterprise Security
- Embed security controls into infrastructure design, code, deployment, and operations.
- Implement least-privilege access, privileged access controls, network segmentation, encryption, secrets management, and centralized logging.
- Partner with security teams to prioritize and remediate vulnerabilities and configuration risks.
- Support audit and compliance requirements through documented controls and verifiable evidence.
- Establish secure configuration baselines and controlled, traceable infrastructure changes.
Team & Technical Leadership
- Manage, coach, and develop infrastructure and cloud engineers while remaining an active technical contributor.
- Set priorities, delegate effectively, and maintain clear ownership of services and deliverables.
- Provide practical guidance through code reviews, design reviews, troubleshooting, and mentoring.
- Balance planned engineering work with operational demands and team capacity.
- Communicate technical risks, priorities, and improvement plans to technical and business stakeholders.
- Coordinate with application, security, and business teams and external service providers., * Secure, reliable infrastructure delivered through repeatable, reviewed, and automated processes.
- Fewer recurring incidents, manual tasks, and unresolved security risks.
- Tested recovery capabilities and clear operational ownership.
- A capable engineering team supported by a manager who contributes directly while delegating effectively.
Interview Expectations
Candidates should be prepared to discuss infrastructure they personally built, code they authored, security controls they implemented, and production incidents they directly resolved. We will explore both individual technical contributions and examples of effective team leadership.
Requirements
Hands-on engineering is an essential, ongoing responsibility of this role. You will personally author and review infrastructure as code, implement security controls, automate operations, and troubleshoot complex production issues while managing and developing engineers.
The successful candidate combines senior-level engineering depth, sound operational judgment, and strong people leadership. Management experience or architectural oversight alone is not sufficient; recent, direct production implementation experience is required., * Senior-level enterprise infrastructure experience: A substantial track record of progressively responsible engineering and operations work supporting business-critical environments.
- Hands-on AWS and Azure expertise: Recent experience personally implementing and operating production infrastructure in both platforms. Depth may vary by service, but working production capability in both is required.
- Infrastructure as code: Strong practical experience authoring, reviewing, deploying, and maintaining infrastructure using Terraform or equivalent tooling, including reusable modules, state management, and configuration drift management.
- Infrastructure delivery practices: Experience with Git, peer review, automated testing or validation, and CI/CD pipelines for infrastructure changes.
- Enterprise security: Direct experience implementing identity and access controls, secure networking, encryption, secrets management, logging, and vulnerability remediation.
- Networking fundamentals: Strong understanding of TCP/IP, DNS, routing, firewalls, load balancing, VPNs, and private cloud connectivity.
- Systems administration: Strong production administration and troubleshooting skills in Linux and/or Windows Server environments.
- Scripting and automation: Proficiency in Python, PowerShell, Bash, or a comparable scripting language.
- Operational reliability: Experience with observability, incident response, root-cause analysis, backup restoration, and disaster recovery testing.
- People management: Demonstrated experience managing engineers, providing performance feedback, developing technical capability, and delivering through a team.
- Technical judgment and communication: Ability to explain tradeoffs, prioritize risk, and translate business requirements into practical infrastructure solutions., * Experience with hybrid infrastructure and enterprise cloud migrations.
- Experience establishing cloud governance, landing zones, and multi-account or multi-subscription environments.
- Experience with policy as code and automated security or compliance checks.
- Experience supporting regulated or audit-intensive organizations.
- Experience with cloud cost allocation, budgeting, and optimization.
- Relevant AWS, Azure, infrastructure automation, or security certifications. Certifications complement-but do not replace-hands-on experience.
Benefits & conditions
- 401(k)
- 401(k) matching
- Dental insurance
- Health insurance
- Life insurance
- Paid time off
- Referral program
- Vision insurance
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence
Is Software Engineering Over-Saturated?
Highest Paying Tech Companies for Developers