Principal Software Engineer-Infrastructure
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+27 more
Job description
This role is responsible for designing, building, and operating the foundational infrastructure that powers the company’s technology platform. This includes compute, storage, networking, and container infrastructure supporting enterprise applications, internal platforms, and hybrid cloud environments., This role focuses on delivering reliable, scalable, and automated infrastructure platforms across data centers and cloud environments. Infrastructure engineers operate the foundational platforms that support modern workloads, including virtualization platforms, storage systems, networking, and container infrastructure such as Kubernetes clusters., Infrastructure Engineering
- Design, deploy, and operate infrastructure platforms including compute, storage, networking, and container infrastructure
- Build and maintain scalable infrastructure across on-premises data centers and cloud environments
- Operate and support Kubernetes clusters and their underlying infrastructure
- Ensure high availability, reliability, and performance of infrastructure systems
- Support hybrid infrastructure environments and platform services that run on top of them
Automation & Infrastructure as Code
- Develop and maintain infrastructure automation using modern programming languages (Go, Python, Java)
- Implement infrastructure provisioning and configuration through infrastructure-as-code tools such as Terraform
- Standardize infrastructure deployment and lifecycle management
- Build tooling that improves operational efficiency and infrastructure reliability
Platform Integration
- Support infrastructure dependencies for container platforms and distributed systems
- Deploy, upgrade, and maintain Kubernetes clusters and related infrastructure components
- Operate infrastructure services including IaaS platforms and storage systems
- Collaborate with platform engineering teams supporting CI/CD, messaging, observability, and developer platforms
Observability & Reliability
- Implement monitoring and observability using Prometheus, Grafana, and OpenTelemetry
- Participate in incident response and root cause analysis
- Contribute to reliability improvements and operational maturity
Security & Access Management
- Implement infrastructure security best practices
- Support identity and access management and secrets management systems
- Collaborate with security teams to ensure infrastructure resilience and compliance
Requirements
Bachelor’s degree in management information systems (MIS), computer science, or related technical field; or equivalent work experience. (Typically four years of related, progressive work experience would be needed for candidates applying for this position who do not possess a bachelor’s degree.)
- Ten or more years of experience in infrastructure, platform, or site reliability engineering roles
- Strong experience working with Linux systems and distributed infrastructure environments
- Proficiency in at least one modern programming language, such as Go, Python, or Java
Core Technical Experience
Experience across several of the following areas:
- Linux systems and core infrastructure fundamentals
- Kubernetes and container orchestration platforms
- Infrastructure as code and declarative system design (e.g., Terraform, GitOps)
- Distributed systems and large-scale infrastructure environments
- Observability and monitoring using open-source tooling (e.g., Prometheus, Grafana, OpenTelemetry)
- Distributed storage systems and concepts (e.g., Ceph or similar technologies)
- Networking fundamentals in distributed and cloud-based environments
- API-driven infrastructure and automation systems
- Infrastructure security practices, including identity, access, and secrets management
- Hybrid infrastructure spanning on-premises data centers and cloud environments, Application Programming Interface (API), Automation, Automation Systems, Best Practices, Cloud Computing, Computer Science, Continuous Deployment/Delivery, Continuous Integration, Distributed Computing, Federal Laws and Regulations, Genetics, Hardware Virtualization, High Availability, High Reliability, Identity Data Management, Incident Response, Infrastructure as a Service (IaaS), Java, Legal, Linux Operating System, Machine Tool, Management of Information Systems/Technology (MIS), Military, Natural Gas, Network Operations Center, Open Source, Operational Improvement, Operational Strategy, Power Generation, Programming Languages, Python Programming/Scripting Language, Reliability Engineering, Root Cause Analysis, Scalable System Development, Software Engineering
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
Highest Paying Tech Companies for Developers
Is Software Engineering Over-Saturated?
The Most Popular IT Jobs on the Market
Top-Paying Tech Jobs (with Salaries)