Senior Platform/DevOps Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+37 more
Job description
We are looking for an experienced, and detail-oriented Senior Platform Engineer to join our growing Edge team. You will be responsible for the design, automation, optimization, and operation of our Kubernetes-based platform supporting our Galleon mobile data centers and Atlas cloud integration. This is a critical role where you will leverage deep technical expertise in cloud infrastructure and Kubernetes while valuing mentorship, collaboration, and open communication. You will work on building and managing resilient, secure, and scalable Kubernetes environments across diverse edge locations and cloud infrastructure, ensuring the reliability of our distributed computing platform., * Architect, design, deploy, configure, and manage highly available Kubernetes clusters on-prem (Galleon data centers) and cloud (AWS, Azure, Google Cloud Platform) environments. This includes designing the cluster layout, resource allocation, and storage configurations
- Administer, maintain, and monitor the health, performance, and capacity of Kubernetes clusters and underlying infrastructure
- Implement and manage Kubernetes networking solutions (CNI plugins, Ingress controllers) and storage solutions (PV/PVC, Storage Classes, CSI drivers)
- Design, deploy, configure, and manage Microsoft Azure Stack Local and HCI environments
- Maintain and monitor containerized platform services running within the clusters and robust monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, ELK stack)
- Drive Infrastructure-as-Code (IaC) initiatives using tools like Terraform, Ansible, Helm, and potentially Kubernetes Operators, promoting automation, repeatability, and reliability
- Support and troubleshoot complex issues related to the Kubernetes platform, containerized services, networking, and infrastructure
- Implement and enforce Kubernetes security best practices (RBAC, Network Policies, Secrets Management, Security Contexts, Image Scanning)
- Automate cluster operations, deployment pipelines (CI/CD integration), and infrastructure provisioning using Infrastructure as Code (IaC) tools (e.g., Terraform, Ansible)
- Optimize Kubernetes clusters for performance, scalability, and resource utilization, particularly in edge environments
- Develop and maintain comprehensive documentation for cluster architecture, configurations, operational procedures, and runbooks
- Work in collaboration with software engineering, DevOps, security teams, and product managers to ensure seamless integration, deployment, and secure operation of applications on Kubernetes
- Evaluate and integrate new technologies from the Kubernetes ecosystem
- Contribute to the operational excellence of the platform, including participating in on-call rotations, incident management, and building self-healing capabilities
Requirements
- ship
- At least 7+ years of experience in DevSecOps/SRE and platform engineering, with a significant focus on building and managing complex production environments
- Minimum of 5 years of hands-on experience designing, deploying, and administering production Kubernetes clusters, with experience specifically in on-premises and bare-metal deployments
- Deep expertise in Linux administration and troubleshooting, demonstrated through at least 3+ years of hands-on experience managing complex Linux environments
- Deep understanding of Kubernetes architecture, core components, operational best practices, and lifecycle management
- Strong understanding and proven experience with Infrastructure as Code (IaC) solutions, particularly Terraform and/or Ansible
- Proficiency in scripting languages (e.g., Python, Bash) for automation
- Experience configuring and managing monitoring/logging tools (e.g., Prometheus, Grafana, ELK Stack)
- Solid understanding of Linux operating system, networking fundamentals (TCP/IP, DNS, Load Balancing, Firewalls, VPNs) and container networking (CNI)
- Strong understanding of Kubernetes security concepts and implementation (RBAC, Network Policies, Secrets)
- Ability to work independently and collaborate effectively with others to debug and solve problems
- A bachelorâs degree in computer science, Engineering, Information Technology, a related technical field, or equivalent practical experience
Preferred Qualifications
- Experience with Red Hat OpenShift Container Platform
- Experience deploying and maintaining CI/CD solutions for DevSecOps, such as GitLab CI or Jenkins
- Strong development experience using Docker, docker-compose, and/or Kubernetes
- Experience developing Ansible playbooks for process automation
- Kubernetes certifications (CKA, CKS, CKAD)
- Experience with Kubernetes operators and Custom Resource Definitions (CRDs)
- Experience with service mesh technologies like Istio or Linkerd
- Experience managing Kubernetes in edge computing or resource-constrained environments, * A go-getter with a growth mindset. Youâre intellectually curious, have strong business acumen, and actively seek opportunities to build relevant skills and knowledge
- A detail-oriented problem-solver. You can independently gather information, solve problems efficiently, and deliver results with a âget-it-doneâ attitude
- Thrive in a fast-paced environment. Youâre energized by an entrepreneurial spirit, capable of working quickly, and excited to contribute to a growing company
- A collaborative team player. You focus on business success and are motivated by team accomplishment vs personal agenda
- Highly organized and results-driven. Strong prioritization skills and a dedicated work ethic are essential for you
Benefits & conditions
For U.S. Based candidates: To ensure fairness and transparency, the starting base salary range for this role for candidates in the U.S. are listed below, varying based on location experience, skills, and qualifications.
In addition to base salary, this role will also be offered equity and subsidized benefits (details available upon request)., * Competitive base salary and equity
- Medical, dental, and vision (subsidized cost)
- Health savings accounts (HSA), flexible spending accounts (FSA), and dependent care FSAs (DCFSA)
- Retirement plan options, including 401(k) and Roth 401(k)
- Unlimited paid time off (PTO)
- 14 paid company holidays per year, $157,596-$196,995 USD
About the company
Armada is the hyperscaler for the edge, delivering modular AI infrastructure from first deployment to AI factory with speed, scale and sovereignty. Named one of Fast Companyâs Most Innovative Companies and to the CNBC Disruptor 50, Armadaâs solutions are deployed in over 60 countries globally for organizations ranging from energy to defense.
With nearly $500 million in funding to date, Armada is backed by leading investors including Founders Fund, Lux, BlackRock and Microsoft (M12), alongside strategic partnerships with Microsoft, Dell, Palantir, NVIDIA, SpaceX, and Skydio. We are building the infrastructure layer for sovereign and edge AI - rugged, deployable compute for customers that cannot rely on centralized cloud.
Working at Armada means taking ownership, driving autonomy, and delivering impact. Youâll tackle challenges that havenât been solved before and help build something transformative from the ground up. What you do here will not only define your career but help further Armadaâs mission to bridge the digital divide for customers around the world.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.dice.comGood distractions
Talks and stories from around this role â technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
DevOps Engineer Salary [2023]
Is Software Engineering Over-Saturated?
Where To Find Software Engineering Jobs
Highest Paying Tech Companies for Developers