DevOps Engineer

Argonne National Laboratory
Lemont, United States of America
5 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Intermediate
Compensation
$ 109K

Job location

Remote
Lemont, United States of America

Tech stack

API
Artificial Intelligence
Amazon Web Services (AWS)
Unix
Software as a Service
Cloud Computing
Cloud Engineering
Code Review
Computer Security
Computer Networks
Continuous Delivery
Continuous Integration
Data as a Services
DevOps
Github
Identity and Access Management
Python
Key Management
Nagios
Network Protocols
Cloud Services
Prometheus
Software Deployment
Load Balancing
Grafana
Multi-Cloud
Gitlab
Git Flow
Kubernetes
Information Technology
Azure
Machine Learning Operations
Api Gateway
Terraform
Devsecops
Workday

Job description

As a DevOps Engineer, you will work within the Data Services and Workflows team at ALCF, with significant interaction with the Infrastructure Services group of AmSC to support all activities on our multi-cloud central hub infrastructure, for development, staging, pre-production, and production environments. Other L2 science service teams are deploying services on top of the infrastructure that the Infrastructure team manages - e.g. data catalogs and repositories, at-scale HPC compute services, user interfaces and APIs, and intelligent operations (AI/MLOps). Your primary job responsibilities will be to support the science teams by building foundational infrastructure and developing CI/CD pipelines to deploy services on that infrastructure., * The service stack is primarily Kubernetes-based. Perform cluster administration and application deployment assistance to users.

  • Build and maintain pipelines for deploying cloud infrastructure and science services.
  • Manage and use image registries such as Harbor.
  • Writing and updating automation for resource provisioning and CI/CD pipelines - e.g. Terraform, GitOps, Python.
  • Implement security controls as defined by Cybersecurity team (DevSecOps).
  • Configure basic instrumentation for infrastructure and core services, to feed into monitoring and alerting systems.
  • Provide primary operational support and engineering for production applications.
  • Define and implement KPIs, processes and drive continuous improvement.
  • Diagnose platform operational problems quickly and effectively.
  • Deploy, manage, and operate managed Kubernetes clusters (Amazon EKS, Azure AKS, Google GKE, or equivalent), including node group lifecycle management, cluster upgrades, cloud-native networking integrations (load balancer controllers, CNI plugins), and multi-environment promotion across dev, staging, and production.
  • Coordinate with vendors to resolve hardware and software problems.

Hybrid Remote Work - Occasionally Onsite: which applies to employees regularly scheduled for some onsite and some remote days, with employees typically working more than 60% of their time remotely. Non-exempt employees should not be permitted to split their time between on-site and remote work on a given workday unless they have advance supervisor approval.

Requirements

  • PT2: Bachelor's Degree in computer science or closely related field and a minimum of 2+ years of experience as a DevOps engineer and/or Cloud Engineer.
  • US citizenship: To perform the essential functions of this position successful applicants must provide proof of U.S. citizenship, which is required to comply with federal regulations and contract
  • Previous experience as a team lead - able to perform task management for DevOps or Cloud Engineering teams.
  • Excellent interpersonal/communication skills, and the ability to work as part of a team.
  • Working knowledge of cloud application architecture patterns and a thorough grasp of common products and managed services for at least one Cloud Service Provider (e.g. AWS).
  • Working knowledge of Kubernetes cluster administration and concepts (CR/CRDs) and application deployment strategies (GitOps, Helm).
  • Working knowledge of Unix system fundamentals and common network protocols.
  • Solid understanding of cloud computing networking concepts.
  • Ability to proactively identify performance issues, problems, and areas for improvement.
  • Ability to identify requirements and to define, plan, and implement requisite solutions.
  • Ability to plan, organize, prioritize tasks, and complete assigned projects with minimal supervision.
  • Experience with continuous integration and continuous deployment software methodologies and strategies.
  • An understanding of code review and familiarity with tools like GitHub and GitLab.
  • Experience using tools such as Nagios, Grafana and Prometheus to monitor systems, metrics, and create dashboards.
  • Experience with OpenTofu or Terraform in multi-account AWS environments, including AWS Organizations, SCPs, and IRSA (IAM Roles for Service Accounts).
  • Hands-on experience with ArgoCD including App of Apps patterns and ApplicationSets for multi-environment GitOps deployments.
  • Familiarity with Tanka, Jsonnet, or equivalent configuration-as-code templating approaches beyond Helm.
  • Experience with Kong Gateway or similar API gateway platforms in Kubernetes environments.
  • Familiarity with secrets management patterns including AWS Secrets Manager, External Secrets Operator, or comparable Kubernetes-native solutions.
  • Ability to model Argonne's Core Values: Impact, Safety, Respect, Integrity, and Teamwork.
  • Exposure to high-speed research networks such as ESnet or Internet2 is a plus.

About the company

The Argonne Leadership Computing Facility's (ALCF) mission is to accelerate major scientific discoveries and engineering breakthroughs for humanity by designing and providing world-leading computing facilities in partnership with the computational science community. We help researchers solve some of the world's largest and most complex problems with our unique combination of supercomputing resources and computational science expertise. The ALCF seeks a DevOps engineer to support infrastructure within ALCF and in support of the Department of Energy's American Science Cloud (AmSC) project. AmSC is a secure, federated, and science-optimized cloud environment that integrates the DOE's world-leading computing and experimental facilities, data resources, and high-performance networks. The AmSC platform enables DOE scientists to create, access, and integrate AI-ready datasets, run scalable AI model training and inference on leadership-class systems, perform distributed simulations, control instruments, and move data efficiently across sites., Argonne employees, and certain guest researchers and contractors, are subject to particular restrictions related to participation in Foreign Government Sponsored or Affiliated Activities, as defined and detailed in United States Department of Energy Order 486.1A. You will be asked to disclose any such participation in the application phase for review by Argonne's Legal Department. All Argonne offers of employment are contingent upon a background check that includes an assessment of criminal conviction history conducted on an individualized and case-by-case basis. Please be advised that Argonne positions require upon hire (or may require in the future) for the individual be to obtain a government access authorization that involves additional background check requirements. Failure to obtain or maintain such government access authorization could result in the withdrawal of a job offer or future termination of employment.

Apply for this position