Cloud System Administrator 2

Wyetech LLC
Annapolis Junction, MD, United States
9 days ago
Apply on www.clearancejobs.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$157,394.0 - $213,138.0
Working hours
Regular working hours

Tech stack

Apache Accumulo Amazon Web Services Amazon S3 Systems Engineering Automation of Tests Bash Shell Cloud Computing Cloud Engineering Cloud Storage Continuous Integration Dynamic Host Configuration Protocol Linux
+50 more
Disaster Recovery RAID Distributed Data Store Distributed Systems Domain Name System (DNS) Perl (Programming Language) General Parallel File Systems Apache Hadoop Hadoop Distributed File System HAProxy Image Management Python (Programming Language) Network Layer Lightweight Directory Access Protocols (LDAP) Linux System Administration Nginx Performance Tuning Reliability Engineering Ansible Prometheus Software Deployment TCP/IP Trivial File Transfer Protocols Virtual Local Area Networks Data Logging Scripting Load Balancing Cloud Platform System Apache Yarn System Availability Grafana Kubernetes Helm Charts Reliability of Systems Firewalls (Computer Science) Cloudformation Containerization Git Flow Kubernetes Infrastructure Automation Frameworks Storage Technologies Information Technology Cassandra Nessus Database Replication Puppet Terraform Splunk Docker Network Optimization Vulnerability Analysis

Job description

We are seeking a highly skilled Senior Cloud & Distributed Systems Engineer to support mission-critical, cloud-based data repositories serving thousands of Intelligence Community users. This role operates in a dynamic, high-tempo operational environment where requirements evolve rapidly in response to global events. The selected candidate will administer and engineer large-scale Hadoop and Accumulo clusters, ensure system reliability and security, and collaborate across infrastructure, networking, and security domains to maintain continuous mission availability.

This is not a traditional system administration role - it is a reliability-focused distributed systems engineering position operating at scale.

Due to federal contract requirements, United States Citizenship and position appropriate security clearance is required. (e.g. Active TS/SCI security clearance with agency appropriate polygraph).

Capabilities

  • Monitor system health, performance, and reliability across large, distributed clusters
  • Troubleshoot and resolve complex hardware, software, network, and cloud platform issues
  • Perform root cause analysis (RCA) and contribute to post-incident reviews
  • Maintain and optimize large-scale Hadoop and Accumulo environments
  • Administer and maintain distributed storage systems
  • Engineer and improve monitoring, observability, and alerting frameworks
  • Participate in architecture and engineering design discussions
  • Create and maintain automation scripts and Infrastructure-as-Code deployments
  • Patch, upgrade, and harden systems in accordance with security compliance standards
  • Administer LDAP-based user and group accounts
  • Maintain hardware inventory and asset tracking
  • Provide after-hours on-call support in a mission-driven environment
  • Interface with hardware, network, infrastructure, and security teams, * AWS Certified SysOps Administrator - Associate
  • AWS DevOps Engineer - Professional
  • Certified Kubernetes Administrator (CKA)

Operational Environment

  • Mission-critical cloud repositories supporting thousands of users
  • High-tempo, operationally responsive environment
  • Daily interaction with infrastructure, hardware, and security teams
  • Requirements shift in response to world events and mission needs
  • Emphasis on automation-first operations and continuous improvement

Requirements

  • TS/SCI with agency appropriate poly
  • Minimum 3 years experience administering large distributed systems
  • 7 years Linux systems administration experience
  • 5 years scripting experience
  • Bachelor’s degree in Engineering, Systems Engineering, Computer Science, Mathematics, or related field highly desired
  • May substitute for two (2) years of experience

Required Technical Skills

Linux Systems Administration (7+ Years)

  • Deep understanding of Linux operating systems and internals
  • User and group account management (LDAP)
  • Configuration and administration of DHCP, DNS, and TFTP
  • System patching, upgrades, and security hardening
  • Performance tuning and resource optimization

Distributed Systems & Cluster Administration (3+ Years) Experience supporting large distributed systems consisting of:

  • Multiple clusters
  • Clusters spanning at least three racks
  • Minimum of 60 nodes per site Experience with:

  • Hadoop (HDFS, YARN tuning)
  • Accumulo (tablet balancing, performance optimization)
  • Cassandra, Scality, Swift, Gluster, Lustre, GPFS, Amazon S3, or comparable technologies

Cloud & Container Technologies

  • Kubernetes orchestration services (CKA-level knowledge preferred)
  • Docker containerization and image management
  • Helm charts and cluster configuration
  • StatefulSets and persistent volume management
  • Cloud-based storage architectures

Automation & Infrastructure as Code

  • 5+ years scripting in Bash, Python, or Perl
  • Experience with configuration management tools: o Puppet o Ansible o Salt

  • Infrastructure as Code: o Terraform or CloudFormation

  • CI/CD pipeline integration and Git-based workflows

Observability & Reliability Engineering

  • Experience implementing and managing monitoring solutions such as: o Prometheus / Grafana o ELK / OpenSearch o Splunk o Cloud-native monitoring platforms

  • Design and tuning of alerting frameworks
  • Experience defining and supporting SLAs/SLOs
  • Incident response participation and documentation
  • Capacity planning and performance analysis

Networking & Infrastructure

  • Understanding of VLANs, port channel bonding, and Layer 2/Layer 3 interactions
  • TCP/IP troubleshooting
  • Load balancing (F5, HAProxy, NGINX)
  • Firewall rule management
  • Network performance analysis

Storage & High Availability

  • RAID and storage architecture knowledge
  • Object storage optimization
  • Data replication and backup strategies
  • Multi-site failover and disaster recovery (DR) planning
  • RPO/RTO considerations
  • Active/Active or Active/Passive cluster design

Security & Compliance

  • System hardening (STIG implementation preferred)
  • Vulnerability scanning tools (e.g., ACAS/Nessus)
  • RMF familiarity
  • Security logging and audit compliance
  • Experience operating in TS/SCI environments, The successful candidate:
  • Thinks like a reliability engineer, not just a system administrator
  • Automates repetitive processes and improves operational maturity
  • Remains calm and analytical during high-impact incidents
  • Understands distributed systems behavior at scale
  • Communicates effectively across technical domains
  • Thrives in mission-driven, dynamic environments

Benefits & conditions

The Benefits Package

  • Wyetech believes in generously supporting employees as they prepare for retirement. The company automatically contributes 20% of each employee’s gross compensation to a Simplified Employee Pension (SEP) IRA, with no requirement for employee matching. All contributions are fully vested from day one, ensuring immediate ownership of retirement funds.

Additional benefits include:

  • Wyetech provides a generous PTO plan of up to 200 hours annually, aligned with applicable state leave regulations. Employees have the flexibility to adjust their PTO allocation at the start of each calendar year, ensuring it meets their evolving needs.

Full-time employees have the option to participate in a variety of voluntary benefit plans including:

  • A Choice of Medical Plan Options, some with Health Savings Account (HSA)
  • Vision and Dental
  • Life and AD&D Benefits
  • Short and Long-Term Disability
  • Hospital Indemnity, Accident, and Critical Illness Insurances
  • Optional Identity Theft and Legal Protection Services

Company Environment & Perks

  • Employee Referral Bonus Eligibility up to $10,000
  • Mobility Among Wyetech-supported Contracts
  • Various contract and work locations throughout Maryland, Virginia, Colorado, Texas, Utah, Alaska, Hawaii and OCONUS
  • Various team-building events throughout the year such as: monthly lunches, summer company picnic, and an annual holiday party.
  • Employees receive two complementary branded clothing orders annually.

$75.67 - $102.47 an hour

Pay Range: $75.67 - $102.47 per hour*

Hourly pay rates listed for this position serve as a general guideline and are not a guarantee of compensation. Compensation will vary dependent upon factors including but not limited to: Government contract rates; education; relevant prior work experience, knowledge, skills, and competencies; certifications, and geographic location. *Hourly pay rates reflect the pre-benefit gross wage amounts.

About the company

At Wyetech, you’ll be at the center of an award-winning corporate culture, breaking technological barriers and solving real-world problems for our federal government customers. We are committed to hiring the best of the best, and in return, we offer a world-class, truly unique employee experience that is rare within our industry.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.clearancejobs.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:44 min

Career transition into cloud native and data management

Michael Cade · LIVE

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

7:28 min

Constructing a new Docker layer from scratch

Oliver Seitz Oliver Seitz · World Congress 2026 Europe

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

3:55 min

Demonstrating .NET installation on Debian and Azure Linux

Silvano Coriani Silvano Coriani · Europe 2026 Virtual

5:02 min

Manual port forwarding configuration using network address translation

Oliver Seitz Oliver Seitz · World Congress 2025

Videos

See all

Related articles

See all