Site Reliability Systems Engineer

MKS2 Technologies
Washington, DC, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience required
1 year minimum
Working hours
Regular working hours
Job source

Tech stack

Microsoft Word Java (Programming Language) Microsoft Excel Microsoft Windows Amazon Web Services Application Performance Management Systems Engineering Microsoft Azure Software as a Service Cloud Engineering Software Quality Cyber Security
+24 more
DevOps Distributed Systems Middleware Monitoring of Systems IBM Websphere Application Server Microsoft Office Network Administration Oracle Databases Platform as a Service (PAAS) Microsoft PowerPoint Reliability Engineering Service Virtualization Software Engineering Data Logging Enterprise Software Applications Test-Driven Development (TDD) Oracle Enterprise Manager Information Technology SolarWinds (Software) Splunk Dynatrace Service Stack Servicenow Microservices

Job description

MKS2 Technologies, LLC, an award-winning high growth small business, creates innovative and customer-centric technology solutions in the areas of Cyber Security, Instructional Design and Training, Software Engineering and IT Support Services to improve the security and well-being of our clients. Our commitment to excellence and our “Mission First” orientation has resulted in steady growth and an expanding client base across government agencies. We have employees nationwide and for the past three consecutive years were named one of the fastest growing Veteran-owned companies in the nation. Please take a moment to browse through our website and learn more about what it means to serve with MKS2.

As a System Engineer, Sr. on our team, your main role is to work with our IST/System Engineering Team (SET) to generate monitoring/observability recommendations through the analysis of monitoring related HPIs/CPIs from initial findings through detailed analysis and generation of actionable insights. Your role expectations and tasks in support of the IST/SET team will include, but not be limited to the following items:

  • Utilize your skills in enterprise-level triage and incident resolution while gaining experience in VA system infrastructure.
  • Use modern system monitoring tools to improve VA enterprise reliability and improve the quality of services provided to veterans.
  • Work with system and application owners to obtain existing design and functionality, leverage comprehension of workflow systems and applications processes within multiple system environments and work across technology and development teams to diagnose outages and recommend changes to increase reliability.
  • Use your hardware and software experience to help strengthen the systems the VA relies on. Your primary focus will be investigation, working with event management, application owners, DevOps teams, and system and network administrators to examine issues across enterprise applications and technology stacks.
  • Partner with system and application owners to understand their platform designs and how they operate across different environments. This insight will help you diagnose outages, trace workflow issues, and recommend changes that enhance stability.
  • Collaborate with developers and identity and access teams when deeper technical investigations are needed.
  • You’ll gain hands-on experience with enterprise-level triage and incident analysis, which will deepen your understanding of the VA’s infrastructure. Tools like SolarWinds, Dynatrace, and Splunk will be part of your daily workflow, giving you the visibility needed to identify reliability concerns and support improvements to the services delivered to veterans.

Requirements

Do you have experience in Technical solutions implementation?, Do you have a High school diploma or GED?, * Deep expertise (3+ years) in two or more of the following tools used for troubleshooting application logging in an enterprise environment (Dynatrace, Splunk, SolarWinds, ServiceNow Operator Workspace)

  • Extensive experience in one or more Technology Areas (Network, Windows, Desktop, Unix/Linux, AWS or Azure Cloud, WebSphere Middleware, Java/JS Development, Microsoft or Oracle Database)
  • 8+ years of experience working with key indicators for IT system operability, reliability, application performance, and code quality
  • 8+ years of experience deploying, maintaining, and troubleshooting complex applications at an enterprise scale while working with cross-functional teams
  • 1+ years of experience in service virtualization, AWS or Azure Cloud technologies, and SaaS and PaaS implementation.
  • Experience with using Microsoft Office, including Word, Excel, and PowerPoint
  • 2+ years independently leading a team to solve difficult technical challenges
  • HS diploma or GED and 20+ years of relevant professional experience or MA or MS degree in computer science, electronics engineering, or other engineering or technical discipline with 10+ years of relevant professional experience

Nice to Have:

  • Experience with test-driven development, distributed systems, microservices and cloud-native application implementation
  • Experience with the following tools: Oracle Enterprise Manager, Riverbed - Aternity, and ServiceNow VTBs
  • Possession of excellent written and verbal communication skills
  • Possession of strong critical thinking and error assessment capabilities
  • Virtual team management
  • Public Trust Clearance

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

1:01 min

Connecting frontend application performance to user retention and revenue

Dani Coll Dani Coll · WWC 2025

1:06 min

Developer experience and project variety at scale

Alexandra Petri · WWC 2023

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

Videos

See all

Related articles

See all