Site Reliability Engineer

Cogent Inc
Columbia, MD, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Application Performance Management Program Optimization Information Systems DevOps Distributed Systems Monitoring of Systems Performance Tuning Reliability Engineering Prometheus Datadog Data Logging
+8 more
Cloud Platform System Grafana Reliability of Systems Infrastructure Automation Frameworks Information Technology Deployment Automation Terraform Splunk

Job description

This role is responsible for implementing observability and automation practices, supporting production systems, and ensuring system performance and availability. The position plays a key role in incident response, root cause analysis, and ongoing system optimization in collaboration with DevOps and development teams., * Support system reliability, monitoring, and operational stability across environments

  • Implement and maintain observability practices, including monitoring, logging, and alerting
  • Contribute to automation efforts that improve system reliability and operational efficiency

Incident Response & Performance Optimization

  • Participate in incident response activities and production support
  • Perform root cause analysis for system issues and outages
  • Support performance optimization and tuning of applications and infrastructure

DevOps & Collaboration

  • Work with DevOps and development teams to maintain production readiness
  • Contribute to continuous improvement of deployment and operational processes
  • Collaborate across engineering teams to support stable and scalable systems

Requirements

Do you have a Bachelor’s degree?, To comply with government contracting requirements, candidates must meet all of the following:

  • Must be a U.S. Citizen, Permanent Resident, or valid EAD holder
  • Must have lived in the United States for at least 3 of the past 5 years
  • Must be currently authorized to work in the U.S. without sponsorship, The ideal candidate will bring experience in system monitoring, DevOps practices, and production support, along with the ability to collaborate across cross-functional engineering teams in a fast-paced environment., * Bachelor’s degree in Computer Science, Information Systems, or a related field, or an equivalent combination of education and experience
  • Experience in system reliability, DevOps, or production support roles
  • Experience with monitoring, logging, and observability tools
  • Understanding of incident management and root cause analysis processes
  • Familiarity with cloud environments and infrastructure concepts
  • Experience supporting automated deployment or operational workflows
  • Strong problem-solving and troubleshooting skills
  • Excellent written and verbal communication skills
  • Ability to work effectively in fast-paced, production-critical environments
  • Strong collaboration skills across development and operations teams

What Will Set You Apart

  • Experience with AWS or other cloud platforms
  • Familiarity with infrastructure-as-code tools (e.g., Terraform or similar)
  • Experience with tools such as Splunk, Datadog, Prometheus, or similar observability platforms
  • Experience with CI/CD pipelines and DevOps automation tools
  • Prior experience supporting enterprise-scale or regulated environments
  • Knowledge of application performance tuning and distributed systems behavior

Benefits & conditions

Pulled from the full job description

  • Health insurance
  • 401(k) matching
  • Paid time off
  • Vision insurance
  • Dental insurance
  • Life insurance
  • Employee assistance program, * Competitive compensation
  • Career growth and professional development opportunities
  • Exposure to complex, mission-critical systems
  • A collaborative and supportive team environment
  • Long-term client engagements with stability and continuity

We are a Certified Great Place to Work, committed to building an inclusive and high-performance culture., * Medical, Dental, and Vision Insurance (comprehensive coverage)

  • 401(k) with company match
  • Company-paid life insurance
  • Short-term and long-term disability coverage
  • Paid Time Off: 3 weeks annually + 10 paid holidays
  • Employee assistance and wellness resources (as applicable)

Compliance Notice

Cogent People Inc. conducts employment verification for all candidates. Misrepresentation of work authorization, residency history, or professional experience will result in disqualification.

We are an Equal Opportunity Employer (EEO) and evaluate all applicants based on qualifications, experience, and role requirements.

We do not engage third-party recruiters for this role unless explicitly stated.

About the company

Cogent People Inc. is a government consulting and technology services firm supporting mission-critical federal and commercial programs. We deliver secure, scalable, and modern digital solutions across complex IT environments.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · WWC Europe 2026

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · WWC 2025

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

Videos

See all

Related articles

See all