> Markdown version of [/jobs/ext/511463-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/511463-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Cogent Inc - **Location:** Columbia, MD, United States - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Application Performance Management, Program Optimization, Information Systems, DevOps, Distributed Systems, Monitoring of Systems, Performance Tuning, Reliability Engineering, Prometheus, Datadog, Data Logging, Cloud Platform System, Grafana, Reliability of Systems, Infrastructure Automation Frameworks, Information Technology, Deployment Automation, Terraform, Splunk - **Published:** June 10, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=b3fb601adf82700d ## About the Role Do you have a Bachelor's degree?, To comply with government contracting requirements, candidates must meet all of the following: * Must be a U.S. Citizen, Permanent Resident, or valid EAD holder * Must have lived in the United States for at least 3 of the past 5 years * Must be currently authorized to work in the U.S. without sponsorship, The ideal candidate will bring experience in system monitoring, DevOps practices, and production support, along with the ability to collaborate across cross-functional engineering teams in a fast-paced environment., * Bachelor's degree in Computer Science, Information Systems, or a related field, or an equivalent combination of education and experience * Experience in system reliability, DevOps, or production support roles * Experience with monitoring, logging, and observability tools * Understanding of incident management and root cause analysis processes * Familiarity with cloud environments and infrastructure concepts * Experience supporting automated deployment or operational workflows * Strong problem-solving and troubleshooting skills * Excellent written and verbal communication skills * Ability to work effectively in fast-paced, production-critical environments * Strong collaboration skills across development and operations teams What Will Set You Apart * Experience with AWS or other cloud platforms * Familiarity with infrastructure-as-code tools (e.g., Terraform or similar) * Experience with tools such as Splunk, Datadog, Prometheus, or similar observability platforms * Experience with CI/CD pipelines and DevOps automation tools * Prior experience supporting enterprise-scale or regulated environments * Knowledge of application performance tuning and distributed systems behavior ## Description This role is responsible for implementing observability and automation practices, supporting production systems, and ensuring system performance and availability. The position plays a key role in incident response, root cause analysis, and ongoing system optimization in collaboration with DevOps and development teams., * Support system reliability, monitoring, and operational stability across environments * Implement and maintain observability practices, including monitoring, logging, and alerting * Contribute to automation efforts that improve system reliability and operational efficiency Incident Response & Performance Optimization * Participate in incident response activities and production support * Perform root cause analysis for system issues and outages * Support performance optimization and tuning of applications and infrastructure DevOps & Collaboration * Work with DevOps and development teams to maintain production readiness * Contribute to continuous improvement of deployment and operational processes * Collaborate across engineering teams to support stable and scalable systems ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [The Best Job Search Websites of 2025](https://www.wearedevelopers.com/magazine/368-the-best-job-search-websites-of-2025) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)