> Markdown version of [/jobs/ext/3054237-application-sre-devops](https://www.wearedevelopers.com/jobs/ext/3054237-application-sre-devops). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Application SRE (DevOps) - **Company:** ELLKAY, LLC. - **Location:** Elmwood Park, NJ, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Microsoft Windows, Amazon Web Services, Microsoft Azure, Bash Shell, Cloud Computing, Continuous Integration, DevOps, Fault Tolerance, Github, Monitoring of Systems, Python (Programming Language), Performance Tuning, Reliability Engineering, Ansible, Prometheus, Datadog, Scripting, Grafana, Cloudformation, Containerization, Infrastructure Automation Frameworks, Puppet, Terraform, Docker, Jenkins, Microservices - **Published:** September 24, 2026 - **Apply:** https://www.juju.com/job/15_d77d44e9 ## About the Role * Strong experience as an SRE, DevOps Engineer, or Production Support Engineer * Solid understanding of Windows, Linux/Unix systems and networking fundamentals * 7 years of experience as an SRE * Hands-on experience with cloud platforms such as AWS, Azure, or GCP * Experience with containerization and orchestration tools like Docker and Kubernetes * Proficiency in CI/CD tools such as Jenkins, GitHub Actions, , or similar * Experience with Infrastructure as Code tools like Terraform, CloudFormation, or ARM * Strong scripting skills in Python, Bash, or similar languages * Experience with monitoring and observability tools (Prometheus, Grafana, ELK, Datadog, etc.) * Understanding of reliability concepts such as SLAs, SLOs, and incident management Preferred Qualifications * Experience supporting microservices-based architectures * Knowledge of security best practices in cloud and DevOps environments * Experience with configuration management tools (Ansible, Chef, or Puppet) * Exposure to chaos engineering or resilience testing practices Soft Skills * Strong problem-solving and troubleshooting skills * Ability to work calmly during incidents and high-pressure situations * Clear communication and collaboration with cross-functional teams * Ownership mindset with a focus on continuous improvement ## Description We are looking for an Application Site Reliability Engineer (SRE) with strong DevOps experience to improve the reliability, scalability, and performance of our applications. The Application Site Reliability Engineer will serve as a technical contact responsible for driving the reliability, performance, and operational maturity of our application ecosystem. This role works across multiple teams to support scalable systems, establish reliability standards, improve observability, and implement automation that reduces operational effort. The SRE will lead complex incident responses, work with engineering teams in best practices, and influence architectural decisions to ensure resilient, high-quality software delivery. You will help define reliability standards, reduce operational toil, and ensure smooth production operations while enabling faster and safer releases., * Own application reliability, availability, performance, and scalability in production and non-production environments * Design, build, and maintain CI/CD pipelines for application deployments * Automate infrastructure provisioning and configuration using Infrastructure as Code * Monitor application health using metrics, logs, and traces; define SLIs, SLOs, and error budgets * Lead incident response, root-cause analysis (RCA), ensuring corrective and preventive actions are completed and communicated. * Improve system resilience through capacity planning, system tuning, and fault tolerance * Partner with development teams to ensure services meet reliability, performance, and scalability objectives. * Reduce manual operational effort through automation and self-healing solutions * Serve as a point of contact for critical Sev1/Sev2 incidents, leading incident command when required. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Automate everything via NodeJS and Puppeteer](https://www.wearedevelopers.com/videos/322-automate-everything-via-nodejs-and-puppeteer) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [The Memory Leak That Ate Our Cluster: A Postmortem](https://www.wearedevelopers.com/videos/2057-the-memory-leak-that-ate-our-cluster-a-postmortem) - [DevOps Maturity Check – a way to balance autonomy and alignment](https://www.wearedevelopers.com/videos/58-devops-maturity-check-a-way-to-balance-autonomy-and-alignment) ## Related Articles - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs)