> Markdown version of [/jobs/ext/1638048-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/1638048-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Exiger LLC - **Location:** Jersey City, NJ, United States (Remote available) - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Amazon Web Services, Apache HTTP Server, Systems Engineering, Cloud Computing, Databases, Data Warehousing, Software Debugging, Linux, File Systems, Distributed Systems, Domain Name System (DNS), Monitoring of Systems, Networking Basics, Routing, Reliability Engineering, Site Reliability Engineering Practices, System Programming, Systems Integration, TCP/IP, Web Services, Load Balancing, Data Storage Technologies, Snowflake, Information Technology, Data Management, Data Pipelines - **Published:** July 9, 2026 - **Apply:** https://diversityjobs.com/main/sendform/8/8/28176/1/17521590?backUrl=%2Fcareer%2F17521590%2FSite-Reliability-Engineer-New-Jersey-Jersey-City ## About the Role * Bachelor's or Master's degree in Computer Science, a related field, or equivalent practical experience. * 6 years of experience in software or systems engineering, including at least 4 years in a dedicated Site Reliability Engineering, production engineering, or platform reliability role. As our first SRE hire, you must have practiced SRE before and be ready to establish the function. * 4 years of experience designing, analyzing, and troubleshooting large-scale distributed systems. * Strong grounding in Unix/Linux internals (filesystems, processes, system calls) and networking fundamentals (TCP/IP, DNS, routing, load balancing). * Hands-on experience establishing core SRE practices from the ground up: SLIs, SLOs, and error budgets, monitoring and observability, capacity planning, and automation that removes repetitive manual work. * A rigorous, empirical mindset: you form hypotheses, measure outcomes, and make metrics-driven decisions rather than relying on intuition or anecdote. * Experience with chaos engineering or fault-injection testing (for example game days, Chaos Monkey, Gremlin, or LitmusChaos) to validate system resilience. * Proven incident management experience: on-call ownership, leading response under pressure, and driving blameless postmortems to root cause. * Experience in troubleshooting and supporting applications like web services, data storage, databases, and data pipelines, with Linux/Unix or other operating systems. * Familiarity with cloud platforms (AWS) and secure system integration. * Comfort integrating AI coding assistants (such as Claude and Codex) into your daily engineering workflow. * Ability to translate ambiguous mission problems into structured technical solutions. * Ability to operate independently in dynamic, high-stakes environments. * Willingness to travel as needed to support customer engagements. Nice to Have: * 4 years of experience programming in Go or C (Java also welcome), with the ability to debug, optimize, and automate rather than just script. * Experience supporting ML or data platforms in production. * Familiarity with data warehouses such as Snowflake, Redshift and/or Apache Iceberg. * Experience operating in FedRAMP or other regulated or government environments. ## Description * Establish the SRE function: define SLIs, SLOs, and error budgets, and set reliability standards that other engineering teams adopt. * Build and own observability: instrument services for availability, latency, and system health, and turn that signal into actionable insight. * Drive decisions with data: form hypotheses, measure the impact of every change, and let metrics rather than intuition set reliability priorities. * Own the reliability of production services from design consulting and launch reviews through steady-state operation. * Eliminate repetitive manual operations through automation and infrastructure as code, replacing them with reliable, self-service tooling. * Plan for scale: capacity planning, performance analysis, and driving changes that improve both reliability and delivery velocity. * Improve resilience through chaos engineering and fault-injection testing, running game days that prove the platform degrades gracefully and recovers from failure. * Lead sustainable, blameless incident response and postmortems, and stand up and participate in an on-call rotation. * Leverage AI-assisted development tooling (such as Codex and Claude) to accelerate automation, tooling, and investigation work, and help the team adopt these tools effectively. ## Related Videos - [An Applied Introduction to eBPF with Go](https://www.wearedevelopers.com/videos/1075-an-applied-introduction-to-ebpf-with-go) - [Creating a routing app with Google Maps API from scratch](https://www.wearedevelopers.com/videos/831-creating-a-routing-app-with-google-maps-api-from-scratch) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [Turning Container security up to 11 with Capabilities](https://www.wearedevelopers.com/videos/718-turning-container-security-up-to-11-with-capabilities) - [A Technical Introduction to Bitcoin's 2nd Layer- The Lightning Network](https://www.wearedevelopers.com/videos/15-a-technical-introduction-to-bitcoin-s-2nd-layer-the-lightning-network) ## Related Articles - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Where to Find Entry-Level Software Engineering Jobs](https://www.wearedevelopers.com/magazine/397-where-to-find-entry-level-software-engineering-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [How Much Does a Software Engineer Make? Realistic Software Engineering Salaries](https://www.wearedevelopers.com/magazine/425-how-much-does-a-software-engineer-make-realistic-software-engineering-salaries)