> Markdown version of [/jobs/ext/734786-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/734786-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Asylon Incorporated - **Location:** Norristown, PA, United States - **Experience:** Experienced - **Salary:** $118,000.0 - $150,000.0 - **Contract:** Permanent contract - **Skills:** Automation of Tests, Bash Shell, Cloud Computing, Software Debugging, DevOps, Distributed Systems, Domain Name System (DNS), Virtual Private Networks (VPN), Python (Programming Language), Linux System Administration, Message Broker, Message Queuing Telemetry Transport (MQTT), Networking Basics, Raspberry Pi, Reliability Engineering, Ansible, Prometheus, Data Streaming, Data Logging, Grafana, Firewalls (Computer Science), Containerization, Kubernetes, Infrastructure Automation Frameworks, Apache Kafka, Build Tools, Codebase, Video Streaming, Software Coding, Terraform, Data Pipelines - **Published:** June 29, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=7170894c28f0f8cb ## About the Role Do you have experience in Tooling?, * 3+ years of professional experience in SRE, DevOps, or infrastructure engineering * Strong with a high-level language such as Python, Go, or Bash for building automation and tooling * Proficient with Kubernetes - deploying, operating, debugging, and scaling containerized workloads * Experience building and operating observability stacks - Prometheus, Grafana, Loki, or similar monitoring, alerting, and logging tools * Background in CI/CD pipelines for automated testing, building, and deploying services * Proficient with Linux systems administration and troubleshooting * Experience with infrastructure-as-code tools such as OpenTofu, Terraform, or Ansible * Comfortable with networking fundamentals - DNS, firewalls, VPNs, and debugging connectivity issues across distributed environments Bonus Points * Experience with K3s or lightweight Kubernetes on edge - running services on resource-constrained hardware in the field * Has worked in airgapped or disconnected environments where systems must operate without cloud dependencies * Experience with on-call rotations and structured incident management processes * Familiarity with message brokers and streaming (MQTT, NATS, Kafka, or similar) for real-time data pipelines * Has worked with robotics or IoT systems, particularly managing fleets of remote devices * Experience with video streaming or processing pipelines in a production environment * Comfortable getting hands-on with hardware - you don't need to be an embedded expert, but you should be the kind of person who's built robots in a college club, tinkered with a Raspberry Pi, or isn't afraid to plug into a serial console and debug a device on a bench * Experience with Bazel or similar build systems for managing complex, multi-language codebases * Experience with capacity planning and performance engineering ## Description Asylon is hiring a site reliability engineer to join our Philadelphia team. You'll be responsible for the reliability, availability, and performance of systems that span cloud infrastructure, on-prem servers in airgapped customer environments, and Kubernetes clusters running on edge devices deployed with our robots in the field. You'll define and maintain SLOs, build observability into every layer of the stack, lead incident response, and drive the automation that keeps our systems running without manual intervention. This role sits at the intersection of infrastructure engineering and operations - you should be as comfortable writing code to eliminate toil as you are triaging an outage on a remote edge device. Due to the nature of the projects worked on in this position, applicants must be a U.S. Person as defined by 22 C.F.R. §120.62. This includes U.S. Citizens, lawful permanent residents, refugees, or asylees. Primary Duties * Own the reliability of production systems across cloud, on-prem, and edge environments - define SLOs, track error budgets, and drive improvements * Build and maintain observability infrastructure - monitoring, alerting, logging, and dashboards - to provide visibility into system health at every layer * Lead incident response, conduct blameless post-mortems, and implement remediation to prevent recurrence * Develop automation to reduce toil, improve deployment reliability, and enable self-healing infrastructure * Build and maintain CI/CD pipelines for service deployment, testing, and infrastructure provisioning * Manage Kubernetes clusters (K3s on edge, on-prem, and managed cloud clusters) - deployments, upgrades, and troubleshooting * Manage infrastructure-as-code for reproducible provisioning across cloud and airgapped on-prem environments * Collaborate with software and robotics engineers to build reliability into systems from the design phase ## Related Videos - [Dev & Test in the Cloud? Deploy your cloud environments with Ansible & Terraform](https://www.wearedevelopers.com/videos/1607-dev-test-in-the-cloud-deploy-your-cloud-environments-with-ansible-terraform) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Eclipse Che for Infrastructure Automation](https://www.wearedevelopers.com/videos/1611-eclipse-che-for-infrastructure-automation) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this) - [Dev Digest 138 - Are you secure about this?](https://www.wearedevelopers.com/magazine/486-dev-digest-138-are-you-secure-about-this) - [Dev Digest 131 - AI'm not sure about OSS](https://www.wearedevelopers.com/magazine/472-dev-digest-131-ai-m-not-sure-about-oss)