> Markdown version of [/jobs/ext/2804967-site-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2804967-site-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Site Reliability Engineer - **Company:** Oxio Corporation - **Location:** United States (Remote available) - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Application Performance Management, Automation of Tests, Microsoft Azure, Bash Shell, Cellular Networks, Cloud Engineering, Configuration Management, System Configuration, Continuous Delivery, Continuous Integration, Information Engineering, Linux, File Systems, Distributed Systems, Domain Name System (DNS), Memory Management, Elasticsearch, Perl (Programming Language), Monitoring of Systems, Hypertext Transfer Protocols (HTTP), Identity and Access Management, Python (Programming Language), Kernel-Based Virtual Machine, Load Testing, NoSQL, Ansible, Prometheus, Ruby, Zero Trust Network Access, SQL Databases, TCP/IP, Virtualization Technology, Workflow Management Systems, Datadog, Circleci, Scripting, Google Cloud, Load Balancing, System Availability, Saltstack, Grafana, Backend, Cloudformation, Containerization, Gitlab-ci, Kubernetes, Infrastructure Automation Frameworks, Deployment Automation, Cassandra, Apache Kafka, Terraform, Splunk, Dynatrace, Docker, Elk Stack, Jenkins, Mobile Data, Vmware - **Published:** September 9, 2026 - **Apply:** https://www.builtincolorado.com/job/site-reliability-engineer/11063834?handler=ApplyRedirect ## About the Role * Understanding of Linux/Unix systems (most systems are Linux-based). * Familiarity with Linux/Unix system internals like process management, filesystems, memory management, and networking. * Proficiency in at least one programming language (Python, Go, or Ruby) and strong skills in scripting (Bash, Perl). * Experience with infrastructure provisioning tools such as Terraform, CloudFormation, or Ansible. * Familiarity with containerization (Docker) and orchestration tools (Kubernetes). * Familiarity with monitoring tools like Prometheus, Grafana, or Datadog. * Knowledge of setting up alerts, analyzing logs, and creating dashboards for observability. * Familiarity with incident management practices (e.g., runbooks, postmortems). * Experience in being part of an on-call rotation and handling incidents. * Experience in setting up and maintaining Continuous Integration/Continuous Delivery pipelines (Jenkins, GitLab CI, CircleCI, etc.). * Hands-on experience with cloud providers (AWS, Google Cloud, Azure). * Knowledge of virtualization technologies (VMware, KVM) and cloud-native architecture. * Understanding of TCP/IP, DNS, HTTP/HTTPS, load balancing, and firewalls. Nice to have * Strong understanding of deployment strategies (canary releases, blue-green deployments, etc.). * Familiarity with high availability and understanding failover mechanisms. * Familiarity with IAM (Identity and Access Management) and zero trust principles. * Experience working with distributed systems (e.g., Kafka, Cassandra, Elasticsearch). * Building custom monitoring tools or writing complex automation scripts. * Functional knowledge of database management (SQL and NoSQL). * Familiarity with distributed tracing (Jaeger, OpenTelemetry) and advanced log aggregation strategies (ELK stack, Splunk). * Familiarity with performance profiling tools and optimizing application performance under heavy load. * Familiarity in load testing and identifying bottlenecks. * Familiarity with Configuration Managment using SaltStack for maintaining server configurations. ## Description Hiring Remotely in USA Entry level Remote Hiring Remotely in USA Entry level Design and operate cloud platforms supporting backend telecom services. Automate deployments, scaling, recovery, and infrastructure provisioning; monitor production systems; maintain observability, alerting, and dashboards; support incident response and on-call operations; manage CI/CD pipelines; and enable engineering, telecom, and data teams through reliable tools and infrastructure. The summary above was generated by AI Site Reliability Engineer OXIO is the first NeoTelco. We are building the world's largest, most accessible, and insightful Telecom network. Our platform empowers anyone to spin up their own carrier from a browser, scaling and supporting you as you scale your network to millions of users. We ensure that users and devices are connected, and stay connected wherever they go: Cross- country, carrier, or cellular technology. We help them pay less for mobile data. This technology is provided through our Carrier-as-a-Service platform: BrandVNO, a fully customizable telecom service. In addition, we enable clients of our service to extract the value from telecom data - enriching their customer experience, business intelligence, and product understanding in the many markets in which we operate. Come join us in creating a modern technology platform with a group of engineers dedicated to advancing our vision. Our team is passionate about what we build, open to new ideas and challenges, and has our sights set on the future of connectivity. Responsibilities * Design and implement platform on the cloud to support OXIO backend services * Automate technical operations: deployments, scaling, recovery, etc. * Monitor and maintain mission-critical production infrastructure to ensure maximum uptime * Participate in an on-call rotation and culture of continuous improvement through blameless postmortems * Enable the Engineering/Telecom/Data Engineering teams by providing them the tools to operate the service they build ## Related Videos - [Coffee with Developers: David Heinemeier Hansson](https://www.wearedevelopers.com/videos/875-coffee-with-developers-david-heinemeier-hansson) - [Leveraging Real time data in FSIs](https://www.wearedevelopers.com/videos/806-leveraging-real-time-data-in-fsis) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [DevOps at Netflix](https://www.wearedevelopers.com/videos/270-devops-at-netflix) - [Coroutine explained yet again 60 years later](https://www.wearedevelopers.com/videos/690-coroutine-explained-yet-again-60-years-later) - [Remote Driving on Plant Grounds with State-of-the-Art Cloud Technologies](https://www.wearedevelopers.com/videos/251-remote-driving-on-plant-grounds-with-state-of-the-art-cloud-technologies) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [The 8 Best Code Testing Tools](https://www.wearedevelopers.com/magazine/402-the-8-best-code-testing-tools)