> Markdown version of [/jobs/ext/1135501-software-engineer-sre](https://www.wearedevelopers.com/jobs/ext/1135501-software-engineer-sre). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer SRE - **Company:** ARROWCORE GROUP - **Location:** Austin, TX, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Intelligent Platform Management Interface, Computer Engineering, Data Centers, Relational Databases, Data Center Infrastructure Management (CIM), Noise Reduction, Python (Programming Language), PostgreSQL, MySQL, Reliability Engineering, Prometheus, SQL Databases, Data Logging, Grafana, Infrastructure Automation Frameworks, Information Technology, Bare Metal, Restful APIs, Splunk - **Published:** June 30, 2026 - **Apply:** https://www.dice.com/job-detail/c4f12f35-ec38-444c-a048-b0eaf10a5802 ## About the Role Experience: 8+ years of experience in site reliability engineering, production operations, or data center infrastructure management operations. Education: Bachelor's Degree in Computer Science, Computer Engineering, or a related technical field is highly preferred. Bare-Metal & Hardware Expertise: Deep hands-on experience troubleshooting, provisioning, and managing enterprise power, bare-metal hardware and server architectures. Inventory & Asset Management: Strong proficiency using NetBox (or similar DCIM tools) for managing rack space, device lifecycle, and asset tracking. Data & API Capabilities: Extensive SQL experience (e.g., PostgreSQL, MySQL) for querying relational data infrastructure and deep familiarity consuming/building RESTful APIs to integrate infrastructure tools. Monitoring & Tooling: Strong expertise with Prometheus for metrics collection, Grafana for visualization, and Splunk for enterprise logging. Infrastructure Protocols: Proficient with IPMI and server out-of-band management protocols, alongside a strong understanding of data center PDU management and power feed architecture. Facilities Knowledge: Practical understanding of data center physical infrastructure, specifically power feed distribution systems and cooling infrastructure (e.g., HVAC, liquid cooling, hot/cold aisle containment, air handling units). Automation: Strong scripting capabilities (Python, Shell) and experience managing infrastructure across highly distributed on-premise environments. ## Description We are seeking an experienced Sr. Data Center Site Reliability Engineer to automate operations and maximize the uptime, efficiency, and scalability of data center, facility power/cooling infrastructure, and software automation. In this role, you will manage, monitor, and optimizing both server reliability and the critical power and cooling infrastructure that sustains our distributed production systems., Enhance data center observability, logging, and alerting solutions using Grafana, Splunk, and Prometheus, building dashboards that correlate server health, network telemetry, facility power and cooling performance. Develop automation scripts for hardware incident triage, alert noise reduction, log correlation, and operational workflows, converting recurring manual bare-power/cooling infrastructure investigation patterns into reusable tooling. Maintain our NetBox data center inventory, building automated pipelines via APIs to track physical infrastructure, rack layouts, and cable topologies. Build and tune Grafana dashboards with complex queries spanning multiple data sources (including Prometheus metrics) for server health visualization, bare-metal hardware bottleneck identification, and data center capacity monitoring using power feed and cooling infrastructure metrics. Utilize Splunk and relational databases for infrastructure analytics, writing extensive SQL queries and SPL queries to troubleshoot server production issues, identify infrastructure bottlenecks, and surface environmental insights via IPMI interfaces into dashboards. Lead incident response and on-call rotations for high-severity data center infrastructure events, directing triage, root cause analysis, mitigation, and resolution for both server-level and facility-level power feed or environmental anomalies. Develop and maintain runbooks, hardware operational playbooks, and process documentation for common facility, power feed, and server failure scenarios, standardizing infrastructure SOPs across the SRE organization. Collaborate closely with development, hardware engineering, and facility operations teams to integrate observability best practices into the infrastructure lifecycle and embed monitoring into new compute, storage, power and cooling system rollouts. ## Related Videos - [SRE Methods In an Agency Environment](https://www.wearedevelopers.com/videos/348-sre-methods-in-an-agency-environment) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Our journey with Spring Boot in a microservice architecture](https://www.wearedevelopers.com/videos/511-our-journey-with-spring-boot-in-a-microservice-architecture) - [MySQL Protocol Features You Should Be Aware Of](https://www.wearedevelopers.com/videos/100267-mysql-protocol-features-you-should-be-aware-of) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [All your telemetry data from any source in one place](https://www.wearedevelopers.com/videos/57-all-your-telemetry-data-from-any-source-in-one-place) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Where To Find Software Engineering Jobs](https://www.wearedevelopers.com/magazine/396-where-to-find-software-engineering-jobs) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Find a Developer Job: 12 Best Job Sites For Developers](https://www.wearedevelopers.com/magazine/165-find-a-developer-job-12-best-job-sites-for-developers)