Software Engineer SRE

ARROWCORE GROUP
Austin, TX, United States
2 months ago
Apply on dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
8 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Intelligent Platform Management Interface Computer Engineering Data Centers Relational Databases Data Center Infrastructure Management (CIM) Noise Reduction Python (Programming Language) PostgreSQL MySQL Reliability Engineering Prometheus
+8 more
SQL Databases Data Logging Grafana Infrastructure Automation Frameworks Information Technology Bare Metal Restful APIs Splunk

Job description

We are seeking an experienced Sr. Data Center Site Reliability Engineer to automate operations and maximize the uptime, efficiency, and scalability of data center, facility power/cooling infrastructure, and software automation. In this role, you will manage, monitor, and optimizing both server reliability and the critical power and cooling infrastructure that sustains our distributed production systems., Enhance data center observability, logging, and alerting solutions using Grafana, Splunk, and Prometheus, building dashboards that correlate server health, network telemetry, facility power and cooling performance.

Develop automation scripts for hardware incident triage, alert noise reduction, log correlation, and operational workflows, converting recurring manual bare-power/cooling infrastructure investigation patterns into reusable tooling.

Maintain our NetBox data center inventory, building automated pipelines via APIs to track physical infrastructure, rack layouts, and cable topologies.

Build and tune Grafana dashboards with complex queries spanning multiple data sources (including Prometheus metrics) for server health visualization, bare-metal hardware bottleneck identification, and data center capacity monitoring using power feed and cooling infrastructure metrics.

Utilize Splunk and relational databases for infrastructure analytics, writing extensive SQL queries and SPL queries to troubleshoot server production issues, identify infrastructure bottlenecks, and surface environmental insights via IPMI interfaces into dashboards.

Lead incident response and on-call rotations for high-severity data center infrastructure events, directing triage, root cause analysis, mitigation, and resolution for both server-level and facility-level power feed or environmental anomalies.

Develop and maintain runbooks, hardware operational playbooks, and process documentation for common facility, power feed, and server failure scenarios, standardizing infrastructure SOPs across the SRE organization.

Collaborate closely with development, hardware engineering, and facility operations teams to integrate observability best practices into the infrastructure lifecycle and embed monitoring into new compute, storage, power and cooling system rollouts.

Requirements

Experience: 8+ years of experience in site reliability engineering, production operations, or data center infrastructure management operations.

Education: Bachelor’s Degree in Computer Science, Computer Engineering, or a related technical field is highly preferred.

Bare-Metal & Hardware Expertise: Deep hands-on experience troubleshooting, provisioning, and managing enterprise power, bare-metal hardware and server architectures.

Inventory & Asset Management: Strong proficiency using NetBox (or similar DCIM tools) for managing rack space, device lifecycle, and asset tracking.

Data & API Capabilities: Extensive SQL experience (e.g., PostgreSQL, MySQL) for querying relational data infrastructure and deep familiarity consuming/building RESTful APIs to integrate infrastructure tools.

Monitoring & Tooling: Strong expertise with Prometheus for metrics collection, Grafana for visualization, and Splunk for enterprise logging.

Infrastructure Protocols: Proficient with IPMI and server out-of-band management protocols, alongside a strong understanding of data center PDU management and power feed architecture.

Facilities Knowledge: Practical understanding of data center physical infrastructure, specifically power feed distribution systems and cooling infrastructure (e.g., HVAC, liquid cooling, hot/cold aisle containment, air handling units).

Automation: Strong scripting capabilities (Python, Shell) and experience managing infrastructure across highly distributed on-premise environments.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

2:18 min

Scaling MySQL databases for massive user growth

Johannes Nicolai Johannes Nicolai +1 · LIVE

1:04 min

Visualizing Keycloak performance via standard Grafana troubleshooting dashboards

Alexander Schwartz Alexander Schwartz · World Congress 2025

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

Videos

See all

Related articles

See all