Site Reliability Engineer II (SRE II) - Data & Intelligence

Noblesoft Technologies
Dallas, TX, United States
16 days ago
Apply on www.careerjet.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours

Tech stack

Java (Programming Language) JavaScript (Programming Language) Artificial Intelligence Amazon Web Services Data Analysis Systems Engineering Microsoft Azure Batch Processing C Sharp (Programming Language) Cloud Engineering Continuous Integration Data as a Services
+47 more
Information Engineering Data Governance Data Infrastructure Extract Transform Load (ETL) Data Systems DevOps Disaster Recovery Distributed Systems Domain Name System (DNS) Monitoring of Systems Python (Programming Language) NoSQL Reliability Engineering Prometheus Azure Machine Learning Azure Data Lake Software Engineering SQL Databases Data Streaming Datadog Data Logging Load Balancing Cloud Platform System Data Ingestion Cloud Monitoring System Availability Snowflake Grafana Reliability of Systems Infrastructure as Code (IaC) Containerization Data Lakes Kubernetes Infrastructure Automation Frameworks Information Technology Deployment Automation Performance Monitor Bicep Apache Kafka Data Management Machine Learning Operations Terraform Splunk Azure Synapse Analytics Data Pipelines Databricks Golang

Job description

The Site Reliability Engineer II (SRE II) is responsible for ensuring the reliability, scalability, performance, security, and operational excellence of data platforms, analytics systems, AI/ML services, and business intelligence applications. This role combines software engineering, systems engineering, automation, and operational expertise to build resilient and highly available data services while driving continuous improvement through observability, automation, and reliability engineering practices. The ideal candidate is passionate about large-scale distributed systems, cloud-native technologies, data platforms, and operational excellence. They partner closely with Data Engineering, Data Science, Analytics, Platform Engineering, and Product teams to maintain and improve critical business services., Reliability & Operations Ensure the availability, performance, scalability, and reliability of data and intelligence platforms. Manage production environments supporting data ingestion, processing, transformation, storage, analytics, and AI/ML workloads. Participate in on-call rotations and incident response activities. Lead troubleshooting efforts for complex production issues and drive root cause analysis (RCA). Develop and implement service level indicators (SLIs), service level objectives (SLOs), and error budgets. Automation & Engineering Design and develop automation to improve operational efficiency and system reliability. Build self-healing solutions and automate routine operational tasks. Create tools and scripts to monitor, deploy, and manage large-scale distributed systems. Improve deployment processes through CI/CD pipelines and Infrastructure as Code (IaC). Observability & Monitoring Design and maintain monitoring, logging, tracing, and alerting solutions. Create dashboards and actionable alerts to proactively identify service degradation. Analyze system performance metrics and recommend optimization opportunities. Drive observability standards across data services and platforms. Platform & Infrastructure Management Support cloud-based infrastructure and platform services across Azure, AWS, or GCP environments. Optimize compute, storage, networking, and data platform resources. Work with containerized and Kubernetes-based workloads. Ensure high availability and disaster recovery capabilities are implemented and tested. Data Platform Reliability Support modern data ecosystems including data lakes, warehouses, streaming platforms, and analytics environments. Monitor ETL/ELT pipelines, batch processing, real-time streaming, and data orchestration services. Partner with Data Engineers to improve pipeline reliability and data quality monitoring. Ensure platform scalability for growing data volumes and user demands. Security & Compliance Implement security best practices and operational controls. Support compliance requirements related to data governance and privacy. Collaborate with security teams to remediate vulnerabilities and improve platform security posture. Continuous Improvement Conduct post-incident reviews and drive corrective and preventive actions. Identify reliability risks and implement long-term improvements. Promote a culture of operational excellence, resilience, and automation. Contribute to engineering standards, runbooks, knowledge sharing, and best practices.

Requirements

Bachelor’s degree in computer science, Information Technology, Engineering, or related field, or equivalent practical experience., 3+ years of experience in Site Reliability Engineering, DevOps, Systems Engineering, Cloud Engineering, or a related role. Experience supporting production environments with high availability requirements. Experience managing cloud infrastructure and distributed systems. Technical Skills Strong knowledge of Linux systems administration. Proficiency in one or more programming languages such as Python, Java, Go, C#, or JavaScript. Experience with CI/CD tools and deployment automation. Experience with Infrastructure as Code tools such as Terraform, ARM, or Bicep. Experience with Kubernetes and container technologies. Knowledge of monitoring and observability technologies such as Prometheus, Grafana, Datadog, Azure Monitor, Splunk, or OpenTelemetry. Understanding of networking, DNS, load balancing, and distributed systems concepts. Experience supporting data platforms such as: o Azure Data Lake o Azure Synapse Analytics o Databricks o Snowflake o Kafka o SQL/NoSQL databases o Data orchestration platforms Preferred Qualifications Experience supporting AI/ML platforms and MLOps environments. Experience with Azure cloud-native services. Familiarity with data governance and data quality frameworks. Knowledge of reliability engineering best practices and SRE methodologies. Experience implementing SLOs, SLIs, and error budgets. Experience supporting large-scale analytics and business intelligence environments. Azure, AWS, Kubernetes, Terraform, or DevOps certifications. Key Competencies Problem-solving and analytical thinking Incident management and troubleshooting Automation-first mindset Collaboration and stakeholder management Strong communication skills Continuous learning and innovation Customer-focused approach Operational excellence and accountability Success Measures Maintaining high availability and reliability targets for critical services. Reducing operational toil through automation. Improving platform observability and incident response effectiveness. Meeting service-level objectives and performance goals. Enhancing deployment reliability and operational efficiency. Driving measurable improvements in system resilience, scalability, and customer experience., FIRE PROTECTION ENGINEER JOB DESCRIPTION Position Summary: Allied Fire Protection is seeking a Fire Protection Engineer with a minimum of 5 years of experience in fire protec…

  • 5 hours ago

About the company

TekWissen

  • Dallas, TX Overview: TekWissen is a global workforce management provider headquartered in Ann Arbor, Michigan that offers strategic talent solutions to our clients worldwide. Our client is …

  • 14 hours ago

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:56 min

Provisioning a secure container infrastructure with Bicep

Matthias Falkenberg +1 · World Congress 2022

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

6:16 min

Event-driven Golang backend architecture and cloud deployment

Irina Branovic Irina Branovic · World Congress 2026 Europe

4:09 min

Selecting infrastructure tools and determining proper abstraction layers

Alayshia Knighten Alayshia Knighten · World Congress 2024

Videos

See all

Related articles

See all