Site Reliability Engineer II

CIS Technologies Inc.
United States
2 days ago
Apply on www.dice.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Working hours
Regular working hours
Job source

Tech stack

Artificial Intelligence BigQuery Cloud Computing Distributed Data Store Monitoring of Systems Python (Programming Language) Reliability Engineering Data Driven Tests Datadog Computer Networking Systems Google Cloud Large Language Models
+5 more
Data Management New Relic (SaaS) Dynatrace Servicenow Programming Languages

Job description

We are seeking a Site Reliability Engineer to join an SRE team focused on observability, monitoring, and technical consulting across Google Cloud Platform-based data platforms. This role is responsible for ensuring the availability, reliability, and performance of cloud and network systems and services through automation, monitoring, troubleshooting, and continuous optimization., * Collaborate with infrastructure teams to implement critical solutions by automating routine tasks.

  • Monitor and manage production environments, proactively identifying and resolving issues.
  • Participate in building advanced tooling for system access monitoring, log session recording, and reliability administration across multiple geographically distributed data centers.
  • Engage with engineering teams to improve on-call efficiencies, incident management, and post-mortem analysis.
  • Perform capacity planning and optimization to support growing demands and traffic patterns.
  • Maintain monitoring and alerting systems for proactive system health checks.
  • Continuously improve system performance, stability, and security through data-driven analysis and optimization.
  • Create and maintain comprehensive documentation and diagrams to facilitate knowledge sharing.
  • Work hands-on with cloud infrastructure, BigQuery workloads, CI/CD pipelines, and enterprise monitoring tools to maintain critical systems at scale.

Requirements

  • Bachelor’s degree.
  • 4+ years of experience in IT.
  • 3+ years of development experience.
  • Practitioner-level experience with at least one coding language or framework.
  • Hands-on experience with Google Cloud Platform (Google Cloud Platform).
  • Experience with BigQuery.
  • Experience with Dynatrace.
  • Proficiency with monitoring and observability tools, ideally Dynatrace or comparable tools such as Datadog or New Relic.
  • Familiarity with ITSM tools such as ServiceNow, including incident, problem, and change management., * Experience with Google Cloud Platform Cloud Run.
  • Experience with Python.
  • Strong troubleshooting and problem-solving skills.
  • Familiarity with AI tools, including agents, skills, LLMs, and copilots.
  • Experience defining and tracking SLAs, SLOs, and SLIs.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

2:30 min

Leveraging BigQuery ML for scalable SQL-based segmentation experiments

Julian Joseph · LIVE

1:01 min

Connecting frontend application performance to user retention and revenue

Dani Coll Dani Coll · World Congress 2025

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

1:08 min

Analyzing error logs and root causes using artificial intelligence

Nishil Patel Nishil Patel · World Congress 2025

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all