Sr. Site Reliability Engineer

VDart, Inc.
Frisco, TX, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$103,000.0 - $147,000.0
Working hours
Regular working hours

Tech stack

Java (Programming Language) Agile Methodology Artificial Intelligence Amazon Web Services Audit Trail Microsoft Azure Bash Shell Oracle WebLogic Server Unix Cloud Engineering Computer Programming Databases
+51 more
Continuous Integration Information Engineering Data Governance Data Systems Cursor (Graphical User Interface Elements) DevOps Disaster Recovery Middleware Monitoring of Systems Python (Programming Language) Key Management Log Analysis Windows Servers MySQL Oracle Data Service Integrator Oracle (Applications) Productivity Software RabbitMQ Reliability Engineering Ansible Standard Sql Software Engineering SQL Databases Private Cloud Environment Data Logging Scripting Cyberark System Availability Delivery Pipeline Large Language Models Snowflake Prompt Engineering Reliability of Systems Generative AI Gitlab Containerization Kubernetes Infrastructure Automation Frameworks Information Technology Deployment Automation Apache Kafka Virtual Agents Restful APIs Terraform Splunk Appdynamics Data Pipelines Docker Jenkins Databricks Microservices

Job description

Senior Engineer, Systems Reliability (SRE) - Privacy ensures the stability, performance, and reliability of IT services and infrastructure. This role combines software engineering and operations expertise to build and maintain highly available, scalable systems. As a leader in DevOps and cloud reliability practices, the engineer supports continuous improvement of automation, deployment pipelines, observability, and incident management, while mentoring junior engineers and optimizing production workflows. The position plays a critical part in enabling software to be delivered faster, better, and more reliably to support business and customer needs. What You ll Do

  • Build and maintain CI/CD pipelines for data engineering deployments using GitLab and Azure DevOps Design and maintain CI/CD pipelines and DevOps automation solutions for REST APIs and microservices.
  • Implement robust monitoring, alerting, and logging for data pipelines, Snowflake and Azure services.
  • Respond to production incidents, troubleshoot failures and restore services quickly.
  • Perform root cause analysis and implement preventive measures.
  • Ensure high availability and disaster recovery planning for critical data systems.
  • Tune SQL queries, Snowflake features and Databricks clusters for optimal performance and cost efficiency.
  • Automate operational tasks to improve deployment reliability and reduce manual intervention. Manage secrets and credentials using Azure Key Vault and CyberArk.
  • Hands-on experience with Terraform, Helm, or Ansible for infrastructure provisioning
  • Working knowledge of containerization (Docker) and Kubernetes orchestration Hands-on experience with cloud platforms (Azure; AWS or GCP)
  • Understanding of deployment strategies (blue/green, rolling, canary), GitOps, and artifact management
  • Ensure compliance with data governance, privacy regulations and organizational security standards.
  • Work closely with data engineers, analysts and cloud teams to ensure smooth operations.
  • Maintain detailed runbooks, operational documentation and incident reports.
  • Perform regular OS patching on Unix and Windows servers to address security vulnerabilities and maintain system stability.
  • Apply critical and cumulative updates for middleware components such as Oracle Data Integrator (ODI), WebLogic and related software to mitigate risks and enhance performance.
  • Coordinate patching schedules with application and infrastructure teams to minimize downtime and ensure business continuity.
  • Use AI productivity tools daily (Claude and Cursor or similar IDE) across the SRE lifecycle including pipeline development, scripting, runbook authoring, log analysis, and incident response Design, build, and operate AI agents to automate SRE tasks such as incident triage, root cause analysis, alert correlation, runbook execution, and patching workflows
  • Apply foundation models, prompt engineering, and RAG patterns to operational use cases such as querying runbooks, summarizing incidents, and surfacing remediation guidance etc but not limited to these areas.
  • Implement audit logging, observability, and human-in-the-loop controls for AI agents and AI-assisted workflows operating in Tier-0 production environments
  • Build and host AI agents, identify gaps and convert them into AI agent use cases, and implement solutions to further modernize the SRE platform

Requirements

  • Bachelor s degree in computer science, Engineering, or equivalent practical experience 5-7 years of experience in systems reliability, software engineering, DevOps, or related technical roles
  • Experience working in Agile and DevOps delivery environments Demonstrated ability to mentor engineers and influence technical outcomes
  • Strong problem-solving skills with a systems-level perspective Strong automation, and agentic AI skills.
  • Familiarity with foundation models, prompt engineering, retrieval-augmented generation (RAG), and AI agent development applied to SRE and operational use cases

Must Have Skills

  • CI/CD tooling and automation experience (gitlab, azure devops, jenkins)
  • Experience working in public or private cloud environments Proficiency in one or more programming or scripting languages (Python, Java, Shell, etc.)
  • Experience with monitoring, logging, and APM tools such as AppDynamics, Splunk, or equivalents Strong understanding of system reliability concepts including scalability, performance, availability, and resilience Strong experience in writing SQLs, analyzing logs and troubleshooting issues.
  • Databases: SQL (Oracle/My SQL/ Snowflake)
  • Messaging: Kafka, Rabbit MQ
  • Hands-on experience with AI productivity tools (Claude and Cursor or similar IDE) and working knowledge of foundation models, prompt engineering, RAG, and AI agent development
  • Experience with containerization and orchestration technologies such as Docker and Kubernetes

Nice to Have

  • Experience migrating systems to cloud-native architectures
  • Familiarity with reliability metrics, service monitoring, or operational dashboards
  • Exposure to platform engineering or shared services environments

Key Skills: CI/CD, automation, AppDynamics, Splunk, Kafka, Rabbit, AI productivity tools

About the company

Bright Vision Technologies

  • Little Elm, TX
  • $100,000-150,000 per year Bright Vision Technologies is a forward-thinking software development company dedicated to building innovative solutions that help businesses automate and optimize their operations…

  • 2 days ago

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.careerjet.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:03 min

Microsoft integrating native Unix coreutils into Windows environments

Chris Heilmann +2 · LIVE

2:18 min

Scaling MySQL databases for massive user growth

Johannes Nicolai Johannes Nicolai +1 · LIVE

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

2:04 min

Defining timestamps and the international standard format

Denny Biasiolli Denny Biasiolli · Europe 2026 Virtual

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

Videos

See all

Related articles

See all