Site Reliability Engineer - SRE

Bitscopic Inc.
United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$140,000.0 - $190,000.0
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Microsoft Azure Backup Devices Bash Shell Cloud Computing Cloud Engineering System Configuration Extract Transform Load (ETL) Disaster Recovery Monitoring of Systems Virtual Private Networks (VPN) Python (Programming Language)
+23 more
Network Architecture Release Management Reliability Engineering Cloud Services Ansible Prometheus Software Engineering SQL Databases Data Streaming Virtual Local Area Networks System Availability Grafana Software Security Reliability of Systems Kubernetes Information Technology Performance Monitor Fortinet Terraform Data Pipelines Docker Elk Stack Jenkins

Job description

We are seeking a skilled Site Reliability Engineer to join our team. This position is ideal for an individual with a strong background in software engineering and operations who has experience working with federal agencies such as the Department of Veterans Affairs(VA). The candidate will be responsible for ensuring the reliability and scalability of our services by effectively monitoring, deploying, and maintenance of our infrastructure., * Develop, maintain, and continuously improve CI/CD pipelines, branch management practices, and release management processes.

  • Act as primary system administrator for infrastructure in VA environments, ensuring system availability, performance, and compliance with federal security standards.
  • Design, implement, and operate monitoring and alerting solutions; validate and triage alerts as the first line of defense for operational issues.
  • Proactively identify and resolve scalability, reliability, and performance risks across applications, infrastructure, and data pipelines.
  • Lead and participate in incident response, including on-call and off-hour support, root cause analysis, and post-incident remediation.
  • Manage ETL pipelines and data flow operations, including health checks, restart automation, recovery workflows, and operational validation.
  • Coordinate closely with VA’s IT team, federal stakeholders, and internal development teams to resolve infrastructure and operational incidents.
  • Own and evolve disaster recovery, backup, and resilience strategies.
  • Document and maintain runbooks, SOPs, system configurations, deployment procedures, and escalation protocols.
  • Evaluate and introduce new tools and technologies to improve system reliability, security, and operational efficiency.

Requirements

Do you have experience in System performance monitoring?, * Bachelor’s degree in Computer Science, Engineering, or a related field.

  • Minimum of 3 years of experience as a Site Reliability Engineer or similar role, particularly in environments subject to federal compliance requirements.
  • Experience working directly with federal agencies, preferably the Department of Veterans Affairs, including navigating their compliance, security, and operational protocols.
  • Strong understanding of monitoring tools and software (e.g., Prometheus, Grafana, ELK stack).
  • Experience with backup, disaster recovery, and failover planning.
  • Proficient with containerization and orchestration technologies (e.g., Docker, Kubernetes).
  • Experienced with cloud services (e.g., AWS, Azure) and managing cloud infrastructure.
  • Proficient in scripting languages (e.g., Python, Bash) for automation and SQL.
  • Excellent problem-solving skills and ability to work under pressure.

Bonus if you have:

  • Certification in cloud technologies and security standards.
  • Experience with infrastructure as code (e.g., Terraform, Ansible).
  • Experience working with Jenkins and/or other CI/CD pipelines management tools
  • Familiarity with VPNs, VLANs, and secured network architecture - particularly for healthcare or government systems.
  • Experience with cloud and application security platforms, such as WAF(e.g. SignalScience/Fastly), RASP (e.g. TCell/Rapid7, Tenable), CNAPP (e.g. Lacework/Fortinet, Wiz) etc.

Benefits & conditions

Pulled from the full job description

  • 401(k)
  • Health insurance
  • Paid time off
  • Vision insurance
  • Health savings account
  • Dental insurance
  • Life insurance, * Mission-Driven - We are self-funded, owned by the team, and entirely focused on making a meaningful impact in healthcare while maintaining collaborative, transparent communication and ego-free team culture.
  • Team Culture - We all value work-life balance, personal and professional growth, competitive compensation, continuous learning, and fun. We build essential products to positively impact and save lives. On our team, everyone’s voice is a vital contribution to the mission. Most of our team members have been with us for over six years.
  • Freedom to Focus - Leverage flexibility, autonomy, and schedule to do your best work. Control your approach to achieve results, along with the responsibility of delivering for the team and our customers.
  • Location Independent - Work from anywhere. We’ve been a distributed team since our founding in 2012.

Pay: $140,000.00 - $190,000.00 per year, * 401(k)

  • Dental insurance
  • Health insurance
  • Health savings account
  • Life insurance
  • Paid time off
  • Vision insurance

About the company

Bitscopic is a data analytics company at the intersection of artificial intelligence, bioinformatics, and decision science. Our tailored software suite empowers healthcare organizations to reduce costs, improve patient care outcomes, and mitigate the spread of infectious diseases. Bitscopic also provides solutions for managing complex clinical trials, operational efficiency, and other bioinformatics areas.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · WWC 2025

1:02 min

Applying an ETL methodology to infrastructure configuration management

Axel Barbier · WWC 2023

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · WWC 2023

2:34 min

Docker sandbox architecture and microVM environment integration

Manuel de la Peña Manuel de la Peña · WWC Europe 2026

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

Videos

See all

Related articles

See all