Software Development Engineer in Test

Socure Inc.
San Francisco, CA, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$150,000.0 - $175,000.0
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) JavaScript (Programming Language) Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Application Release Automation Automation of Tests Computer Programming Disaster Recovery Distributed Systems Python (Programming Language) Log Analysis
+14 more
Machine Learning Software Engineering TypeScript Datadog Cloud Platform System Build Management Playwright Build Tools Cloudwatch Restful APIs Splunk New Relic (SaaS) SDET Microservices

Job description

We are seeking a Senior SDET to support the next evolution of quality engineering by combining automated functional validation, production health monitoring, and AI-driven failure analysis.

This role is focused on ensuring that our Disaster Recovery (DR) and production environments are not only available, but fully functional, continuously validated, and increasingly capable of self-diagnosis. The ideal candidate will bring strong automation skills, systems thinking, and a passion for improving reliability across complex distributed systems., As a Senior SDET, you will design and build automated quality and validation systems that strengthen confidence in both production and disaster recovery readiness. You will partner closely with QA, SRE, and Engineering teams to validate critical business workflows, improve observability, reduce alert noise, and accelerate incident detection and resolution., * Design and implement automated functional health checks for DR and production environments using synthetic transactions and API validation.

  • Build continuous validation pipelines that verify end-to-end business workflows such as authentication, transactions, and system integrations.
  • Develop intelligent alerting mechanisms based on functional failures and customer-impacting behavior, not solely infrastructure metrics.
  • Integrate observability signals including logs, metrics, and traces with automated test frameworks to improve system visibility and diagnosis.
  • Develop AI/ML-driven approaches to detect failure patterns, correlate issues across services, and identify probable root causes.
  • Build systems that recommend or trigger automated remediation actions to support early-stage self-healing capabilities.
  • Partner cross-functionally with QA, SRE, and Engineering teams to improve service reliability, incident response, and recovery readiness.
  • Define, measure, and report on functional SLAs, service health indicators, and quality metrics.
  • Contribute to disaster recovery drills, readiness exercises, and automated validation efforts that improve resilience over time.

Requirements

Do you have experience in Tooling?, * 5+ years of experience in QA Automation, SDET, Software Engineering, or a related technical discipline.

  • Strong experience building and maintaining automated test frameworks, including tools such as Playwright, Jest, SuperTest, and REST API testing frameworks.
  • Experience working in cloud environments, preferably AWS.
  • Familiarity with observability and monitoring platforms such as Datadog, New Relic, CloudWatch, Splunk, or similar tools.
  • Strong programming skills in TypeScript, Python, Java, or similar languages.
  • Experience designing end-to-end test strategies for distributed systems and production-like environments.
  • Strong problem-solving skills, with the ability to analyze failures across application, infrastructure, and workflow layers., * Experience with synthetic monitoring, production validation, or proactive health-checking systems.
  • Exposure to AI/ML techniques for anomaly detection, log analysis, or failure correlation.
  • Experience with CI/CD pipelines, release automation, and validation gates.
  • Understanding of microservices architecture, distributed system failure modes, and incident management practices.
  • Familiarity with SRE concepts such as SLIs, SLOs, error budgets, or production-readiness practices.

About the company

Socure is building the identity trust infrastructure for the digital economy - verifying 100% of good identities in real time and stopping fraud before it starts. The mission is big, the problems are complex, and the impact is felt by businesses, governments, and millions of people every day.

We hire people who want that level of responsibility. People who move fast, think critically, act like owners, and care deeply about solving customer problems with precision. If you want predictability or narrow scope, this won’t be your place. If you want to help build the future of identity with a team that holds a high bar for itself - keep reading.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:38 min

Establishing comprehensive monitoring and log management

Michael Eder +1 · LIVE

5:34 min

Managing token budgets and enterprise usage of coding agents

Chris Heilmann +2 · LIVE

2:03 min

Automating the complete functional software testing development lifecycle

April Yoho April Yoho · WWC Europe 2026

3:10 min

Correlating dispersed logs using structured request tracing

Michael Eder +1 · LIVE

2:16 min

Structuring technical interviews to evaluate candidate culture and observability

Tejas Chopra Tejas Chopra · LIVE

1:08 min

Analyzing error logs and root causes using artificial intelligence

Nishil Patel Nishil Patel · WWC 2025

Videos

See all

Related articles

See all