> Markdown version of [/jobs/ext/2116712-software-engineer](https://www.wearedevelopers.com/jobs/ext/2116712-software-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # software engineer - **Company:** ยท Sentinelone - **Location:** Atlanta, GA, United States (Remote available) - **Experience:** Expert - **Salary:** $156,000.0 - $215,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Microsoft Windows, Artificial Intelligence, Amazon Web Services, Apple Mac Systems, Software as a Service, Cloud Computing, Software Quality, Databases, Data Integrity, Linux, Distributed Systems, Github, Internet Security, Python (Programming Language), PostgreSQL, Enterprise Messaging Systems, Octopus Deploy, Redis, Reliability Engineering, Requirements Management, Server Administration, Software Engineering, Scripting, Apache Cassandra, System Availability, Backend, Data Layers, Kubernetes, Cassandra, Apache Kafka, Graphql, Data Management, Vertica, Restful APIs, Code Restructuring, Multiplatform, SentinelOne Expertise, Docker, Golang - **Published:** August 19, 2026 - **Apply:** https://www.careerbuilder.com/job-details/staff-backend-software-engineer-agent-platform-atlanta-ga--eb679170-ecce-4116-bb64-6c69fba6e800 ## About the Role * 8+ years of professional backend software engineering experience with deep expertise in at least one of Java, Go, or Python, and the willingness to work across all three as needed. * A reliability-first mindset with a proven ability to solve complex production incidents, perform systematic root-cause analysis, and implement durable engineering solutions that prevent recurrence. * Demonstrated ability to quickly understand unfamiliar distributed systems, navigate large codebases, and troubleshoot complex failures across service boundaries. * Strong hands-on experience designing, building, and operating large-scale distributed systems, with a deep understanding of failure modes, performance trade-offs, resilience patterns, and operational excellence. * Experience with AWS, GCP, or similar cloud platforms, as well as Docker, Helm, and Kubernetes. * Experience with messaging systems and data platforms such as Kafka, PostgreSQL, Redis, Cassandra, ClickHouse, or similar technologies. * Excellent communication and collaboration skills, with the ability to work effectively across engineering teams, Product, Technical Account Managers, and other stakeholders, influence technical direction, and mentor fellow engineers. * A high degree of ownership, autonomy, curiosity, and the ability to drive ambiguous technical problems to successful outcomes. * Experience in an enterprise SaaS or cybersecurity software company is highly desirable., Amazon Web Services (AWS), Apache Cassandra, Apache Kafka, Architectural Services, Artificial Intelligence (AI), Automation, Cancer, Cellular Telephone, Cloud Computing, Coaching, Communication Skills, Continuous Improvement, Cross-Functional, Customer Support/Service, Data Quality, Distributed Computing, Docker, Editing, Employee Assistance Plan, Endpoint Security, Engineering, Flexible Spending Accounts, GCP (Good Clinical Practices), GitHub, Go Programming Language (Golang), GraphQL, Health Insurance, High Availability, High Tech Industry, High Throughput, Identify Issues, Improvement Metrics, Incident Response, Insurance, International Business, Internet Security, Java, Large-Scale Systems, Legal, Linux Operating System, Mac Operating System, Mentoring, Messaging Technology, Microsoft Windows Operating System, Multiplatform/Cross-Platform, Operational Improvement, Policy Development, PostgreSQL, Problem Solving Skills, Production Systems, Productivity Management, Python Programming/Scripting Language, REST (Representational State Transfer), Redis, Refactoring, Reimbursement, Reliability Engineering, Requirements Management, Root Cause Analysis, Sales Management, Security Attacks, Software Administration, Software Engineering, Software as a Service (SaaS), Stock Purchase Plans, Team Player, Technical Delivery, Technical Leadership, Test Plan/Schedule, Willing to Travel ## Description As a Staff Software Engineer on the Agent Platform team, you will help secure tens of millions of devices across Windows, Linux, and macOS while processing billions of security events every day as part of SentinelOne's Endpoint Protection product line. We are looking for an excellent software engineer with a reliability-first mindset who thrives on solving complex distributed systems challenges in production and turning those learnings into durable platform improvements and long-term engineering ownership. You'll help drive customer-critical incident response across the Agent Platform, rapidly build context across multiple platform services, and partner with engineering teams to continuously improve the platform's reliability, operability, and scalability as a whole. You will build and evolve the high-throughput, highly available services responsible for policy, configuration, and command distribution to millions of agents worldwide, while contributing to the agent platform protocols that enable other SentinelOne teams to deliver new security capabilities safely and at scale. This role offers broad technical ownership, the opportunity to work across service boundaries, and the ability to influence the design and reliability of critical platform capabilities used throughout the Agent Platform at SentinelOne. You'll solve some of the company's most complex production challenges while helping shape the future of our platform. What Will You Do? Primary responsibilities include: * Drive rapid response to customer-critical incidents by diagnosing, triaging, and resolving complex production issues that span multiple services across the Agent Platform. * Lead systematic root-cause analysis and translate incident learnings into durable reliability, scalability, observability, and operability improvements. * Quickly build a systems-level understanding of unfamiliar services and codebases, collaborating closely with engineering teams to resolve complex cross-service issues. * Design, develop, test, document, deploy, and operate large-scale, high-volume, low-latency distributed systems processing millions of events per second. * Understand, maintain, and continuously improve existing services through refactoring, feature development, and architectural enhancements. * Maintain application stability and data integrity by monitoring key metrics, improving operational visibility, and continuously strengthening the codebase. * Translate business and functional requirements into robust, scalable, and operable technical solutions. * Partner closely with engineering teams across SentinelOne to solve cross-functional problems, influence technical direction, and deliver solutions that scale. * Continuously evaluate and adopt technologies that improve the platform's scalability, reliability, and operational excellence. Your Toolkit * Our backend services are primarily developed in Java, with Go and Python also playing important roles. * We use gRPC, REST, GraphQL, and Kafka for service communication. * Our data layer includes Redis, PostgreSQL, Cassandra, ClickHouse, and our own columnar time-series database. * Our services run across multiple AWS and GCP regions on Kubernetes, using Docker, GitHub, and ArgoCD. * We provide engineers with modern AI-powered tools to improve both engineering productivity and software quality. ## Related Videos - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Accelerating Authentication Architecture: Taking Passwordless to the Next Level](https://www.wearedevelopers.com/videos/733-accelerating-authentication-architecture-taking-passwordless-to-the-next-level) - [Psychological Safety in Software Engineering - Jenny-Margrethe Vej & Alexandra Hou Aldershaab](https://www.wearedevelopers.com/videos/2142-psychological-safety-in-software-engineering-jenny-margrethe-vej-alexandra-hou-aldershaab) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Best Countries for Software Engineers](https://www.wearedevelopers.com/magazine/267-best-countries-for-software-engineers) - [Why Upskilling And Reskilling is Important For Developers](https://www.wearedevelopers.com/magazine/428-why-upskilling-and-reskilling-is-important-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)