> Markdown version of [/jobs/ext/2727917-software-development-engineer](https://www.wearedevelopers.com/jobs/ext/2727917-software-development-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Development Engineer - **Company:** Workday, Inc. - **Location:** United States (Remote available) - **Experience:** Experienced - **Salary:** $148,000.0 - $222,000.0 - **Contract:** Permanent contract - **Skills:** Query Performance, Java (Programming Language), Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Amazon S3, Apache HTTP Server, Computing Platforms, Architectural Patterns, Big Data, Continuous Integration, Data Infrastructure, Data Security, Distributed Data Store, Distributed Systems, Elasticsearch, Fault Tolerance, Github, Python (Programming Language), Load Testing, Online Transaction Processing, Query Optimization, Runbook, Software Engineering, Data Logging, Data Ingestion, System Availability, Grafana, Apache Spark, Backend, Data Lakes, AI Platforms, Integration Tests, Information Technology, Code Testing, Apache Flink, Apache Kafka, Bitbucket, Vertica, Api Design, Software Version Control, Data Pipelines, Dynatrace, Workday, Golang - **Published:** September 5, 2026 - **Apply:** https://startup.jobs/software-development-engineerdistributed-systems-workday-9874136 ## About the Role 5+ years experience in software development engineering. 3+ years experience specifically focused on designing, building, and operating complex distributed system architectures, evidenced by successful deployment of systems with high availability (e.g., 99.9% uptime) and fault tolerance. 5+ years experience with at least two of the following programming languages Java, Python, Go, including experience in writing production-level code for distributed systems. Bachelor's degree in a relevant field such as Computer Science, Engineering, or a related discipline; a Master's degree (e.g., MS in Computer Science, Distributed Systems, or related field) is strongly preferred or equivalent practical experience. Other Qualification * Strong ability in Algorithmic to build highly efficient and scalable solutions for complex high-throughput data ingestion and sub-second query performance challenges. * Solid experience in API Development, including an understanding of gRPC, REST, and OpenTelemetry (OTLP), with practical experience designing and building scalable distributed APIs for observability data. * Strong understanding of Code Testing methodologies, such as distributed load testing and integration testing, and experience contributing to end-to-end telemetry pipeline testing and CI/CD automation. * Solid understanding of Distributed Systems Software principles, including data partitioning, eventual consistency, and fault tolerance mechanisms, with hands-on experience in Kafka, Spark, Flink, or ClickHouse. * Experience implementing and maintaining High Availability strategies for critical distributed systems, including multi-AZ deployments, robust retry mechanisms, and automated failover. * Practical experience with Large Scale Data Processing technologies and frameworks such as Apache Kafka, Spark, Flink, and Apache Iceberg within complex distributed architectures. * Good understanding of Large Scale Systems design principles, including distributed data sharding, replication, and query optimization, and experience working on observability pipelines or data lake platforms. * Strong understanding of Object-Oriented Design (OOD) principles and architectural patterns for building highly scalable and maintainable distributed systems. * Experience with Source Control Management (SCM) tools such as Git and GitHub/Bitbucket, and following best practices for collaborative distributed development workflows. * Strong understanding of System Security principles and best practices relevant to securing distributed environments, including mutual TLS (mTLS), multi-tenant authorization (authz), and data encryption. * Proven ability to actively collaborate within and across distributed software development teams and contribute constructively to architectural discussions and system designs. * Strong skills in creating Technical Writing Documentation for runbooks, system design specs, and API documentation related to distributed systems architecture and design. Workday Pay Transparency Statement ## Description We're obsessed with making hard work pay off, for our people, our customers, and the world around us. As a Fortune 500 company and a leading AI platform for managing people, money, and agents, we're shaping the future of work so teams can reach their potential and focus on what matters most. The minute you join, you'll feel it. Not just in the products we build, but in how we show up for each other. Our culture is rooted in integrity, empathy, and shared enthusiasm. We're in this together, tackling big challenges with bold ideas and genuine care. We look for curious minds and courageous collaborators who bring sun-drenched optimism and drive. Whether you're building smarter solutions, supporting customers, or creating a space where everyone belongs, you'll do meaningful work with Workmates who've got your back. In return, we'll give you the trust to take risks, the tools to grow, the skills to develop and the support of a company invested in you for the long haul. So, if you want to inspire a brighter work day for everyone, including yourself, you've found a match in Workday, and we hope to be a match for you too. About the Team The Data Platform and Observability Engineering (DPOE) team is building Workday's next-generation, multi-petabyte scale Observability Platform. We own the libraries, distributed services, and infrastructure that power ingestion, storage, and query across the observability stack - Iceberg, ClickHouse, Tempo, Grafana, S3, Kafka, and Elasticsearch - serving traces, metrics, and logs for every workload at Workday. Our roadmap directly shapes how the company detects, diagnoses, and eventually predicts operational issues at scale. About the Role As a Software Development Engineer, you will build and scale core features of Workday's distributed tracing platform. Working with ClickHouse, Grafana Tempo, and a modern big-data pipeline (Kafka, Spark/Flink, Iceberg, S3) on AWS, you will deliver high-quality, high-performance code capable of handling massive data loads. This is a hands-on role for an engineer who excels at building resilient backend services and is eager to grow their expertise in distributed systems and big data while contributing to the future of Observability AI., * Develop the Platform: Write high-quality, well-tested code to build and scale features for the distributed tracing platform on ClickHouse/Tempo, focusing on robust ingestion and fast query execution. * Maintain Big-Data Pipelines: Build and support high-throughput data pipelines using Kafka, Spark/Flink, and Iceberg-on-S3. * Tune and Optimize: Write optimized code and participate in profiling and resolving latency or throughput issues in production environments. * Ensure Reliability: Implement robust error handling, retries, and failover mechanisms to ensure high availability for tracing services. * Secure the Data: Apply necessary security controls to ensure multi-tenant data access is properly handled according to platform architecture. * Operational Support: Maintain the health of the platform through comprehensive monitoring, logging, and alerting, and participate in the team's on-call rotation. * Support Observability AI: Build reliable data ingestion paths that will serve as the foundation for AI-driven anomaly detection and root-cause analysis. * Collaborate and Learn: Write clear technical documentation, participate actively in system design reviews, and work closely with Senior and Principal engineers to level up your distributed systems knowledge. ## Related Videos - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Destigmatizing the Workplace: Building Real Inclusion](https://www.wearedevelopers.com/videos/1492-destigmatizing-the-workplace-building-real-inclusion) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [How building an industry DBMS differs from building a research one](https://www.wearedevelopers.com/videos/768-how-building-an-industry-dbms-differs-from-building-a-research-one) - [Scoring 2000 Products per Request: Performance Pitfalls in Golang](https://www.wearedevelopers.com/videos/2073-scoring-2000-products-per-request-performance-pitfalls-in-golang) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [How Much FAANG Companies Actually Pay Software Engineers in 2025](https://www.wearedevelopers.com/magazine/230-how-much-faang-companies-actually-pay-software-engineers-in-2025) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)