Senior Software Engineer - Platform

Dremio
Santa Clara, CA, United States
23 days ago
Apply on arc.dev
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
2 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Artificial Intelligence Amazon Web Services Apache HTTP Server Application Frameworks Microsoft Azure C++ (Programming Language) Software as a Service Concurrency Controls Continuous Integration Dependency Injection DevOps
+31 more
Distributed Systems Domain Name System (DNS) Hypertext Transfer Protocols (HTTP) Python (Programming Language) MongoDB Networking Basics Node.Js NoSQL Role-Based Access Control Redis Distributed Caching SAP (Applications) SAP HANA Software Engineering SQL Databases TCP/IP Rust (Programming Language) Multithreading Cloud Platform System Caching Kubernetes Information Technology Database Replication Asynchronous Programming Terraform Heap (Data Structure) Docker Jenkins Golang Programming Languages Microservices

Job description

The Observability and Core Services team owns the platform foundation that Dremio is built on - and that every product team depends on to ship. As a core member, you’ll design and deliver the critical building blocks: telemetry, messaging, KV store and persistence, application frameworks, caching, coordination, and events and audit services, along with Dremio’s growing integration with the SAP platform. When these systems are rock-solid, the whole engineering org moves faster.

You’ll work at the intersection of DevOps/SRE and application development - designing solutions that are highly scalable and genuinely easy to use - while collaborating with core product teams to ship safely and partnering with Field teams to support customers at enterprise scale.

This is also where Dremio’s hardest distributed systems problems live. You’ll tackle consistent persistence, distributed caching, networking, containerized services, private links, Terraform-driven infrastructure, and access control at scale. The scope is broad by design: we believe the best platform engineers are the ones who can hold the full picture and own problems end-to-end.

You’ll grow by taking on complex challenges, shipping high-quality systems at massive scale, and building the kind of foundational work that compounds across everything Dremio builds.

What You’ll Be Doing

  • Build and own an observability platform that ingests customer telemetry - events, logs, and traces - using OpenTelemetry and industry best practices
  • Deliver dashboarding and visualization capabilities that give Support and Field teams real-time visibility into system health, emerging issues, and trends - and help customers stay resilient at scale
  • Design and implement a notification and alerting system driven by telemetry data and heuristics with AI agents and lakehouse intelligence
  • Write infrastructure-as-code with Helm and Terraform to simplify deployment and operations within Dremio Software
  • Build and maintain application frameworks that standardize development across teams - tackling cross-cutting challenges like dependency injection and role-based access control at scale
  • Partner with our SRE and DevOps teams to define best practices and deliver standardized solutions for common engineering workflows
  • Engage directly with the developer community to understand their pain points and turn those insights into platform improvements
  • Help define and deliver roadmap and integration milestones with SAP HANA/BDC platform to deliver a cohesive experience, Dremio is not responsible for any fees related to unsolicited resumes and will not pay fees to any third-party agency or company that does not have a signed agreement with the Company.

Requirements

  • Bachelor’s, Master’s, or higher in Computer Science or a related field
  • 5+ years of software engineering experience
  • 2+ years experience using Java
  • 2+ years hands-on with infrastructure-as-code (Helm, Terraform, or equivalent)
  • Deep experience with a major cloud platform (AWS, Azure, or GCP)
  • Hands-on with OpenTelemetry or comparable observability platforms
  • Experience with Kubernetes and Docker
  • Solid understanding of networking fundamentals (TCP/IP, DNS, HTTP)
  • Desire to learn new technologies and proven problem solver

Bonus points if you have

  • Experience building SaaS microservices and distributed systems
  • Hands-on with multi-threaded or asynchronous programming models
  • Proficient with NoSQL databases (MongoDB, RocksDB) or Redis
  • Background in CI/CD tooling (Jenkins, ArgoCD)
  • Other programming languages such as C++, Go, Python, Node.js, Rust, etc
  • Experience in query processing, distributed systems, concurrency control, data replication, or storage systems
  • Java GC/heap tuning, Apache Arrow, SQL operators, or data-path optimization

About the company

About Dremio

51-200

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on arc.dev
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:55 min

Demonstrating semantic routing thresholds with the Redis vector library

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

1:08 min

Building solutions with open source GoLang infrastructure tools

Jad Wahab · LIVE

2:14 min

Solving complex platform architecture challenges at an enterprise scale

Maria Apazoglou · Coffee With Developers

3:42 min

Comparing in-memory and Redis storage for cache scalability

Simone Sanfratello · World Congress 2022

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all