> Markdown version of [/jobs/ext/2715290-systems-engineer](https://www.wearedevelopers.com/jobs/ext/2715290-systems-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # systems engineer - **Company:** Amplitude, Inc. - **Location:** New York, NY, United States (Remote available) - **Experience:** Expert - **Salary:** $198,000.0 - $299,000.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Amazon Web Services, Amazon Elastic Compute Cloud, Amazon S3, Big Data, BigQuery, C++ (Programming Language), Cloud Computing, Program Optimization, Profiling, Code Review, Encodings, Databases, Data Structures, Data Systems, Software Debugging, Distributed Data Store, Distributed Systems, Memory Management, Amazon DynamoDB, Middleware, Failover, Java Virtual Machine (JVM), Python (Programming Language), Online Analytical Processing, Performance Tuning, Redis, Data Streaming, Usage Analysis, Parquet, Multithreading, Amazon ElastiCache, Snowflake, Concurrency, Backend, Kubernetes, Druid, Apache Kafka, Presto, Vertica, Terraform, Data Pipelines - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/staff-software-engineer-infrastructure-amplitude-9636331 ## About the Role * 7+ years of industry experience in backend or infrastructure engineering, with depth in distributed data systems. * Hands-on experience building or extending analytical/OLAP systems - query engines, columnar storage, large-scale data processing frameworks, or equivalent. * Track record of driving significant cost optimization on cloud infrastructure at scale (compute, storage, network). * Strong computer science fundamentals: distributed systems (partitioning, replication, consistency, failover), data structures and algorithms, concurrency and multi-threading, performance optimization. * Production experience with modern cloud infrastructure - AWS (S3, DynamoDB, EC2), Kafka, Redis/ElastiCache, Kubernetes, Terraform - or strong equivalents. * Proficiency in Java, C++, or Python. * Demonstrated technical influence beyond your immediate team: leading design discussions, driving cross-team alignment, mentoring engineers. Nice to Have * Experience with specific OLAP or query engine systems: Druid, ClickHouse, Presto/Trino, BigQuery, Snowflake, or similar. * Deep JVM expertise - GC tuning, profiling, memory optimization at production scale. * Experience with columnar data formats and encodings (Arrow, Parquet, ORC, or custom formats). * Familiarity with product analytics, experimentation platforms, or event-driven data systems. * Contributions to open-source data infrastructure projects or published work in the data systems space. ## Description * Work across Nova's query execution engine and distributed compute layer: query planning, columnar storage formats, encoding and compression, caching, and cluster-level resource management. * Design and implement new capabilities as Nova expands to support more warehouse-imported data types, such as metrics, profiles, and dimensions. * Design for high-throughput automated query workloads - as AI agents become a primary source of queries, ensure Nova's architecture supports sustained, concurrent, and programmatic query patterns at scale. Drive cost and performance at scale * Own and execute projects that materially reduce infrastructure cost - compute, storage, network, and memory - while maintaining or improving latency and throughput. * Profile and optimize JVM performance: GC tuning, memory management, concurrency, and data layout decisions that compound at our scale. * Build guardrails and observability to catch expensive or pathological queries before they impact the system. Improve reliability and operational excellence * Strengthen Nova's reliability posture: identify systemic failure modes, drive durable fixes, and raise the bar on how we detect and respond to production issues. * Participate in on-call rotation to root-cause incidents and turn one-off fixes into architectural improvements. * Contribute to capacity planning, safe rollout practices, and the operational tooling that keeps Nova healthy. Influence through technical leadership * Lead the design and execution of multi-month projects that improve Nova's architecture, performance, or capabilities. * Contribute to technical direction through design docs, architecture discussions, and code reviews - helping the team make principled tradeoffs. * Mentor senior engineers on distributed systems thinking, production debugging, and system design. * Collaborate with Product, Middleware, Data Pipeline, and other engineering teams to ensure Nova's capabilities translate into customer value., * Gets energy from working deep inside a complex distributed system - understanding how data flows through it, where the bottlenecks are, and how to make it meaningfully better. * Has built or significantly extended an OLAP engine, columnar database, query processor, or large-scale data processing system - not just operated one. * Thinks about cost, performance, and reliability as interconnected concerns, not separate workstreams. * Communicates clearly about technical tradeoffs and earns influence through the quality of your work and ideas, not through title. * Finds it natural to help other engineers level up - through pairing, design reviews, or just being the person who explains the "why" behind a system's design. ## Related Videos - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [How building an industry DBMS differs from building a research one](https://www.wearedevelopers.com/videos/768-how-building-an-industry-dbms-differs-from-building-a-research-one) - [Accelerating Authentication Architecture: Taking Passwordless to the Next Level](https://www.wearedevelopers.com/videos/733-accelerating-authentication-architecture-taking-passwordless-to-the-next-level) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [Event based cache invalidation in GraphQL](https://www.wearedevelopers.com/videos/433-event-based-cache-invalidation-in-graphql) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers)