> Markdown version of [/jobs/ext/2728320-confluent-senior-software-engineer-flink-autopilot](https://www.wearedevelopers.com/jobs/ext/2728320-confluent-senior-software-engineer-flink-autopilot). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Confluent - Senior Software Engineer - Flink Autopilot - **Company:** IBM - **Location:** Armonk, NY, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Amazon Web Services, Data Analysis, Systems Engineering, Microsoft Azure, Cloud Engineering, Data Stores, Software Design Documents, Distributed Data Store, Distributed Systems, Fault Tolerance, Open Source Technology, Systems Integration, Autoscaling, System Availability, Control Structures, Kubernetes, Apache Flink, Deployment Automation, Production Code, Apache Kafka, Software Coding, Stream Processing, Confluent - **Published:** September 5, 2026 - **Apply:** https://arc.dev/remote-jobs/j/redirect/phpxybr1ei ## About the Role Master's Degree Required Technical And Professional Expertise * Deep expertise involving 6-8+ years of relevant experience in stream processing or large-scale distributed systems, and some familiarity with autoscaling, resource management, or scheduling in distributed data systems. * Strong fundamentals in distributed systems or stream processing, with a track record of independently designing, building, and shipping complex systems end to end * Demonstrated ability to implement consistency and high availability directly in code - reasoning about concurrency, failure and recovery, state, and coordination, and getting correctness right without relying on a higher-level framework to provide it * Proficiency in Java and/or Scala (the core languages of the Flink engine), with the ability to contribute production code from day one * Hands-on experience building and operating mission-critical systems in a public cloud environment (AWS, GCP, or Azure) * Ability to own a technical area, make sound trade-offs, and align a small group on direction Preferred Technical And Professional Experience * Experience implementing low-level reliability mechanisms yourself (for example, leader election, distributed coordination, reconciliation loops, consensus, or state replication) rather than only consuming them as a service * Experience with Kubernetes and operating distributed applications in production * Open source engagement: recognized, impactful technical contributions to open-source stream processing projects, particularly Apache Flink * Experience working with Go in a cloud control-plane context ## Description The Stream Processing & Analytics (SPA) team is building an elastic, reliable, durable, cost-effective, and performant stream processing engine based on Apache Flink for Confluent Cloud. Within SPA, the Flink Autopilot team owns the systems that make Flink a true cloud-native, "zero-knob" experience - autoscaling, resource management, and worry-free operations that deliver the right amount of stream processing at any given moment, so customers can focus on their use case instead of managing infrastructure. As a senior engineer on Autopilot, you'll independently drive complex engineering projects end to end within our domain, become a go-to expert for one or more areas of the system, and help the team raise the quality and operational health of the platform. Read This First: The Kind of Engineering We Do This is a low-level systems engineering role. The hard part of Autopilot is not integrating Flink, Kafka, or Kubernetes, but it's writing the code underneath them that has to be correct under concurrency, failure, and partial state. You will build control loops that make autoscaling and resource decisions, and that code must remain consistent and highly available on its own (through crashes, restarts, leader changes, and racing events), without a higher-level framework quietly handling correctness for you. When we say "scalable, fault-tolerant, distributed systems," we mean the mechanics of consistency and high availability implemented in your own code, not the ability to assemble existing platforms that provide those properties. If your strength is designing high-level services that delegate reliability to Kafka/Flink/Kubernetes, this role is probably not the right fit, and we'd rather both sides know that now than discover it mid-process. What we mean vs. what we don't mean * Consistency: We mean reasoning about and hand-writing the logic that keeps state correct across concurrent updates, retries, and failures (idempotency, ordering, reconciliation, avoiding races). We do not mean "point at a datastore that promises consistency and move on" * High availability: We mean writing code that survives process crashes, leader changes, and restarts; recovering and reconciling its own state. We do not mean "deploy multiple replicas and let the platform handle failover" * Fault tolerance: We mean anticipating partial failure in your own control loops and coordination logic and handling it explicitly. We do not mean relying on retries and health checks configured at the infrastructure layer. * Distributed systems: We mean the coordination, state, and correctness problems inside the system you're building. We do not mean operating distributed applications as a user of them Deep, effective use of Kafka, Flink, and Kubernetes is genuinely valuable here, but it is a complement to the above, not a substitute for it. What You Will Do * Drive complex projects end to end: Independently take projects from design through production within the Autopilot domain (for example: autoscaling logic, resource management, or job/task manager resilience). Break large efforts into clear milestones and deliver them with high quality * Build correctness into the code: Design and implement control loops and coordination logic that guarantee consistency and high availability directly (handling concurrency, restarts, leader changes, and partial failure explicitly, rather than delegating those guarantees to an underlying platform) * Master a domain: Develop deep expertise in one or more areas of Autopilot and become the person the team relies on for that area, representing your components in cross-functional design and architecture discussions * Articulate your design thinking: Explain the failure modes, trade-offs, and correctness arguments behind your systems through clear design documents, one-pagers, and proposals * Raise the bar through reviews and mentorship: Review PRs and designs with constructive feedback, coach junior engineers and interns, and contribute to the team's interview question pool and hiring * Communicate clearly: Demonstrate strong, succinct written and verbal communication to align a small group on technical direction and keep stakeholders informed. ## Related Videos - [Let's Get Started With Apache Kafka® for Python Developers](https://www.wearedevelopers.com/videos/565-let-s-get-started-with-apache-kafka-for-python-developers) - [What If We've Been Scaling Stream Processing Wrong All Along?](https://www.wearedevelopers.com/videos/100139-what-if-we-ve-been-scaling-stream-processing-wrong-all-along) - [Fault Tolerance and Consistency at Scale: Harnessing the Power of Distributed SQL Databases](https://www.wearedevelopers.com/videos/1520-fault-tolerance-and-consistency-at-scale-harnessing-the-power-of-distributed-sql-databases) - [Understanding Kubernetes in a visual way](https://www.wearedevelopers.com/videos/100085-understanding-kubernetes-in-a-visual-way) - [Inside Bitpanda's Tech Stack: Scaling a European Fintech Leader - Markus Dorner](https://www.wearedevelopers.com/videos/1979-inside-bitpanda-s-tech-stack-scaling-a-european-fintech-leader-markus-dorner) - [Why and when should we consider Stream Processing frameworks in our solutions](https://www.wearedevelopers.com/videos/1085-why-and-when-should-we-consider-stream-processing-frameworks-in-our-solutions) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [How Much FAANG Companies Actually Pay Software Engineers in 2025](https://www.wearedevelopers.com/magazine/230-how-much-faang-companies-actually-pay-software-engineers-in-2025) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Résumé-Driven Development: How IT trends affect the job market for software developers](https://www.wearedevelopers.com/magazine/59-resume-driven-development-how-it-trends-affect-the-job-market-for-software-developers)