> Markdown version of [/jobs/ext/2717616-distributed-systemsevent-driven-architecturesgenerative-aigoion-xtpjava](https://www.wearedevelopers.com/jobs/ext/2717616-distributed-systemsevent-driven-architecturesgenerative-aigoion-xtpjava). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Distributed SystemsEvent-Driven ArchitecturesGenerative AiGoIon XtpJava - **Company:** Domino’s - **Location:** United States (Remote available) - **Experience:** Expert - **Salary:** $185,000.0 - $210,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Applicant Tracking Systems, Database Queries, Distributed Systems, Python (Programming Language), Knowledge Management, Load Testing, Software Reliability Testing, Prometheus, Systems Integration, Chatbots, Grafana, Multi-Cloud, Kubernetes, Infrastructure Automation Frameworks, Data Analytics, New Relic (SaaS), Servicenow - **Published:** September 4, 2026 - **Apply:** https://www.builtincolorado.com/job/staff-performance-quality-engineer/8537719?handler=ApplyRedirect ## About the Role * Background in SRE, platform engineering, or infrastructure with hands-on experience operating and troubleshooting distributed systems in production Kubernetes environments * Strong proficiency in Python and comfort working in a large, modular codebase that spans orchestration, infrastructure automation, and systems integration * Experience with observability stacks (Prometheus, Grafana, New Relic, or similar) - writing queries, building dashboards, and using metrics to diagnose performance and reliability issues at the systems level * Demonstrated ability to go beyond detection to resolution: profiling services, identifying resource bottlenecks, and working with engineering teams to ship durable fixes * Familiarity with performance and load testing methodologies (e.g., Locust, k6, or similar) as part of a broader infrastructure or reliability practice * Clear ownership mindset - self-directed, accountable, and able to communicate priorities and status effectively in a remote, async environment ## Description The Automation Team at Domino acts as a force multiplier for engineering, building the tools and systems that enable teams to ship code confidently and consistently. A core part of this mission is Tempest, an in-house platform that orchestrates realistic, long-duration workloads against live Kubernetes clusters and validates the results against real observability data. Today, when scale testing surfaces a bottleneck, a resource misconfiguration, or a regression in system behavior, the team can identify and report the issue - but we need someone who can take the next step: profiling services, tracing root causes through Prometheus and New Relic data, and partnering with platform engineers to drive durable fixes. Focused on iteration and continuous improvement, the team looks for targeted enhancements that create outsized impact, and this role will close the gap between detection and resolution at the infrastructure level. What your impact will be In your first year, you will: * Serve as the technical owner of Tempest, Domino's scale and reliability platform, ensuring it remains reliable, extensible, and aligned with evolving infrastructure needs * Diagnose and drive resolution of performance bottlenecks and resource misconfigurations surfaced by scale testing - working directly with platform and infrastructure teams to ship fixes, not just file tickets * Deliver accurate, data-driven sizing recommendations for customer-facing documentation based on rigorous empirical testing across deployment sizes * Strengthen observability across scale testing by improving Prometheus and New Relic instrumentation, making it faster to pinpoint root causes during and after multi-day load runs * Establish and operationalize scale testing on cloud platforms, ensuring appropriate sizing and configuration guidance for this increasingly divergent product line * Partner with platform teams to enable effective scale and reliability testing across additional cloud providers, helping position Domino for future multi-cloud success * Increase the efficiency and leverage of a small team by building infrastructure automation that scales operationally as the product and customer base grow, Leads architecture and hands-on engineering across a futures commission merchant platform, including market data, order routing, settlement, reporting, risk, and liquidation. Drives replatforming away from legacy vendors, improves reliability and technical quality, partners with product, risk, operations, finance, and compliance, and mentors engineers. Requires extensive backend distributed-systems experience, direct FCM or derivatives expertise, regulated-finance experience, and proficiency in Go, Java, or similar technologies., Artificial Intelligence * Machine Learning * Natural Language Processing * Software * Conversational AI Lead and coach a recruiting team while owning day-to-day talent acquisition execution across engineering, research, and go-to-market functions. Run full-cycle recruiting, support complex searches, build AI-assisted workflows and scalable recruiting systems, use funnel data to diagnose performance, develop recruiters, and improve hiring quality through structured interviewing and stakeholder coaching. Top Skills: AIAshbyJuiceboxMetaview PNC Bank ## Related Videos - [Applying Agile Principles to Incident Management ](https://www.wearedevelopers.com/videos/101-applying-agile-principles-to-incident-management) - [Chatbots are going to destroy infrastructures and your cloud bills](https://www.wearedevelopers.com/videos/1130-chatbots-are-going-to-destroy-infrastructures-and-your-cloud-bills) - [5 steps for running a Kubernetes environment at scale](https://www.wearedevelopers.com/videos/88-5-steps-for-running-a-kubernetes-environment-at-scale) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [AI in Production: applied AI & enterprise use cases](https://www.wearedevelopers.com/videos/100130-ai-in-production-applied-ai-enterprise-use-cases) - [Testing AI Agents: Automated Evaluation for Chatbots & RAG Systems](https://www.wearedevelopers.com/videos/100300-testing-ai-agents-automated-evaluation-for-chatbots-rag-systems) ## Related Articles - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [WeAreDevelopers Dev Digest Issue 116 - The new search wars…](https://www.wearedevelopers.com/magazine/445-wearedevelopers-dev-digest-issue-116-the-new-search-wars) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift)