Senior AI/ML Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
- Design and build AI-powered features — LLM agents, retrieval, and event intelligence — that operate on high-volume, real-time event streams, from problem framing through production deployment and monitoring.
- Architect and own the systems behind them: agent and prompt orchestration, retrieval pipelines, tool/API integrations, and low-latency inference and evaluation at scale.
- Reason about consistency, throughput, fault tolerance, and cost across services that must stay reliable under bursty, unpredictable load.
- Take AI features from prototype to production, establishing the evaluation, guardrail, observability, and improvement loops that keep them accurate and trustworthy over time.
- Partner with platform, product, and applied-research teams to define what “good” looks like and to integrate AI cleanly into existing services.
- Raise the bar through example, reviews and mentorship, and help shape the team’s technical direction.
Requirements
What you’ll bring
- 5+ years of software engineering experience, with meaningful time spent building and operating production distributed systems (high-throughput services, streaming/event-driven architectures, or large-scale data platforms).
- Hands-on experience building and shipping AI systems in production — LLM-powered applications, agents, or retrieval — including the surrounding orchestration, serving, and evaluation, not just prototypes.
- Strong programming fundamentals and comfort moving between systems and AI/application code.
- Solid grounding in applied AI fundamentals: prompting, retrieval, agent patterns, and how to evaluate and guardrail LLM behavior.
- Experience with cloud infrastructure (AWS, GCP, or Azure), containers, and orchestration (Kubernetes).
- A pragmatic, reliability-minded mindset: you optimize for systems that work correctly at scale, and you can articulate the trade-offs behind your choices.
- Strong communication and collaboration skills, and a track record of raising the quality of the teams and systems around you.
Nice to have
- Experience with LLMOps tooling and patterns — evaluation harnesses, prompt/version management, tracing and observability for agents, and online/offline eval consistency.
- Deep experience serving LLM-based systems in production, including retrieval-augmented generation, multi-step agents, and tool use.
- Background in anomaly detection, event correlation, or applied problems in observability, AIOps, or reliability.
- Familiarity with the ecosystem — e.g. LLM APIs and frameworks such as LangChain or LlamaIndex, vector databases, and distributed data/compute tools such as Kafka, Airflow, or Spark.
- Contributions to open-source AI or distributed-systems projects.
Benefits & conditions
What we offer
As a global organization, our total rewards approach is competitive with industry standards and aligned with local laws and regulations. Learn more, including country-specific offerings, on our benefits site.
Your package may include:
- Competitive salary
- Comprehensive benefits package
- Flexible work arrangements
- Company equity*
- ESPP (Employee Stock Purchase Program)*
- Retirement or pension plan*
- Generous paid vacation time
- Paid holidays and sick leave
- Dutonian Wellness Days & HibernationDuty - companywide paid days off in addition to PTO
- Paid parental leave: 22 weeks for pregnant parent, 12 weeks for non-pregnant parent (some countries have longer leave standards and we comply with local laws)*
- Paid volunteer time off: 20 hours per year
- Company-wide hack weeks
- Mental wellness programs
*Eligibility may vary by role, region, and tenure
About the company
PagerDuty is transforming critical work for modern enterprises. The AI-powered PagerDuty Operations Cloud empowers business resilience and drives operational efficiency. With a generative AI assistant at its core, PagerDuty empowers teams to detect and resolve issues in real time, orchestrate complex workflows, and drive continuous improvement across their digital operations.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
What Are Large Language Models?
MLOps And AI Driven Development
MLOps – What’s the deal behind it?
Dev Digest 121 - AI goes offline