Software Development Manager, Data Center - GenAI

Amazon.com, Inc.
Bellevue, WA, United States
23 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$184,900.0 - $250,200.0
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Data Analysis Code Review Continuous Integration Data Centers Distributed Systems Amazon DynamoDB Information Retrieval Knowledge-Based Systems Natural Language Processing
+24 more
Software Product Management Azure Machine Learning Search Technologies Software Engineering Software Systems Large Language Models Multi-Agent Systems Prompt Engineering Distributed Programming Session Description Protocol Security Descriptions (SDES) Generative AI AWS Lambda Web Filtering Servicebus AWS Data Analytics Build Process Front End Software Development Virtual Agents Cloudwatch Api Gateway Software Coding Software Version Control Api Management Serverless Computing

Job description

AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS global infrastructure. In other words, we’re the people who keep the cloud running. We support all AWS data centers and all of the servers, storage, networking, power, and cooling equipment that ensure our customers have continual access to the innovation they rely on. We work on the most challenging problems, with thousands of variables impacting the supply chain - and we’re looking for talented people who want to help.

You’ll join a diverse team of design engineers, quality/reliability engineers, supply chain specialists, field engineers, and other vital roles. You’ll collaborate with people across AWS to help us deliver the highest standards for quality and reliability while providing seemingly infinite capacity at the lowest possible cost for our customers. And you’ll experience an inclusive culture that welcomes bold ideas and empowers you to own them to completion.

You’ll join a team of Software Development Engineers building an agentic AI platform that serves a broad customer base of design engineers, quality/reliability engineers, supply chain specialists, field engineers, and other vital roles across AWS data center operations. You’ll collaborate with people across AWS to help us deliver the highest standards for quality and reliability while providing seemingly infinite capacity at the lowest possible cost for our customers. And you’ll experience an inclusive culture that welcomes bold ideas and empowers you to own them to completion.

As the Software Development Manager for the Data Center Agentic AI Platform team, you will lead a team of Software Development Engineers building AWS data center’s agentic GenAI platform that powers AI-assisted operations across the global data center infrastructure. You will own the technical vision and strategic roadmap for the platform, driving investments across agentic AI systems, full-stack engineering, search and knowledge systems capabilities. Your leadership will shape the direction of a next-generation AI/ML platform that orchestrates physical work processes, automates decision-making, and enhances operational efficiency for a 30K+ globally distributed user base. You will champion platform thinking building reusable primitives, APIs, and extensible components that dozens of teams across the Data Center Community build upon.

In this role, you will drive the design and delivery of production-grade agentic systems including LLM orchestration, tool-calling patterns, agent frameworks, multi-agent orchestration, and intelligent workflow automation. You will partner closely with cross-functional stakeholders including data center operations, controls engineering, product management, and peer engineering teams to translate complex operational needs into scalable AI-powered solutions. You will establish and raise the bar on engineering practices including code reviews, CI/CD, progressive deployment, observability, and operational readiness for AI systems in production. You will also own hiring strategy and talent development, building a high-performing engineering team with deep expertise in generative AI, distributed systems, and full-stack development, while communicating platform strategy, technical roadmaps, and business impact to senior leadership with clarity and conviction., * Lead and mentor a team of SDEs building and operating the Data Center Agentic AI Platform, fostering a culture of ownership, innovation, and operational excellence

  • Own the end-to-end technical roadmap for the platform, balancing investments across agentic AI capabilities, platform infrastructure, frontend experiences, search/knowledge systems workstreams
  • Drive the architecture and delivery of agentic AI systems including LLM orchestration, prompt engineering, skills, harness, tool-calling patterns, semantic search, and agent frameworks leveraging technologies such as Amazon Bedrock, AgentCore, and multi-agent orchestration patterns
  • Lead the development of full-stack serverless solutions leveraging AWS Lambda, API Gateway, DynamoDB, EventBridge, CDK, and related services to deliver scalable, production-grade platform capabilities
  • Own the design of search and knowledge systems including vector embeddings, hybrid retrieval, document processing pipelines, and semantic chunking to power the platform’s intelligent responses
  • Define and implement evaluation frameworks, guardrails, and safety mechanisms for agentic AI systems, including LLM output quality evals, agent behavior testing, content filtering, and responsible AI controls to ensure reliable and trustworthy platform behavior at scale
  • Build and evolve platform primitives and reusable components that enable dozens of teams across the Data Center Community to build AI-powered capabilities on top of the platform
  • Partner with data center operations, controls engineering, product management, and peer engineering teams to identify high-impact use cases and translate them into platform features
  • Establish and enforce engineering excellence including CI/CD pipeline design, progressive deployment, synthetic monitoring, observability (CloudWatch, X-Ray, OpenTelemetry), and operational readiness reviews
  • Own hiring, performance management, and career development for the team, building a diverse pipeline of engineers with expertise in GenAI, distributed systems, and full-stack development
  • Communicate platform strategy, project status, and business impact to senior leadership, driving alignment on priorities and resource allocation

A day in the life

  • Conduct 1:1s and team stand-ups to unblock engineers, review progress, and align priorities across agentic AI, platform infrastructure, frontend, and search/knowledge workstreams
  • Review technical designs, architecture proposals, and code reviews - ensuring high standards for agentic system design, API contracts, prompt engineering patterns, and infrastructure-as-code
  • Triage and prioritize incoming requests from cross-functional stakeholders (data center operations, controls engineering, product managers, peer platform teams) against Eva’s roadmap and strategic pillars
  • Monitor operational health of Eva’s production systems including agent orchestration services, RAG pipelines, search infrastructure, APIs, and frontend applications - driving rapid resolution of any issues
  • Collaborate with product and program managers to refine requirements, scope agentic AI features, and plan sprint/iteration deliverables
  • Participate in hiring activities including resume reviews, phone screens, on-site interviews, and calibration sessions to build and maintain a strong engineering talent pipeline
  • Engage in strategic planning discussions with senior leadership on Eva’s platform direction, GenAI technology adoption, resource allocation, and long-term technical investments
  • Coach and mentor engineers, providing career development guidance, actionable feedback, and growth opportunities in emerging areas like agentic AI and large-scale platform engineering

Requirements

  • 3+ years of engineering team management experience
  • 7+ years of working directly within engineering teams experience
  • Knowledge of engineering practices and patterns for the full software/hardware/networks development life cycle, including coding standards, code reviews, source control management, build processes, testing, certification, and livesite operations
  • Experience partnering with product or program management teams
  • 3+ years of developing large-scale, multi-tiered distributed software systems using distributed programming experience, * Experience delivering products against plan in a fast-paced, multi-disciplined, distributed-responsibility and often ambiguous environment
  • Experience in recruiting, hiring, mentoring/coaching and managing teams of Software Engineers to improve their skills, and make them more effective, product software engineers
  • Knowledge of ML, NLP, Information Retrieval and Analytics
  • Experience working with fast-moving, high-performance teams and driving innovative solutions tailored to unique business environments
  • Experience leading teams building AI/ML or generative AI systems in production, including LLM-based applications, agentic architectures, or RAG systems

Benefits & conditions

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits .

USA, WA, Bellevue - 184,900.00 - 250,200.00 USD annually

About the company

Amazon Web Services (AWS) is the world’s most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating - that’s why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses.

Diverse Experiences

Amazon values diverse experiences. Even if you do not meet all of the preferred qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn’t followed a traditional path, or includes alternative experiences, don’t let it stop you from applying.

Work/Life Balance

We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there’s nothing we can’t achieve in the cloud.

Inclusive Team Culture

Here at AWS, it’s in our nature to learn and be curious. Our employee-led affinity groups foster a culture of inclusion that empower us to be proud of our differences. Ongoing events and learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences, inspire us to never stop embracing our uniqueness.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:45 min

Executing event-driven code using AWS Lambda functions

Sebastien Stormacq Sebastien Stormacq · LIVE

3:13 min

Navigating the GenAI observability dashboard in Amazon CloudWatch

Yasemin Aktürk Yasemin Aktürk · Europe 2026 Virtual

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

49 sec

Automated distributed tracing on AWS Lambda functions

Developersteve · LIVE

2:08 min

Essential engineering roles in the generative AI space

Mary Grygleski Mary Grygleski · LIVE

55 sec

Resolving DynamoDB hot keys using in-memory caching

Irina Branovic Irina Branovic · WWC Europe 2026

Videos

See all

Related articles

See all