Principal Software Engineer, AI Infra Management and Ops

Microsoft
Redmond, WA, United States
1 day ago
Apply on www.techcareers.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
6 years minimum
Compensation
$142,800.0 - $274,800.0
Working hours
Regular working hours

Tech stack

C (Programming Language) Java (Programming Language) JavaScript (Programming Language) Artificial Intelligence Computing Platforms Microsoft Online Services C Sharp (Programming Language) C++ (Programming Language) Cloud Computing Continuous Integration Data Centers Distributed Systems
+20 more
Ethernet Network Topologies InfiniBand Python (Programming Language) Network Interface Remote Direct Memory Access Cloud Services Software Engineering Rust (Programming Language) Data Processing Graphics Processing Unit (GPU) Computer Networking Systems Cloud Platform System High Performance Computing Infrastructure as Code (IaC) Information Technology Deployment Automation Bare Metal Serverless Computing Programming Languages

Job description

Help shape the future of cloud infrastructure and artificial intelligence platforms at unprecedented scale. Our team is building the foundational platform capabilities that power large-scale distributed systems, accelerated computing environments, and next-generation cloud services. You will work at the intersection of software, infrastructure, networking, and datacenter technologies to enable reliable, secure, and scalable platforms that support a broad range of business-critical workloads. You will collaborate with engineering teams across Microsoft to advance platform capabilities, accelerate innovation, and improve operational excellence across diverse infrastructure environments.

As a Principal Software Engineer, you will lead the design and evolution of cloud-native platforms, distributed systems, and infrastructure services that support service lifecycle management, resource orchestration, observability, reliability, and automation. You will help guide technical strategy across multiple engineering investments, contribute to architecture decisions involving compute, networking, storage, and platform services, and collaborate across organizations to deliver capabilities that improve platform efficiency and developer productivity. This opportunity will allow you to deepen your expertise in distributed systems and cloud infrastructure, expand your influence across engineering organizations, and contribute to innovations in Artificial Intelligence (AI) infrastructure, graphics processing unit (GPU) platforms, high-performance networking, and large-scale datacenter systems.

Microsoft’s mission is to empower every person and every organization on the planet to achieve more. We cultivate a culture grounded in respect, integrity, accountability, inclusion, collaboration, and continuous learning, enabling individuals and teams to grow, contribute, and make meaningful impact.

Responsibilities

  • Lead the design, development, and evolution of cloud-native services, distributed systems, and platform capabilities that support large-scale production environments.
  • Collaborate across engineering organizations to help define technical direction and architecture for platform infrastructure, cloud services, networking, storage, compute, and operational excellence investments.
  • Guide the development of platform engineering capabilities that improve deployment automation, service lifecycle management, observability, reliability, scalability, and developer productivity.
  • Advance software engineering practices including Continuous Integration and Continuous Delivery (CI/CD), Infrastructure as Code (IaC), testing, security, monitoring, and operational readiness throughout the engineering lifecycle.
  • Contribute to the design and operation of Artificial Intelligence (AI) infrastructure, High Performance Computing (HPC) platforms, bare-metal infrastructure, graphics processing unit (GPU) environments, and large-scale distributed computing systems.
  • Collaborate on datacenter architecture and networking solutions, including rack-scale systems, network topology design, Ethernet fabrics, InfiniBand fabrics, Remote Direct Memory Access (RDMA), Smart Network Interface Cards (SmartNICs), Data Processing Units (DPUs), and accelerated computing platforms.
  • Foster a culture of technical excellence through architecture collaboration, mentorship, knowledge sharing, design review participation, and support for engineering excellence across the broader organization.

Requirements

  • Bachelor’s Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience., Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. This includes passing the Microsoft Cloud background check upon hire/transfer and every two years thereafter., * Recognized for delivering technical solutions that span multiple engineering teams, organizations, or business areas.
  • Demonstrated ability to simplify complex technical challenges and translate them into practical, scalable, and maintainable solutions.
  • Proven ability to influence technical direction through collaboration, architectural leadership, technical mentoring, and cross-organizational partnerships.
  • Knowledge of large-scale distributed systems, cloud-native platforms, or infrastructure services supporting business-critical production environments.
  • Familiarity with Artificial Intelligence (AI) infrastructure, High Performance Computing (HPC) environments, accelerated computing platforms, or large-scale graphics processing unit (GPU) deployments.
  • Knowledge of datacenter architecture, rack-scale infrastructure, network topology design, and modern datacenter networking technologies.
  • Familiarity with InfiniBand fabrics, NVLink, NVSwitch, Ethernet networking, Remote Direct Memory Access (RDMA), Smart Network Interface Cards (SmartNICs), or Data Processing Units (DPUs) supporting large-scale compute environments.
  • Understanding of bare-metal infrastructure, hardware lifecycle management, fleet operations, infrastructure telemetry, or platform operations at scale.
  • Proficiency in one or more modern programming languages such as Go, Rust, C#, Java, or Python.

AIINFRA

Software Engineering IC5 - The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.techcareers.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

1:57 min

Routing cross-rack traffic seamlessly with NCCL

Kevin Klues Kevin Klues · World Congress 2025

1:11 min

Running high-performance edge computing on bare metal servers

Josip Stuhli Josip Stuhli · Coffee With Developers

4:52 min

Connecting namespaces with local virtual ethernet pairs

Oliver Seitz Oliver Seitz · World Congress 2025

1:51 min

Bypassing the CPU stack with remote direct memory access

Lerna Ekmekcioglu Lerna Ekmekcioglu · Europe 2026 Virtual

3:23 min

The AI workload technology stack and its components

Lerna Ekmekcioglu Lerna Ekmekcioglu · Europe 2026 Virtual

Videos

See all

Related articles

See all