Software Engineer II - AI Infrastructure
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+12 more
Job description
As part of the AI Infra team, you will work on large-scale infrastructure that supports some of the world’s most demanding cloud services and AI workloads. The infrastructure you will work on powers OpenAI and OSS model hosting, large-scale inferencing, and the backend for Microsoft Copilots - across the largest capacity fleets in the industry. You will help design and build foundational systems that enable reliability, performance, and operational excellence at global scale., As a Software Engineer II, you will design, develop, and operate distributed systems that power mission-critical cloud services. You will collaborate across engineering disciplines to build resilient, scalable, and highly available platforms, leveraging strong software engineering fundamentals and data-driven decision making. You will contribute throughout the software development lifecycle, from architecture and implementation to deployment, monitoring, and continuous improvement.
This opportunity will allow you to:
-
Accelerate your technical and career growth by solving complex distributed systems challenges at cloud scale.
-
Develop deep expertise in large-scale infrastructure, reliability engineering, system architecture, and operational excellence.
-
Hone your collaboration and leadership skills while working across teams to deliver high-impact solutions that serve millions of users and workloads.
Microsoft Culture
Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees, we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day, we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond
Responsibilities
Responsibilities
-
Design, develop, test, deploy, and operate large-scale distributed systems and platform services that deliver reliable, scalable, secure, and high-performance experiences.
-
Build and enhance infrastructure capabilities, resource management services, and workload orchestration solutions that optimize system efficiency, utilization, availability, and performance.
-
Develop software and platform features that support capacity management, service scalability, policy-driven decision making, workload placement, prioritization, and operational flexibility across distributed environments.
-
Collaborate with engineers, product stakeholders, and cross-functional partners to define technical requirements, influence architecture decisions, and deliver high-quality solutions that address customer and business needs.
-
Develop, test, and maintain control plane services written in C#, hosted on Kubernetes (AKS) clusters. .
-
Analyze complex production issues, identify root causes, and implement sustainable solutions that enhance scalability, maintainability, efficiency, and service health. Provide operational support and DRI (on-call) responsibilities for the service.
-
Contribute to engineering best practices through technical design reviews, code quality initiatives, knowledge sharing, mentorship, and a culture of innovation, accountability, inclusion, and continuous learning.
Requirements
- Bachelor’s Degree in Computer Science or related technical field AND 2+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
- OR equivalent experience., * ‘Master’s Degree in Computer Science or related technical field AND 3+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
- OR Bachelor’s Degree in Computer Science or related technical field AND 5+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
- OR equivalent experience.
- OOP (Object Oriented Programming) proficiency and practical familiarity with common code design patterns
- 2+ years of experience with service development in a distributed environment, in a dev-ops role, including concurrency management and stateful resource management
- Hands-on experience with public cloud services at the IaaS level
AIINFRA
Software Engineering IC3 - The typical base pay range for this role across the U.S. is USD $102,100 - $202,200 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $133,800 - $219,200 per year.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
How to Become an AI Engineer
Top-Paying Tech Jobs (with Salaries)
How Much Does a Software Engineer Make? Realistic Software Engineering Salaries
Best Paying Jobs in Technology