Software Engineer II

Microsoft
United States
3 days ago
Apply on apply.careers.microsoft.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Starter
Experience required
2 years minimum
Compensation
$102,100.0 - $202,200.0
Working hours
Regular working hours

Tech stack

C (Programming Language) Java (Programming Language) JavaScript (Programming Language) Bioinformatics Microsoft Online Services C Sharp (Programming Language) C++ (Programming Language) Python (Programming Language) Log Files Software Engineering Supercomputing Information Technology
+1 more
Data Pipelines

Job description

Microsoft Azure High Performance Computing & AI Engineering (HPC & AI Eng) team is responsible for managing the core platform & fleet of AI & High Performance Computing products that customers use to run their most performant and demanding workloads. The AI Customer Experience (AICE) engineering team within the HPC & AI Eng. team is on the frontlines managing the flagship supercomputers and infrastructure used by top tier AI customers that enable breakthroughs such as ChatGPT and are highlighted in Top500, MLPerf and Graph500 rankings.

We run lean, obsess about customer experience and use evidence-based approach to decision making. We have live-site first, metrics-driven culture that prevents us from accumulating debt and necessity to put out fires on daily basis. You will be in a position that carries a ton of responsibility and provides opportunities to directly impact customers satisfaction.

As a Supercomputing Software Engineer on the AICE team, you will design & develop capabilities needed to monitor & efficiently operate across the infrastructure & fleet of supercomputers at scale. To enable first to know of critical incidents impacting customer capacity, you will create end to end data pipelines that process & synthesize large volume of telemetry, log files and other data sources to create actionable alerts.

Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond. Responsibilities

  • Contribute to improving key metrics such as Job Mean Time to Interrupt, Nodes in Service, Mean Time to Resolve on flagship supercomputers.
  • Manages operations of supercomputers by responding quickly to mitigate issues.
  • Implements systemic solutions and mitigations to more complex issues impacting performance or functionality of supercomputers
  • Reviews and writes incident postmortem and presents insights that drive changes to reduce or eliminate incidents.
  • Independently improves troubleshooting guides (TSGs), wikis, tests, and telemetry, adding comprehensive observability and monitoring capabilities.
  • Proactively seeks new knowledge and adapts to new trends, technical solutions, and patterns that will improve the availability, reliability, efficiency, observability, and performance of supercomputers while also driving consistency in monitoring and operations at scale.

Requirements

  • Bachelor’s Degree in Computer Science or related technical field AND 2+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
  • OR equivalent experience.

Other Requirements:

  • Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings:
  • Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter., * Bachelor’s Degree in Computer Science
  • OR related technical field AND 4+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, OR Python
  • OR Master’s Degree in Computer Science or related technical field AND 2+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
  • OR equivalent experience.

Software Engineering IC3 - The typical base pay range for this role across the U.S. is USD $102,100 - $202,200 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $133,800 - $219,200 per year.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on apply.careers.microsoft.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:11 min

Enhancing manual debugging through model-assisted log analysis

Michael Niebisch Michael Niebisch · World Congress 2024

2:04 min

Insights on transitioning from supercomputing to technical education

Andrew Holway · LIVE

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

3:41 min

Understanding the daily challenges of massive log volumes

Michal Bojko Michal Bojko · World Congress 2025

4:18 min

Prioritizing communication and structural awareness over strict tool mastery

Liam Hurrel +1 · World Congress 2021

3:52 min

Bridging Fortran and Python in modern computing

Stephen Jones · Coffee With Developers

Videos

See all

Related articles

See all