Senior Software Engineer, Fleet-level ML Performance

Google LLC
Sunnyvale, CA, United States
13 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$174,000.0 - $252,000.0
Working hours
Regular working hours
Job source

Tech stack

Systems Engineering C++ (Programming Language) Cloud Computing Encodings Computer Engineering Data Centers Python (Programming Language) Software Requirements Analysis Systems Architecture AI Infrastructure Deep Learning Information Technology

Job description

  • Perform fleet-level performance analysis of key ML workloads (e.g., Gemini) on future TPU systems using advanced simulation tools to evaluate hardware/software trade-offs and guide next-generation chip architecture.
  • Partner with teams across the ML stack including model researchers, compiler developers, systems engineers, and TPU architects to analyze and optimize performance across the design space.
  • Design and implement a unified ML Accelerator Platform Performance Estimation Methodology using C++ and Python to enable scalable performance and TCO projections across Google.
  • Collaborate cross-functionally with data center, hardware architecture, and framework teams to define key hardware and software requirements for future AI infrastructure.

Requirements

Experience driving progress, solving problems, and mentoring more junior team members; deeper expertise and applied knowledge within relevant area., * Bachelor’s degree in Computer Science, Electrical Engineering, Computer Engineering, or a related field.

  • 5 years of experience in systems architecture or computer architecture, power and performance trade-off analysis, or data center, cloud, infrastructure hardware optimization .
  • Experience with Reliability, Availability, and Serviceability (RAS) features, paradigms, or architecture.

Preferred qualifications:

  • Master’s degree or PhD in Electrical Engineering, Computer Engineering or Computer Science, with an emphasis on computer architecture.
  • Knowledge of deep learning workloads, including embedding architectures and their hardware execution characteristics.

About the company

Google Cloud’s mission is to make every business successful through AI by combining cutting-edge technology, infrastructure, and talent. AI/ML software engineers in Cloud bridge the gap between pioneering models and a massive product vehicle reaching billions. Our talent density and AI-powered tools drive rapid development, rooted in a culture of empowerment and a bias to action. In this role, you aren’t just building technology; you’re shaping the frontier of enterprise and driving the evolution of advanced models.

The TPU Chip Architecture and Performance team bridges Google’s machine learning workloads and custom silicon architectures. Through hardware/software co-design, we define and shape the future TPU platforms required to meet Google’s ambitious AI goals. In this role, you will conduct end-to-end performance analysis of critical ML workloads including Gemini, YouTube, Ads, and key 3P models on next-generation TPU systems. Using advanced simulation and projection methodologies, you will evaluate hardware features and software optimizations to balance performance and cost trade-offs. You will drive high-impact architecture decisions that maximize TPU efficiency while accounting for system-wide constraints such as power, reliability (RAS), and scheduling.

The AI and Infrastructure team is redefining what’s possible. We empower Google customers with breakthrough capabilities and insights by delivering AI and Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers, and billions of Google users worldwide.

We’re the driving force behind Google’s groundbreaking innovations, empowering the development of our cutting-edge AI models, delivering unparalleled computing power to global services, and providing the essential platforms that enable developers to build the future. From software to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and much more.Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $174000 - $252000 (USD) + 15% bonus target + equity + benefits, Google is proud to be an equal opportunity and affirmative action employer. We are committed to building a workforce that is representative of the users we serve, creating a culture of belonging, and providing an equal employment opportunity regardless of race, creed, color, religion, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition (including breastfeeding), expecting or parents-to-be, criminal histories consistent with legal requirements, or any other basis protected by law. See alsoGoogle’s EEO Policy (https://www.google.com/about/careers/applications/eeo/) ,Know your rights: workplace discrimination is illegal (https://careers.google.com/jobs/dist/legal/EEOC_KnowYourRights_10_20.pdf) ,Belonging at Google (https://about.google/belonging/) , andHow we hire (https://careers.google.com/how-we-hire/) .

If you have a need that requires accommodation, please let us know by completing ourAccommodations for Applicants form (https://goo.gl/forms/aBt6Pu71i1kzpLHe2) .

Google is a global company and, in order to facilitate efficient collaboration and communication globally, English proficiency is a requirement for all roles unless stated otherwise in the job posting.

To all recruitment agencies: Google does not accept agency resumes. Please do not forward resumes to our jobs alias, Google employees, or any other organization location. Google is not responsible for any fees related to unsolicited resumes.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:17 min

Optimizing character encoding with Kim variable byte encoding

Douglas Crockford Douglas Crockford · World Congress 2024

51 sec

Repurposing hardware and operating underwater data centers

Chris Heilmann +1 · LIVE

1:25 min

Distinguishing artificial intelligence from deep learning

Sam Witteveen · Coffee With Developers

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

4:12 min

Distilling cross-encoder models into smaller efficient sentence embedding models

Marek Suppa · LIVE

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

Videos

See all

Related articles

See all