Tech Lead Manager, Staff Software Engineering, XProf

Google LLC
Sunnyvale, United States of America
yesterday

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Languages
English
Experience level
Senior
Compensation
$ 301K

Job location

Sunnyvale, United States of America

Tech stack

C++
Computer Programming
Software Debugging
Memory Management
Python
Machine Learning
Software Engineering
High Performance Computing
Information Technology

Job description

  • Be focused around driving continuous improvements to the machine learning software/hardware stacks through providing insightful performance debugging. Provide insights by summarizing different views of captured profile data such as trace timelines, memory usage, HLO profiles, ML graph summaries.
  • Learn and build an intuitive understanding of existing data collection, analysis, and visualization workflows.
  • Support new and exciting ML paradigms (such as horizontal scaling for upcoming TPU chips) by making contributions across the end to end stack and analysis tools.
  • Partner with product area leads to understand model optimization use cases, drive cross functional efforts to deliver on chip profiling requirements, and propose new hardware features.
  • Collaborate across Hardware, Driver, Runtime, and Performance Analysis teams and many other stakeholders.

Google is proud to be an equal opportunity workplace and is an affirmative action employer. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. See also Google's EEO Policy and EEO is the Law. If you have a disability or special need that requires accommodation, please let us know by completing our Accommodations for Applicants form.

Requirements

  • Bachelor's degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 5 years of experience in programming, debugging, computer architecture, and C++.
  • 3 years of experience in a technical leadership role.
  • Experience in a people management, supervision/team leadership role., * Master's degree or PhD in Engineering, Computer Science, or a related technical field.
  • Experience with Python, standard ML, high-performance computing, and embedded systems.
  • Background in machine learning, accelerator architectures (driver/run times), or related fields.

About the company

With technical and leadership expertise, you manage engineers across multiple teams and locations, a large product budget and oversee the deployment of large-scale projects across multiple sites internationally. In this role, you will have unique and exciting opportunity to work at the "heart of machine learning" and learn about basic to advanced ML paradigms for training/inference while delivering impact across different compute infrastructure, Cloud, and Open Source environments. You will focus on hardware/software interactions with accelerators (current and chips under design) and integration with the rest of the profiling system. The AI and Infrastructure team is redefining what's possible. We empower Google customers with breakthrough capabilities and insights by delivering AI and Infrastructure at unparalleled scale, efficiency, reliability and velocity. Our customers include Googlers, Google Cloud customers, and billions of Google users worldwide. We're the driving team behind Google's groundbreaking innovations, empowering the development of our cutting-edge AI models, delivering unparalleled computing power to global services, and providing the essential platforms that enable developers to build the future. From software to hardware our teams are shaping the future of world-leading hyperscale computing, with key teams working on the development of our TPUs, Vertex AI for Google Cloud, Google Global Networking, Data Center operations, systems research, and much more. Individual pay is determined by factors including job-related skills, experience, and relevant education or training. US: $207000 - $301000 (USD) + 20% bonus target + equity + benefits

Apply for this position