Senior Software Engineer, Gemini Audio, DeepMind
Role details
Job location
Tech stack
Job description
- Design, build, and maintain training infrastructure to support Gemini Audio encoder and pretraining models, focusing on Accelerated Linear Algebra (XLA) optimization for training on Tensor Processing Units (TPUs).
- Implement tools to analyze and track audio model training efficiency and health.
- Contribute and work closely with Gemini to build and maintain audio related components.
- Collaborate with research teams across Gemini Audio to improve the model training and inference efficiency.
Google is proud to be an equal opportunity workplace and is an affirmative action employer. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements. See also Google's EEO Policy and EEO is the Law. If you have a disability or special need that requires accommodation, please let us know by completing our Accommodations for Applicants form.
Requirements
- Bachelor's degree or equivalent practical experience.
- 5 years of experience with software development in Python.
- 3 years of experience testing, maintaining, or launching software products, and 1 year of experience with software design and architecture.
- Experience with compiler optimization, code generation, and runtime systems for popular accelerators, including GPU or TPU.
- Experience with Machine Learning, Machine Learning Optimization, Performance Optimization, and Large Language Models., * Master's degree or PhD in Computer Science or related technical field.
- Experience with TPU architectures, memory hierarchies, performance bottlenecks, and Accelerated Linear Algebra.
- Experience tailoring algorithms and ML models to exploit TPU architecture strengths and minimize weaknesses.
- Experience with various stack layers: compilers (OpenXLA, MLIR), serving libraries/frameworks (vLLM, sglang) and ML frameworks (JAX, PyTorch).
- Debugging experience to improve performance of single-mode or multi-mode (distributed) systems.
Benefits & conditions
(part of Google) 5.05.0 out of 5 stars New York, NY $174,000 - $253,000 a year - Full-time, Members of the team work on research and develop audio representations that best captures both the semantic and acoustic information within the audio signal to enable Gemini models to understand audio inputs and also generates natural audio sounds. This builds the foundation for audio-to-audio dialog models, text-to-speech, speech-to-speech translation, duplex audio dialogs.
Artificial intelligence will be one of humanity's most transformative inventions. At Google DeepMind, we are a pioneering AI lab with exceptional interdisciplinary teams focused on advancing AI development to solve complex global challenges and accelerate high-quality product innovation for billions of users. We use our technologies for widespread public benefit and scientific discovery, ensuring safety and ethics are always our highest priority. We are pushing the boundaries across multiple domains. Our global teams offer diverse learning opportunities and varied career pathways for those driven to achieve exceptional results through collective effort. Individual pay is determined by factors including job-related skills, experience, and relevant education or training.
US: $174000 - $253000 (USD) + 15% bonus target + equity + benefits
Learn more about benefits at Google.