> Markdown version of [/jobs/ext/2732382-machine-learning-infrastructure-engineer](https://www.wearedevelopers.com/jobs/ext/2732382-machine-learning-infrastructure-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Machine Learning Infrastructure Engineer - **Company:** Echo Neurotechnologies Corporation - **Location:** San Francisco, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Training Data, Java (Programming Language), Artificial Intelligence, Data Analysis, Artificial Neural Networks, Big Data, C++ (Programming Language), Cloud Computing, Cloud Engineering, Computer Clusters, Nvidia CUDA, Databases, Computer Engineering, Information Engineering, Data Files, Extract Transform Load (ETL), Data Security, Data Systems, Distributed Systems, Fault Tolerance, Python (Programming Language), Machine Learning, Metadata, Performance Tuning, Systems Development Life Cycle, Azure Machine Learning, Software Engineering, Time Tracking Software, Rust (Programming Language), Scripting, Graphics Processing Unit (GPU), Pytorch, Build Management, Kubernetes, Information Technology, Data Lineage, Data Management, Machine Learning Operations, Data Pipelines, Docker, Golang - **Published:** September 5, 2026 - **Apply:** https://www.careerbuilder.com/job-details/ml-infrastructure-engineer-san-francisco-ca--0e6945cf-68ef-4ce8-801e-cdbd3b4cd428 ## About the Role * Bachelor's degree in Computer Science, Electrical Engineering, or a related technical discipline * 5+ years of industry experience in software engineering, large-scale data infrastructure, or systems ML * Extensive proficiency in Python * Familiarity with PyTorch * Experience designing, building, and maintaining high-throughput data pipelines for large and diverse datasets * Experience working with distributed-training frameworks (e.g. FSDP, DeepSpeed, Megatron-LM, Ray, etc.) * Experience building or optimizing ML training pipelines for transformers or other large neural-network models * Demonstrated ability to partner closely with research and modeling teams to productionize workflows * Excellent communication and collaboration skills to work effectively on cross-functional and interdisciplinary teams * Experience having technical ownership over at least one successfully implemented collaborative project, * Advanced degree (MS or PhD) in Computer Science, Electrical Engineering, or a related technical discipline * Proficiency in C++, Go, CUDA, Rust, and/or Java * Experience in data engineering and systems ML for time-series data * Deep understanding of the fundamentals of distributed systems, including scalability, fault tolerance, monitoring, observability, scheduling, performance tuning, and resource management * Experience with cloud-native environments and orchestration (Kubernetes, Docker, etc.) * Experience scaling foundation-model training infrastructure or multi-cluster computing environments, Artificial Intelligence (AI), Best Practices, C++ Programming Language, CUDA (Compute Unified Device Architecture), Cloud Architecture, Cloud Computing, Communication Skills, Computer Engineering, Computer Science, Computer Systems, Cross-Functional, Data Analysis, Data Formats, Data Management, Data Modeling, Data Partitioning, Data Sets, Distributed Computing, Docker, Documentation, Ecosystems, Electrical Engineering, GPU (Graphics Processing Unit), Go Programming Language (Golang), High Throughput, Java, Machine Learning, Metadata, Neural Networks, Neurology, Performance Tuning/Optimization, Process Improvement, Python Programming/Scripting Language, Realtime Communications, Research & Development (R&D), Resource Management, Rust Programming Language, Scalable System Development, Software Engineering, Startup, Systems Scalability, Team Player, Time Tracking, Training Data Sets ## Description We are seeking a Senior Machine Learning Infrastructure Engineer to join our team. The person who fills this role will design, build, and scale infrastructure to power massive-scale data, modeling, and analysis platforms, playing a critical role in shaping a high-performance, production-grade ML ecosystem to support rapid experimentation with diverse datasets spanning neural signals, behavior, and more. This person will have significant ownership over the ML R&D platform, working closely with domain experts to architect new cloud infrastructure, data pipelines, and modeling flows. The work will ultimately enable the development of cutting-edge models for neuroscientific discovery and neural decoding, empowering brain-computer interface technology to improve the lives of patients living with severe neurological conditions. PRIMARY ROLE AND RESPONSIBILITIES * Create flexible and performant ML infrastructure + Design and build systems ML cloud infrastructure to enable massive-scale modeling and analytics + Support diverse model exploration, hyperparameter optimization, pretraining, fine-tuning, and evaluation processes + Design and optimize scalable distributed training pipelines, with support for features such model sharding, cross-GPU communication, and real-time training monitoring + Create, operate, and maintain robust ML platforms and services across the model lifecycle + Make informed architecture decisions that balance performance, cost, reliability, and scalability * Build diverse and scalable data platforms + Design, build, and optimize massive-scale databases and data pipelines for scalable, flexible, and reliable data access + Explore research-driven, tailored data solutions using existing and simulated data, comparing performance and efficiency across solutions for typical data-access patterns + Create infrastructure and pipelines for ingesting internal and external datasets with varied shapes, formats, and associated metadata + Design and assess custom data formats for efficient storage and slicing of high-dimensional time-series data + Enable efficient data movement, preprocessing, and artifact management for data lineage and modeling reproducibility * Meet company standards for delivered solutions + Establish best practices for reliability, observability, reproducibility, and operational excellence across the ML ecosystem + Make informed and collaborative decisions with domain experts across the software & ML teams + Foster visibility and reproducibility within the company by maintaining extensive documentation of design decisions, evaluations of viable alternatives for selected solutions, pipeline assessments, etc. + Support ML R&D operations while preparing for eventual incorporation into product pipelines ## Related Videos - [Introduction to Azure Machine Learning](https://www.wearedevelopers.com/videos/368-introduction-to-azure-machine-learning) - [Photonic Computing: Programming a New Class of AI Accelerators (incl. Live Coding)](https://www.wearedevelopers.com/videos/100196-photonic-computing-programming-a-new-class-of-ai-accelerators-incl-live-coding) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Go with the Flow: Stop the Leaks Before Your Memory's a Waterfall!](https://www.wearedevelopers.com/videos/100073-go-with-the-flow-stop-the-leaks-before-your-memory-s-a-waterfall) - [Geometric deep learning for drug discovery](https://www.wearedevelopers.com/videos/264-geometric-deep-learning-for-drug-discovery) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) ## Related Articles - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again)