> Markdown version of [/jobs/ext/1365741-director-core-infrastructure-engineering](https://www.wearedevelopers.com/jobs/ext/1365741-director-core-infrastructure-engineering). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Director, Core Infrastructure Engineering - **Company:** Oracle - **Location:** Sacramento, CA, United States - **Experience:** Expert - **Salary:** $121,500.0 - $306,400.0 - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Artificial Intelligence, Amazon Web Services, Microsoft Azure, Ubuntu (Operating System), CentOS, Cloud Computing, Computer Networks, Debian Linux, Linux, Python (Programming Language), Oracle (Applications), Package Management Systems, Performance Tuning, Red Hat Enterprise Linux, Ansible, Shell Script, Software Engineering, Oracle Linux, High Performance Computing, Containerization, Kubernetes, Slurm, Machine Learning Operations, Terraform, Oracle Cloud Infrastructure, Docker, Programming Languages - **Published:** July 21, 2026 - **Apply:** https://www.juju.com/job/00000000gi60y2 ## About the Role 7+ years of senior software engineering leadership or related experience. Strong communication and collaboration skills, with the ability to work effectively in cross-functional teams and convey technical concepts to non-technical stakeholders. Demonstrated leadership and people management skills. Proven at building and managing distributed/cloud software engineering solutions. Demonstrated ability to: + Develop short, medium, and long-term plans to achieve strategic objectives. + Interact across functional areas with senior management or executives to ensure unit objectives are met. + Influence thinking and gain acceptance from others in sensitive situations. BS or MS degree or equivalent experience relevant to functional area. Experience using tools like Ansible, Terraform, Python, containerization technologies (e.g., Docker, Kubernetes) and orchestration tools Solid understanding of networking concepts, security principles, and best practices. Excellent problem-solving skills, with the ability to troubleshoot complex issues and drive resolution in a fast-paced environment. Strong Linux skills with hands-on experience in Oracle Linux/RHEL/CentOS, Ubuntu, and Debian distributions, including system administration, package management, shell scripting, and performance optimization. Strong proficiency in at least one of the programming languages such as Python, Rust, Go, Java, or Scala Proven experience designing, implementing, and managing infrastructure for AI/ML or HPC workloads. ## Description We are seeking an experienced Core Infrastructure Engineering Leader to lead a high-performing engineering team responsible for delivering healthy infrastructure with optimal performance configuration. This leader will drive end-to-end technical customer execution including; POCs, technical troubleshooting, post-sale customer management/support, build automated solutions for provisioning, configuring, and monitoring AI/ML infrastructure to streamline operations and enhance productivity and optimize infrastructure performance., Lead, mentor, and develop a team of Core Infrastructure Engineers responsible for designing, implementing, and maintaining the infrastructure that supports our largest GPU/AI/ML customers. Drive the design, development, testing, validation, and deployment readiness of our automated GPU Cluster deployment tool (like AWS Parallel Cluster, Azure Cycle Cloud) with Slurm and/or Oracle Kubernetes Engine (OKE) to streamline operations and enhance productivity. Build collaborative relationships with OCI Services team, customer and sales team to deliver reliable, scalable infrastructures. Act as a technical liaison between customers, core engineering teams, and support. Work with OCI Strategic customers to grow our business in pre/post sales stages in a technical infra expert role. Take ownership of problems and work to identify solutions. Ability to think through the solution and identify/document potential issues impacting your customers. Optimize infrastructure performance by tuning parameters, optimizing resource utilization, and implementing caching and data pre-processing techniques. Troubleshoot infrastructure performance, scalability, and reliability issues and implement solutions to mitigate risks and minimize downtime. Document infrastructure designs, configurations, and procedures to facilitate knowledge sharing and ensure maintainability. As a trusted customer advocate, you will help customers/partners understand best practices around advanced GPU solutions, and how to migrate their workloads to the cloud. Educate customers of all sizes on the value proposition of Oracle Cloud and participate in deep architectural discussions to ensure solutions are designed for successful deployment in the cloud., Oversee and guide multiple teams on managing complex projects or initiatives, monitoring timelines, deliverables, and budgets when applicable to ensure strategic objectives are met. Serves as a role model for appropriately delegating work, setting priorities, and ensuring alignment with business needs. Coaches others on adjusting resources or project timelines in anticipation of business changes. Collaboration & Partnership: Role models leading cross-functional collaborative efforts to ensure alignment of expectations and strategic objectives. Empowers team to build and maintain partnerships with business leaders, stakeholders, and/or customers to address barriers and contribute to organizational success. Drives transparency and inclusivity by modeling actively seeking, listening to, and leveraging diverse perspectives. Problem Solving: Shares problem-solving strategies across teams, providing oversight on complex operational and/or technical issues, as needed. Coaches teams on analyzing highly complex data and/or information to identify solutions to ambiguous issues and provides direction on identifying root causes to prevent recurrence of issues. Continuous Learning: Pursues strategic learning opportunities to maintain expertise and apply best practices at the organizational level. Creates opportunities for team members and leaders to build their expertise in new areas, coaching them to build innovative skills. Identifies skill gap trends across the organization, and upholds a culture that places significant emphasis on sharing knowledge and pursuing learning opportunities that advance the organization. Evaluates efficiency of learning strategies and recommends adjustments as needed. Continuous Improvement: Empowers team to own the development and implementation of ideas that increase the efficiency and effectiveness of processes, protocols, and workflows across the department. Coaches teams to gain buy-in for ideas and to seek feedback on approaches and methods for continued improvement. Prioritizes and reviews the roadmap of improvement initiatives to ensure alignment with strategic direction and maximize return on investments. Performance and Development: Serves as a role model for driving performance across teams through tailored feedback and coaching in alignment with performance management processes, guidelines, and expectations. Drives consistency in the application of talent development procedures and socializes performance expectations across the organization. Ensures that individual development goals are aligned with organizational strategic initiatives. Collaborates with HR to implement talent strategy through hiring and promotion processes. Disclaimer: Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements. ## Related Videos - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Running Secure Life Science Research at Scale using Hybrid GPU HPC and Kubernetes 🧬](https://www.wearedevelopers.com/videos/100355-running-secure-life-science-research-at-scale-using-hybrid-gpu-hpc-and-kubernetes) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Docker exec without Docker](https://www.wearedevelopers.com/videos/1094-docker-exec-without-docker) - [Docker build without Docker](https://www.wearedevelopers.com/videos/100114-docker-build-without-docker) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [What Are The Top Skills Required For Azure Developers?](https://www.wearedevelopers.com/magazine/77-what-are-the-top-skills-required-for-azure-developers) - [7 Cloud Computing Trends Coming in 2025 for Developers](https://www.wearedevelopers.com/magazine/412-7-cloud-computing-trends-coming-in-2025-for-developers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [Stephan Gillich - Bringing AI Everywhere](https://www.wearedevelopers.com/magazine/489-stephan-gillich-bringing-ai-everywhere) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer)