Principal Software Engineer - AI Infra Compute

Oracle
Richmond, VA, United States
about 1 month ago

Role details

Contract type
Permanent contract
Employment type
Part-time (≤ 32 hours)
Experience level
Expert
Compensation
$114,600.0 - $234,600.0
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Apple Mac Systems Microsoft Azure Bash Shell Cloud Computing Databases Data Governance Data Warehousing Software Debugging
+30 more
Linux Perl (Programming Language) InfiniBand Python (Programming Language) Machine Learning Ruby Swagger Software Engineering Systems Integration TypeScript Openapi AI Infrastructure Scripting Graphics Processing Unit (GPU) Google Cloud Application Enhancement Tool High Performance Computing Chatbots Software Security Event Driven Architecture Containerization Data Analytics Apache Kafka Data Management Api Design Api Gateway Restful APIs Stream Processing Oracle Cloud Infrastructure Docker

Job description

OCI (Oracle Cloud Infrastructure) AI Infrastructure is at the forefront of building a cutting-edge, ultra-high-performance GPU platform designed to support AI/ML/HPC workloads. This is your chance to be part of the AI revolution, creating systems that allow customers to scale from tens to thousands of GPUs without compromising performance.

Our team is the GPU Availability and Monitoring team in the Compute Org. we are responsible for designing and developing architectural changes for GPU delivery, health monitoring, triage automation, and diagnostic services. These are essential for running distributed AI/ML/HPC workloads across thousands of GPUs, leveraging technologies like RoCE and Infiniband.

We’re excited to meet a talented Senior Software Engineer like you, who shares our enthusiasm for innovation and excellence. As a valued member of our software engineering division, you’ll have the opportunity to shape the future of our technology stack and drive meaningful change in the world of cloud infrastructure and automation.

In this role, you’ll design, develop, troubleshoot, and debug software programs for databases, applications, tools, networks, and other cloud infrastructure components. Your expertise in AI and ML will help us stay ahead of the curve, and your passion for collaboration will inspire our team to achieve greatness together.

If you’re ready to unleash your full potential and make a lasting impact, we’d love to hear from you!

Responsibilities

Responsibilities

  • Work independently in ambiguous situations to ensure adherence to published standards and practices.
  • Design, develop, troubleshoot, and debug software programs for various cloud infrastructure components, including databases, applications, tools, and networks.
  • Take an active role in defining and evolving standard practices and procedures for software engineering, with a focus on AI-driven development.
  • Define and develop software for tasks associated with developing, designing, and debugging software applications or operating systems, leveraging AI and ML techniques.
  • Lead the development of critical initiatives, including:
  • Design and implement spike detection mechanisms for provisioning failures to minimize operational disruptions using ML algorithms.
  • Expand integrations with Kafka to enable near real-time actions supporting 1-Day SLO objectives for hardware repairs, utilizing event-driven architecture and stream processing.
  • Developing an automated ticket routing framework to streamline workflows, enhance efficiency, and reduce operational overhead, powered by NLP and ML.
  • Accelerate dedicated initiatives through collaborative efforts with cross-functional teams and customers, applying AI-driven insights and recommendations.
  • Harness the power of AI and ML to create innovative tools and frameworks that automate testing, simulate complex environments, and reproduce incidents, freeing up human ingenuity to focus on higher-value tasks and amplifying our ability to deliver exceptional customer experiences.
  • Collaborate and lead technical discussions across multiple teams to ensure seamless integrations and effective problem-solving.
  • Provide direction and mentoring to junior engineers, sharing knowledge and expertise to promote growth and development.

Requirements

  • Programming languages: Python, Java, TypeScript
  • Development methodologies: Agile Principles
  • Data management: data modeling, data warehousing, data governance
  • Cloud infrastructure: OCI, AWS, Azure, Google Cloud Platform (GCP)
  • Operating Systems: Linux, MacOS
  • Scripting languages: Bash, Perl, Ruby
  • Familiarity with containerization technologies such as Docker
  • API design and development: RESTful APIs, API gateways, API security
  • Familiarity with API documentation tools such as Swagger/OpenAPI
  • Experience with AI-powered tools and platforms: chatbots, virtual assistants, predictive analytics

Benefits & conditions

US: Hiring Range in USD from: $114,600 to $234,600 per annum. May be eligible for bonus, equity, and compensation deferral.

Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle’s differing products, industries and lines of business.

Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.

Oracle US offers a comprehensive benefits package which includes the following:

  1. Medical, dental, and vision insurance, including expert medical opinion
  2. Short term disability and long term disability
  3. Life insurance and AD&D
  4. Supplemental life insurance (Employee/Spouse/Child)
  5. Health care and dependent care Flexible Spending Accounts
  6. Pre-tax commuter and parking benefits
  7. 401(k) Savings and Investment Plan with company match
  8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
  9. 11 paid holidays
  10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
  11. Paid parental leave
  12. Adoption assistance
  13. Employee Stock Purchase Plan
  14. Financial planning and group legal
  15. Voluntary benefits including auto, homeowner and pet insurance

The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.

Career Level - IC4

About the company

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.

True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on dejobs.org

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

50 sec

Why developer happiness matters in web frameworks

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:07 min

Inspecting default bridge architectures and custom Docker networks

Oliver Seitz Oliver Seitz · World Congress 2025

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

3:30 min

Falling in love with Ruby and creating Basecamp

David Heinemeier Hansson David Heinemeier Hansson +1 · Coffee With Developers

2:14 min

Exploring internal AI product initiatives and global engineering roles

Maria Apazoglou · Coffee With Developers

Videos

See all

Related articles

See all