Senior Network Operations Engineer - Nashville TN

Oracle
Springfield, IL, United States
2 days ago
Apply on www.techcareers.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Part-time (≤ 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$81,100.0 - $187,000.0
Working hours
Regular working hours

Tech stack

Border Gateway Protocol Cloud Computing Computer Programming Databases Dynamic Host Configuration Protocol Software Debugging Domain Name System (DNS) Ethernet Monitoring of Systems InfiniBand IPv4 IPv6
+29 more
Virtual Private Networks (VPN) Multi-protocol Systems Python (Programming Language) Network Architecture Network Planning and Design Open Shortest Path First (OSPF) Oracle (Applications) Remote Direct Memory Access Reliability Engineering Broadcom Ansible Software Engineering SQL Databases TCP/IP Technical Data Management Systems Network Routers Scripting Transport Layer Security Computer Network Operations High Performance Computing Computer Network Technologies Firewalls (Computer Science) Juniper Information Technology Performance Monitor Puppet Open Network Automation Platform Oracle Cloud Infrastructure Cisco

Job description

In this role, you will monitor and troubleshoot network events, collect and analyze technical data, triage and mitigate incidents, coordinate escalations, and help drive continuous operational improvement. You will work alongside experienced engineers, partner teams, and vendors to maintain and optimize the infrastructure that supports Oracle customers and services worldwide.

Our mission is to keep OCI’s global network highly available, performant, and resilient-while delivering exceptional service to customers and dependable operational support to our engineering and technical teams.

For GNOC engineers, that mission translates into a broad, fast-moving role with real operational impact. You will help centrally manage OCI’s network infrastructure, respond to and resolve complex events, and develop automated solutions that reduce recurring operational work and improve reliability at scale., NOC Operations

  • Use established procedures and operational tooling to plan, implement, and safely complete network changes.
  • Mentor, onboard, and train junior network engineers.
  • Participate in operational rotations and provide break/fix and incident-response support.
  • Identify and triage actionable incidents through monitoring systems; analyze and mitigate network events; conduct or support root-cause analysis (RCA); and coordinate follow-up actions with internal support teams and vendors.
  • Provide on-call support as required, exercising sound independent judgment in a varied and complex operational environment.
  • Participate in major incident calls and use technical and analytical skills to resolve network issues affecting Oracle customers and services.
  • Manage fault detection, response, and escalation for OCI systems and networks, collaborating with third-party suppliers through resolution.

Leadership

  • Collaborate with GNOC Shift Leads and management to ensure the efficient and timely completion of daily GNOC responsibilities.
  • Lead, contribute to, and participate in the identification, development, and evaluation of projects and tools that improve GNOC effectiveness.
  • Drive runbook audits and updates to maintain compliance and align operational processes with partner service teams.
  • Conduct interviews and participate in hiring junior-level engineers.
  • Lead and/or represent the GNOC in vendor meetings, service reviews, and governance boards.
  • Automation and Scripting
  • Collaborate with network automation teams to integrate and improve operational support tooling.
  • Develop scripts and automation to reduce manual effort and improve the reliability of routine operational tasks.
  • Preferred experience with Python, Puppet, SQL, Ansible, network automation, and databases.
  • Project Delivery
  • Lead technical initiatives, including the development and improvement of runbooks, methods of procedure (MOPs), operational processes, and team onboarding materials.
  • Support the implementation of short-, medium-, and long-term plans to achieve project objectives.
  • Regularly engage senior management and network leadership to ensure team priorities and project objectives are met., Capacity Ingestion and Management:

-Takes proactive steps to design and architect infrastructure and/or service according to terms for reliability and functionality.

-Forecasts demands for infrastructure and responds to capacity needs, ensuring systems have sufficient resources to handle current and future workloads.

-Collaborates with the software development team to develop infrastructures and features that are reliable and scalable according to deployment requirements.

-Independently identifies opportunities for and drives prototyping (e.g., testing new applications or infrastructures, assisting in onboarding).

Incident and Service Lifecycle Management:

-Performs data collection, triage, technical analysis, and redirection to maintain and optimize operations and infrastructure reliability.

-Independently monitors services, maintains up-to-date knowledge of their performance, and documents their condition.

-Leverages comprehensive knowledge to perform incident response, root cause analyses, and/or maintenance on assigned services (e.g., software installs, version upgrades, security updates, backup and recovery).

-Provides health and performance reporting and takes appropriate actions based on trends in data.

-May independently perform provisioning to support infrastructure, applications, and services.

-May perform standard and non-standard decommissioning (e.g., shutting down servers, removing data from databases) to remove objects that are no longer needed.

Automation:

-Identifies opportunities for automation and assesses potential benefits.

-Develops automation tools or scripts to provide solutions, gather metrics, monitor, analyze, mitigate, or remediate issues/defects within infrastructures.

-Independently conducts testing to ensure automation performs the task correctly and produces expected results.

Technical Communication and Guidance

-Communicates the scale, capacity, security, performance attributes, and requirements of services and technology within and sometimes beyond immediate team.

-Identifies and explains the potential impact of infrastructure, feature, and tool changes, considering their impact on team operations.

Troubleshooting and Resolution:

-Provides operational support for technology, escalating incidents and other standard and non-standard issues arising within Oracle services.

-Participates in on-call shifts to address issues.

-Resolves technical issues spanning various services, investigating and debugging products in order to reach SLOs (service level objectives).

-Documents incidents and performs root cause analyses according to standard reporting methods.

-Independently performs post-mortem procedures to prevent incident reoccurrence.

Innovation and Improvement:

-Experiments with new tools and technologies to assess their potential impact on and improve infrastructure performance and reliability, ensuring adherence to security standards.

-Independently identifies and executes improvements for performance bottlenecks and deployments to ensure efficient resource usage, speed, and scalability.

-Develops knowledge of site reliability trends and shares new information with team members, management, and beyond to help others build, test, deploy and run services.

-Performs standard and non-standard analyses and provides clear data on production to contribute to business development decisions (e.g., design changes).

Core Responsibilities

Planning & Execution:

Independently manages work, monitoring timelines and deliverables to ensure projects or initiatives stay on track and meet requirements. Proactively prioritizes work and adapts to resource or timeline shifts, suggesting adjustments to maintain project efficiency.

Collaboration & Partnership:

Collaborates across teams to align on expectations and achieve shared objectives. Builds and maintains a comprehensive understanding of business, stakeholder, and/or customer needs to build and support effective partnerships. Actively listens to diverse perspectives and asks questions to ensure understanding of others.

Problem Solving:

Independently identifies and addresses standard and non-standard issues in accordance with standard practices, escalating more complex issues as appropriate. Analyzes data and/or information from multiple sources to troubleshoot standard and non-standard errors. Contributes to knowledge sharing and best practices.

Continuous Learning:

Embraces continuous learning by actively seeking to build knowledge and new skills and/or tools and staying current with industry trends and best practices. Seeks out and leverages feedback and training to improve skills. Contributes to a culture of continuous learning and knowledge sharing with team members.

Continuous Improvement:

Develops ideas and recommends updates to increase the efficiency and effectiveness of processes, protocols, and workflows within a team. Seeks input from team members on alternative approaches and methods for improving work.

Requirements

  • Strong knowledge of networking protocols and technologies, including BGP, OSPF, IS-IS, TCP/IP, IPv4/IPv6, DNS, DHCP, MPLS, VPNs, and TLS.
  • Broad hands-on experience with at least three of the following: Juniper, Cisco, Arista, InfiniBand, firewalls, routers, switches, circuit management, and optical/network transport services.
  • Strong analytical skills, including the ability to gather, correlate, and interpret data from multiple sources.
  • Ability to diagnose, prioritize, resolve, or appropriately escalate network alerts and faults.
  • Experience in a large ISP, cloud provider, or similarly complex enterprise network environment.
  • Exposure to commodity Ethernet hardware and networking ASICs, including Broadcom and NVIDIA/Mellanox.
  • Cisco, Arista and Juniper certifications are desirable.

GPU, RDMA, and HPC

  • Experience supporting GPU and RDMA network environments is highly desirable.
  • Experience supporting high-performance computing (HPC) environments is highly desirable.
  • Experience with InfiniBand and NVIDIA networking technologies, including Spectrum, is highly desirable.

Network Design and Lifecycle Management

  • Participate in network lifecycle management, including network build, refresh, and upgrade projects.
  • Participate in network solution design and design-review activities., * Self-motivated, proactive, and able to work independently.
  • Bachelor’s degree preferred, with 3-5 years of relevant network operations or engineering experience.
  • Strong organizational, time-management, verbal, and written communication skills.
  • Comfortable managing a broad range of priorities in a fast-paced operational environment.
  • Experience with incident-response plans, processes, and strategies.
  • Experience supporting large-scale enterprise infrastructure and cloud computing environments in a 24/7 network operations setting, including willingness to work rotational shifts., 8 years of experience in software engineering, infrastructure management, or related field

OR

Bachelor’s Degree in Computer Science, Engineering, or related field AND 4 years of experience in software engineering, infrastructure management, or related field

OR

Master’s Degree in Computer Science, Engineering, or related field AND 2 year of experience in software engineering, infrastructure management, or related field.

OR

Doctorate in Computer Science, Engineering, or related field

Job Skills:

Same skills as prior level plus;

Operating Systems Demonstrated ability in or knowledge of operating systems, including installing, upgrading, and troubleshooting various operating environments.

Automation Experience:

3 years of experience in automation.

Programming Experience:

3 years of experience in programming and/or scripting., 9 years of experience in software engineering, infrastructure management, or related field

OR

Bachelor’s Degree in Computer Science, Engineering, or related field AND 5 years of experience in software engineering, infrastructure management, or related field

OR

Master’s Degree in Computer Science, Engineering, or related field AND 3 years of experience in software engineering, infrastructure management, or related field

OR

Doctorate in Computer Science, Engineering, or related field AND 1 year of experience in software engineering, infrastructure management, or related field.

Automation Experience:

5 years of experience in automation.

Programming Experience:

5 years of experience in programming and/or scripting.

Benefits & conditions

US: Hiring Range in USD from: $81,100 to $187,000 per annum. May be eligible for bonus and equity.

Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle’s differing products, industries and lines of business.

Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.

Oracle US offers a comprehensive benefits package which includes the following

  1. Medical, dental, and vision insurance, including expert medical opinion
  2. Short term disability and long term disability
  3. Life insurance and AD&D
  4. Supplemental life insurance (Employee/Spouse/Child)
  5. Health care and dependent care Flexible Spending Accounts
  6. Pre-tax commuter and parking benefits
  7. 401(k) Savings and Investment Plan with company match
  8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
  9. 11 paid holidays
  10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
  11. Paid parental leave
  12. Adoption assistance
  13. Employee Stock Purchase Plan
  14. Financial planning and group legal
  15. Voluntary benefits including auto, homeowner and pet insurance

The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.

Career Level - IC3

About the company

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.

True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.techcareers.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

1:34 min

The pros and cons of campus-wide IP authentication

Christoph Eicke Christoph Eicke · World Congress 2025

2:22 min

Introducing Skupper for application connectivity

Alex Soto Alex Soto · World Congress 2024

1:29 min

Expanding practical knowledge with community sandboxes and resources

Stuart Clark · LIVE

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

3:05 min

Exploring microcontrollers and communication protocols for amateur hardware

Philipp-Alexander Blum · LIVE

1:06 min

Developer experience and project variety at scale

Alexandra Petri · World Congress 2023

Videos

See all

Related articles

See all