Site Reliability Engineer I- Operations
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+37 more
Job description
Join our team as a Site Reliability Engineer I - Operations and play a critical role in maintaining the reliability, security, and performance of the technology that powers our organization. In this position, youāll work with modern infrastructure, automation tools, and cloud technologies to design and improve resilient systems, troubleshoot complex technical challenges, and help minimize downtime across enterprise platforms. Youāll collaborate with cross-functional teams, contribute to continuous improvement initiatives, and leverage industry-leading tools such as Jira, Confluence, Opsgenie, and CI/CD pipelines to enhance operational efficiency. If you enjoy solving technical challenges, building reliable systems, and making a measurable impact through technology, this role offers an excellent opportunity to develop your expertise while supporting mission-critical services., * Under close supervision, epic plans and executes projects related to the three pillars of IT operations (operational processes, change incident problem, and Ops readiness. Assists in the execution of monitoring systems and alert configurations so that Operations knows about outages before users.
- Collaborates with leadership on the creation, facilitation, and integration of documentation, including installation steps, standard operating procedures, incident runbooks, and disaster recovery documentation into a curated change/incident/problem management library. Assists Network, Application, database, and systems administrators with the enforcement of standard procedures, acts as a remote hand within a secure data center, and maintains all required supplies and tooling for the deployment of physical enterprise equipment.
- As an incident commander, participates in business-hour on-call rotation, evaluating incoming alerts for validity and dispatching the appropriate SME to resolve issues. Executes public communications in accordance with Operational standard procedures, informing stakeholders of possible service disruptions. Maintains the integrity of Runbooks.
- Performs other job-related duties as assigned.
Requirements
Graduation from an accredited college or university with an associateās degree and two yearsā experience OR any combination of education and experience totaling four years., * Knowledge of Linux and Windows Operating systems, TCP/IP fundamentals, firewall management, and anti-virus software.
- Knowledge of best practices for securing operating systems, data center maintenance, and network setup.
- Knowledge of various Monitoring solutions such as Prometheus, PRTG, Site24x7, TestCafe, Selenium, Splunk, NewRelic, Azure Monitor, and AWS CloudWatch.
- Knowledge of storage technologies such as SAN or NAS.
- Knowledge of Azure Active Directory, Active Directory, and LDAP.
- Knowledge of load balancing, clustering, and enterprise server architecture.
- Knowledge of Relational Database principles and databases/languages such as PL/SQL, MySQL, SQL Server, Oracle, Microsoft SQL, or MS Access.
- Knowledge of the Atlassian Suite, including Jira, Confluence, Status Page, and Opsgenie.
- Knowledge of Scrum/Agile principles as applicable to a DevOps Team., * Communicate effectively in normal and high-pressure situations verbally and through written mediums.
- Perform basic server, system, and application procedures such as managing user access, performing maintenance, and troubleshooting.
- Skills in troubleshooting hardware and software problems and researching technical issues.
- Experience using basic CLI tools in Windows and Linux operating systems to troubleshoot and gather information.
- Skills in customer service and interpersonal communication, both verbally and written.
- Basic scripting and programming skills in languages such as Python, JavaScript, JSON, SQL, Bash, TestCafe, and Selenium.
- Experience with instant communication and team collaboration platforms like MS Teams, Slack, or Jitsi.
- Skills in working in an ITSM solution such as Jira, ServiceNow, Asana.
Abilities
- Ability to identify, research, troubleshoot, and implement solutions for hardware and software problems.
- Ability to work in a customer service, team-oriented, collaborative, Scrum/Agile environment.
- Highly self motivated with the ability to learn quickly and accept feedback from peers.
- Ability to learn the implement process, and maintenance procedures for new technologies, equipment, hardware, and software such as operating systems, ITSM tools, monitoring solutions, and data center management.
- Ability to act as an āon-callā incident commander for communicating outages between customers, subject matter experts, teams, and leaders.
- Ability to create proposals in visually-pleasing and user-friendly language. Ability to think critically and solve complex problems.
- Ability to perform tasks in a timely and professional manner.
Benefits & conditions
4.14.1 out of 5 stars United States $23.95 - $28.18 an hour
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.indeed.comGood distractions
Talks and stories from around this role ā technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
What Are The Top Skills Required For Azure Developers?
Is Software Engineering Over-Saturated?
The Most Popular IT Jobs on the Market
Why Upskilling And Reskilling is Important For Developers