AIOps Engineer
VDart, Inc.
Alpharetta, GA, United States
about 2 months ago
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Apply on www.careerjet.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
7 years minimum
Compensation
$83,200.0 - $104,000.0
Working hours
Regular working hours
Job source
Tech stack
.NET Framework
Artificial Intelligence
Amazon Web Services
Microsoft Azure
Bash Shell
Cloud Computing
Cloud Engineering
Computer Networks
Linux
DevOps
Domain Name System (DNS)
Elasticsearch
+24 more
Monitoring of Systems
Hypertext Transfer Protocols (HTTP)
Information Technology Operations
Python (Programming Language)
Machine Learning
Windows PowerShell
Reliability Engineering
Ansible
Prometheus
TCP/IP
Datadog
Google Cloud
Load Balancing
Grafana
Reliability of Systems
Infrastructure as Code (IaC)
Git
Containerization
Kubernetes
Information Technology
Terraform
Splunk
Dynatrace
Docker
Job description
We are seeking a skilled AI Ops Engineer to design, implement, and maintain AI-driven IT operations solutions. The ideal candidate will leverage automation, machine learning, and observability platforms to improve system reliability, reduce incident resolution time, and optimize infrastructure performance across cloud and on-premises environments., * Design, implement, and manage AI-powered monitoring and observability solutions.
- Develop automation scripts and workflows to streamline IT operations.
- Monitor application, server, and network performance using modern observability tools.
- Identify anomalies, predict system failures, and recommend proactive solutions.
- Perform root cause analysis (RCA) for production incidents and implement preventive measures.
- Build and maintain dashboards, alerts, and reporting systems.
- Collaborate with DevOps, Cloud, Security, and Development teams to improve system reliability.
- Implement Infrastructure as Code (IaC) using tools such as Terraform or Ansible.
- Support Kubernetes and containerized workloads.
- Optimize cloud infrastructure across AWS, Azure, or Google Cloud Platform.
- Maintain operational documentation and standard operating procedures.
- Drive continuous improvement through automation and AI-based operational insights., Join a dynamic team as a PhD Engineer specializing in Electrical, Mechanical, or Chemical disciplines, where your expertise will contribute to a high-impact customer project. This …, As a .NET Engineer, you will play a crucial role in shaping the future of AI systems by providing high-quality, real-world input that influences how models learn, reason, and perfo…
- 1 month ago +
Requirements
- Bachelor’s degree in Computer Science, Information Technology, or a related field.
- 3 7 years of experience in IT Operations, DevOps, Site Reliability Engineering (SRE), or Cloud Engineering.
- Strong knowledge of Linux and Windows administration.
- Proficiency in Python, Bash, or PowerShell scripting.
- Experience with Docker and Kubernetes.
- Hands-on experience with cloud platforms such as AWS, Azure, or Google Cloud.
- Experience with monitoring tools such as Prometheus, Grafana, Datadog, Splunk, Dynatrace, or Elastic Stack.
- Understanding of networking concepts including TCP/IP, DNS, HTTP, and load balancing.
- Familiarity with CI/CD pipelines and Git.
- Knowledge of AI/ML concepts, anomaly detection, and predictive analytics is preferred.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Apply on www.careerjet.com
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
EM
Eli McGarvie
over 3 years ago
CH
Chris Heilmann
Dev Digest 120 - Apple and peers
about 2 years ago
LM
Luis Minvielle
How to Become an AI Engineer
almost 3 years ago
IK
Igor Khokhriakov
How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again
24 days ago
CH
Chris Heilmann
Dev Digest 121 - AI goes offline
over 2 years ago
AJ
Austin Joy
What Are The Top Skills Required For Azure Developers?
over 4 years ago