> Markdown version of [/jobs/ext/2730480-devops-engineer](https://www.wearedevelopers.com/jobs/ext/2730480-devops-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # DevOps Engineer - **Company:** SocialEdge, Inc. - **Location:** New York, NY, United States (Remote available) - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Elastic Compute Cloud, Amazon S3, Automation of Tests, Bash Shell, Cloud Computing, Cloud Computing Security, Cloud Engineering, Configuration Management, Cyber Security, Continuous Integration, DevOps, Disaster Recovery, Monitoring of Systems, Identity and Access Management, Python (Programming Language), Network Security, Linux System Administration, Log Analysis, Machine Learning, Networking Basics, Routing, Reliability Engineering, Software Tools, Prometheus, Azure Machine Learning, Software Deployment, Systems Integration, Tripwire, Software Vulnerability Management, Data Logging, Scripting, Google Cloud, Load Balancing, Cloud Platform System, Istio, System Availability, Delivery Pipeline, Grafana, Model Validation, Multi-Cloud, AWS Lambda, Apigee, Infrastructure as Code (IaC), Amazon Virtual Private Cloud (VPC), Cloudformation, Event Driven Architecture, Amazon Relational Database Service, Containerization, Gitlab-ci, Kubernetes, Infrastructure Automation Frameworks, Deployment Automation, Nessus, Linkerd (Service Mesh), Machine Learning Operations, Functional Programming, Cloudwatch, Api Gateway, Amazon Simple Queue Service (SQS), Terraform, Devsecops, Serverless Computing, Jenkins, Databricks - **Published:** September 5, 2026 - **Apply:** https://startup.jobs/devops-engineer-creatoriq-com-9895808 ## About the Role * 3+ years of experience in DevOps, Cloud Engineering, Site Reliability Engineering (SRE), or a similar infrastructure-focused role. * 2+ years of hands-on experience with AWS services such as EC2, S3, RDS, Lambda, IAM, VPC, SQS, API Gateway, or similar services. * 2+ years of experience working with containerized environments and orchestration platforms such as Kubernetes and Amazon EKS. * Strong experience building and maintaining CI/CD pipelines using tools such as GitLab CI/CD or Jenkins. * Hands-on experience with Infrastructure as Code using Terraform, Terragrunt, CloudFormation, or similar technologies. * Strong Linux system administration and troubleshooting skills. * Solid understanding of networking fundamentals, including routing, load balancing, network security, and related concepts. * Scripting experience with Python, Bash, or similar languages to automate infrastructure and operational tasks. * Hands-on experience using AI tools to improve engineering workflows, automation, troubleshooting, or agentic use cases. * Experience supporting data, ML, or other compute-intensive production workloads. * Experience with Google Cloud would be valuable, particularly for candidates who have worked across multi-cloud environments. * Familiarity with Helm and service mesh technologies such as Istio, Linkerd, Traefik, or similar tools would be beneficial. * Experience with serverless and event-driven architectures using technologies such as AWS Lambda, API Gateway, and SQS is a plus. * Exposure to cloud and infrastructure security practices, including vulnerability management and tools such as Nessus, Prowler, Trivy, firewalls, or similar technologies, would be valuable. * Knowledge of security standards, compliance requirements, and cloud security best practices is beneficial. * Experience with observability, log analysis, and monitoring platforms such as Coralogix, Prometheus, Grafana, or similar solutions is a plus. * FinOps experience, including cloud cost monitoring, optimization, and accountability practices, would be valuable. * Experience with API gateways or API management platforms such as Kong, Apigee, or similar technologies is beneficial. * Experience with MLOps platforms and practices-particularly Databricks, model serving, ML pipelines, and model monitoring-would be an advantage. ## Description The DevOps Engineer is responsible for supporting and improving cloud and ML/AI infrastructure, automating deployments, and maintaining CI/CD pipelines to ensure efficient, secure, and scalable development workflows. This role plays a crucial part in infrastructure automation, monitoring, and cloud security while collaborating with Software and ML Engineers, Product support, QA, and Security teams. As a key member of the DevOps team, the DevOps Engineer helps manage cloud environments, CI/CD pipelines, and Infrastructure as Code (IaC), ensuring high availability and compliance with security best practices., * Support and maintain scalable, highly available, and secure cloud infrastructure in accordance with company policies and standards. * Provision and manage cloud resources using Infrastructure as Code (Terraform, Terragrunt, CloudFormation). * Implement cloud security best practices, including IAM/role-based access controls, encryption, vulnerability management, and secure infrastructure configurations. * Support containerized environments and orchestration platforms. * Apply DevSecOps principles across infrastructure and deployment workflows. * Participate in disaster recovery planning, testing, and recovery activities., * Maintain and optimize CI/CD pipelines using tools such as GitLab CI/CD and Jenkins, supporting application and ML model deployments. * Improve deployment reliability and support zero-downtime deployment strategies. * Automate configuration management, infrastructure provisioning, and routine operational processes. * Troubleshoot deployment and pipeline issues and implement improvements to prevent recurrence. * Develop scripts and automation to reduce manual work and improve engineering efficiency., * Help design, deploy, operate, and secure infrastructure supporting AI and agentic products, including MCP, agents, integrations, internal tooling, and customer-facing use cases. * Use AI-assisted engineering tools, coding copilots, and AI-driven troubleshooting to improve DevOps productivity and reduce repetitive operational work. * Evaluate and adopt practical AI-enabled workflows that improve infrastructure management, troubleshooting, and operational efficiency., * Operate and scale ML platform infrastructure, including Databricks interactive clusters, jobs compute, ML pipelines, and Model Serving endpoints. * Manage production model-serving infrastructure, including compute capacity, provisioned throughput, and autoscaling for high-throughput inference workloads. * Maintain infrastructure-level monitoring for model drift, data quality, inference performance, and serving health, while partnering with ML Engineering on model evaluation, quality thresholds, and model correctness. * Partner with ML Engineering to support reliable CI/CD and production deployment of ML models. Observability, Incident Response & Engineering Collaboration * Maintain monitoring, logging, metrics, and alerting solutions using tools such as Prometheus, Grafana, Coralogix, and CloudWatch. * Support incident response and perform Root Cause Analysis (RCA) for infrastructure and deployment-related issues. * Improve system observability through effective log aggregation, metrics collection, monitoring, and alerting. * Partner with Software Engineers, ML Engineers, QA, and Software Engineers in Test to improve deployment workflows and integrate automated testing into CI/CD pipelines. * Collaborate with IT Security to maintain secure cloud operations and infrastructure policies. * Respond to engineering and Product Support requests in a timely manner and provide technical infrastructure support when needed. * Maintain accurate internal technical and operational documentation. * Collaborate effectively with international teams across multiple time zones ## Related Videos - [From DevOps to Scaled DevOps: How We’re Rebuilding Continuous Delivery as a Platform](https://www.wearedevelopers.com/videos/100018-from-devops-to-scaled-devops-how-we-re-rebuilding-continuous-delivery-as-a-platform) - [Rate-limiting using eBPF and Istio: How to protect your SaaS customers from themselves](https://www.wearedevelopers.com/videos/100220-rate-limiting-using-ebpf-and-istio-how-to-protect-your-saas-customers-from-themselves) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [My journey into DevOps world - How it all started!](https://www.wearedevelopers.com/videos/545-my-journey-into-devops-world-how-it-all-started) - [#90DaysOfDevOps - The DevOps Learning Journey](https://www.wearedevelopers.com/videos/548-90daysofdevops-the-devops-learning-journey) - [Get started with securing your cloud-native Java microservices applications](https://www.wearedevelopers.com/videos/123-get-started-with-securing-your-cloud-native-java-microservices-applications) ## Related Articles - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Is Software Engineering Over-Saturated?](https://www.wearedevelopers.com/magazine/418-is-software-engineering-over-saturated) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs)