DevOps Engineer

SocialEdge, Inc.
New York, NY, United States
1 day ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Working hours
Regular working hours
Job source

Tech stack

Application Programming Interfaces (APIs) Artificial Intelligence Amazon Elastic Compute Cloud Amazon S3 Automation of Tests Bash Shell Cloud Computing Cloud Computing Security Cloud Engineering Configuration Management Cyber Security Continuous Integration
+54 more
DevOps Disaster Recovery Monitoring of Systems Identity and Access Management Python (Programming Language) Network Security Linux System Administration Log Analysis Machine Learning Networking Basics Routing Reliability Engineering Software Tools Prometheus Azure Machine Learning Software Deployment Systems Integration Tripwire Software Vulnerability Management Data Logging Scripting Google Cloud Load Balancing Cloud Platform System Istio System Availability Delivery Pipeline Grafana Model Validation Multi-Cloud AWS Lambda Apigee Infrastructure as Code (IaC) Amazon Virtual Private Cloud (VPC) Cloudformation Event Driven Architecture Amazon Relational Database Service Containerization Gitlab-ci Kubernetes Infrastructure Automation Frameworks Deployment Automation Nessus Linkerd (Service Mesh) Machine Learning Operations Functional Programming Cloudwatch Api Gateway Amazon Simple Queue Service (SQS) Terraform Devsecops Serverless Computing Jenkins Databricks

Job description

The DevOps Engineer is responsible for supporting and improving cloud and ML/AI infrastructure, automating deployments, and maintaining CI/CD pipelines to ensure efficient, secure, and scalable development workflows. This role plays a crucial part in infrastructure automation, monitoring, and cloud security while collaborating with Software and ML Engineers, Product support, QA, and Security teams.

As a key member of the DevOps team, the DevOps Engineer helps manage cloud environments, CI/CD pipelines, and Infrastructure as Code (IaC), ensuring high availability and compliance with security best practices., * Support and maintain scalable, highly available, and secure cloud infrastructure in accordance with company policies and standards.

  • Provision and manage cloud resources using Infrastructure as Code (Terraform, Terragrunt, CloudFormation).
  • Implement cloud security best practices, including IAM/role-based access controls, encryption, vulnerability management, and secure infrastructure configurations.
  • Support containerized environments and orchestration platforms.
  • Apply DevSecOps principles across infrastructure and deployment workflows.
  • Participate in disaster recovery planning, testing, and recovery activities., * Maintain and optimize CI/CD pipelines using tools such as GitLab CI/CD and Jenkins, supporting application and ML model deployments.
  • Improve deployment reliability and support zero-downtime deployment strategies.
  • Automate configuration management, infrastructure provisioning, and routine operational processes.
  • Troubleshoot deployment and pipeline issues and implement improvements to prevent recurrence.
  • Develop scripts and automation to reduce manual work and improve engineering efficiency., * Help design, deploy, operate, and secure infrastructure supporting AI and agentic products, including MCP, agents, integrations, internal tooling, and customer-facing use cases.
  • Use AI-assisted engineering tools, coding copilots, and AI-driven troubleshooting to improve DevOps productivity and reduce repetitive operational work.
  • Evaluate and adopt practical AI-enabled workflows that improve infrastructure management, troubleshooting, and operational efficiency., * Operate and scale ML platform infrastructure, including Databricks interactive clusters, jobs compute, ML pipelines, and Model Serving endpoints.
  • Manage production model-serving infrastructure, including compute capacity, provisioned throughput, and autoscaling for high-throughput inference workloads.
  • Maintain infrastructure-level monitoring for model drift, data quality, inference performance, and serving health, while partnering with ML Engineering on model evaluation, quality thresholds, and model correctness.
  • Partner with ML Engineering to support reliable CI/CD and production deployment of ML models.

Observability, Incident Response & Engineering Collaboration

  • Maintain monitoring, logging, metrics, and alerting solutions using tools such as Prometheus, Grafana, Coralogix, and CloudWatch.
  • Support incident response and perform Root Cause Analysis (RCA) for infrastructure and deployment-related issues.
  • Improve system observability through effective log aggregation, metrics collection, monitoring, and alerting.
  • Partner with Software Engineers, ML Engineers, QA, and Software Engineers in Test to improve deployment workflows and integrate automated testing into CI/CD pipelines.
  • Collaborate with IT Security to maintain secure cloud operations and infrastructure policies.
  • Respond to engineering and Product Support requests in a timely manner and provide technical infrastructure support when needed.
  • Maintain accurate internal technical and operational documentation.
  • Collaborate effectively with international teams across multiple time zones

Requirements

  • 3+ years of experience in DevOps, Cloud Engineering, Site Reliability Engineering (SRE), or a similar infrastructure-focused role.
  • 2+ years of hands-on experience with AWS services such as EC2, S3, RDS, Lambda, IAM, VPC, SQS, API Gateway, or similar services.
  • 2+ years of experience working with containerized environments and orchestration platforms such as Kubernetes and Amazon EKS.
  • Strong experience building and maintaining CI/CD pipelines using tools such as GitLab CI/CD or Jenkins.
  • Hands-on experience with Infrastructure as Code using Terraform, Terragrunt, CloudFormation, or similar technologies.
  • Strong Linux system administration and troubleshooting skills.
  • Solid understanding of networking fundamentals, including routing, load balancing, network security, and related concepts.
  • Scripting experience with Python, Bash, or similar languages to automate infrastructure and operational tasks.
  • Hands-on experience using AI tools to improve engineering workflows, automation, troubleshooting, or agentic use cases.
  • Experience supporting data, ML, or other compute-intensive production workloads.
  • Experience with Google Cloud would be valuable, particularly for candidates who have worked across multi-cloud environments.
  • Familiarity with Helm and service mesh technologies such as Istio, Linkerd, Traefik, or similar tools would be beneficial.
  • Experience with serverless and event-driven architectures using technologies such as AWS Lambda, API Gateway, and SQS is a plus.
  • Exposure to cloud and infrastructure security practices, including vulnerability management and tools such as Nessus, Prowler, Trivy, firewalls, or similar technologies, would be valuable.
  • Knowledge of security standards, compliance requirements, and cloud security best practices is beneficial.
  • Experience with observability, log analysis, and monitoring platforms such as Coralogix, Prometheus, Grafana, or similar solutions is a plus.
  • FinOps experience, including cloud cost monitoring, optimization, and accountability practices, would be valuable.
  • Experience with API gateways or API management platforms such as Kong, Apigee, or similar technologies is beneficial.
  • Experience with MLOps platforms and practices-particularly Databricks, model serving, ML pipelines, and model monitoring-would be an advantage.

Benefits & conditions

What you will get from us:

  • People: Work with talented, collaborative, and friendly people who love what they do.
  • Guidance: Utilize our learning platform to fully get the training and tools you’ll need to become successful here from your first day with us.
  • Work/life harmony: 15 days of vacation, floating and company holidays, wellness benefits, and paid parental leave.
  • Whole Health Package: Comprehensive medical, dental, vision, life, and disability insurance, plus additional wellness benefits.
  • Planning for the future: A 401(k) plan to help you plan ahead.
  • Work from home stipend: To assist you in setting up a home office that works for you., We understand that a comprehensive benefits package plays a significant role in your overall compensation. To gain more insight into the various components of our total compensation, we invite you to review our benefits and perks.

About the company

CreatorIQ is the operating system for creator-led growth trusted by more than 1,300 global brands and agencies.

We’re on a mission to make businesses more human, and humans more impactful. We operate by our values - be intentional, pursue excellence every day, embrace the journey together, and be a good human - every day. CreatorIQ has earned the title of best companies to work for in multiple programs, including BuiltIn LA and NY. It’s been named a Fastest-Growing Company in North America on the Deloitte Technology Fast 500 for four years, was named a leader in IDC MarketScape: Worldwide Influencer Marketing Platforms for Large Enterprises in 2025, was named a Leader by The Forrester New Wave : Influencer Marketing Solutions, and has been consistently recognized by G2 as a Leader, and is rated 5 stars on Influencer MarketingHub. We operate in a flexible work model that combines both in-person and remote work to boost collaboration, enhance innovation, and adapt to individual work styles.

We’re seeking passionate, innovative minds to join our journey. Be a part of our dynamic team and let’s transform the industry together!, CreatorIQ is the operating system for creator-led growth, helping global brands and agencies transform creator marketing into an intelligence-driven growth engine. Powered by the Creator Graph , which processes more than 250 million social posts daily across more than 15 million creators worldwide, CreatorIQ unifies fragmented platform data into a centralized intelligence layer and system of record for creator relationships, performance, governance, and commerce. More than 1,300 organizations-including Dentsu, Delta Air Lines, Google, Beiersdorf, Nestlé, and Wella-rely on CreatorIQ as the infrastructure to run and scale their creator programs globally. CreatorIQ is a global company headquartered in Los Angeles with offices in Austin, New York, San Francisco, London, Manila, and Warsaw. Learn more at www.creatoriq.com and follow us on LinkedIn and Instagram.

At CreatorIQ, we believe that diversity is the key to unlocking our full potential. We are committed to fostering an inclusive, equitable, and empowering work environment where everyone can thrive, regardless of race, ethnicity, gender, sexual orientation, age, religion, disability, or any other characteristic that makes us unique. By embracing our core values of being intentional, pursuing excellence every day, embracing the journey together, being a good human, and staying focused on what’s important, we create an atmosphere that promotes collaboration and growth. Join us to celebrate differences, innovate together, and be a part of a business that is disrupting the marketing industry.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

3:32 min

Shifting to a DevOps career from non-technical backgrounds

Megha Kadur · LIVE

2:17 min

Mapping the maturity roadmap for scaled devops adoption

Dominik Krichbaum Dominik Krichbaum · World Congress 2026 Europe

2:53 min

Configuring dynamic proxy updates with Istio Pilot

Jan Mensch Jan Mensch · World Congress 2026 Europe

1:20 min

Identifying multi-disciplinary talent for developer experience engineering roles

Hazal Mestci +1 · Coffee With Developers

3:18 min

Scaling global network engineering through DevOps culture

Stuart Clark · LIVE

7:15 min

Installing Istio programmatically with bash scripts

Thomas Südbröcker · LIVE

Videos

See all

Related articles

See all