Manager, IT & Infrastructure Operations

Agilence, Inc.
United States
about 1 month ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Compensation
$120,000.0 - $150,000.0
Working hours
Regular working hours
Job source

Tech stack

Microsoft Windows Active Directory Amazon Web Services Server Applications Systems Engineering JIRA Microsoft Azure Backup Devices Bash Shell Business Systems Ubuntu (Operating System) CentOS
+50 more
Software as a Service Cloud Computing Databases Continuous Integration Linux DevOps Disaster Recovery Middleware Monitoring of Systems Hyper-V Identity and Access Management Information Technology Operations Networking Hardware Virtual Private Networks (VPN) Linux Distribution Log Analysis Uptime Microsoft SQL Server Windows Servers Nagios Oracle (Applications) Windows PowerShell Red Hat Enterprise Linux Ansible Prometheus Runbook Salesforce.Com Security Assertion Markup Language (SAML) SQL Databases Systems Integration Transmission Control Protocol (TCP) Virtual Local Area Networks Virtualization Technology Software Vulnerability Management Data Logging Enterprise Software Applications Cloud Platform System Computer Network Technologies Grafana Infrastructure Automation Frameworks Bug Reporting Information Technology CIS Benchmarks Teamcity Cloud Optimization Firewall Services Module Zendesk Terraform Jenkins Vmware

Job description

The Manager, IT & Infrastructure Operations leads the team responsible for the systems, equipment, and applications that keep Agilence’s production environments running. This includes the hybrid infrastructure - Windows Server, Linux, core networking, virtualization, and cloud environments in Azure and AWS - that hosts every product Agilence builds and sells, along with the production application stack itself. The role also owns corporate IT: the identity platforms, endpoints, and business systems the company runs on.

This is a hands-on leadership role reporting to the CTO. The manager is accountable for the availability, performance, security, and cost of both the production and corporate environments, leads a small team of systems administrators and infrastructure engineers, and partners closely with DevOps, Engineering, Security, and Customer Success on day-to-day operations, planned projects, and incident response.

The team is also the technical escalation point for the Customer Success organization. When a customer-reported application issue cannot be resolved at the front line, it comes to this team to troubleshoot, reproduce, and diagnose across the application, database, integration, and infrastructure layers. Issues within the team’s control are resolved directly; those requiring a code change are escalated to the development team with clear, well-documented findings. Customer Success owns the customer relationship and end-user support for Agilence applications; this role owns the technical diagnosis behind it and the platforms those products run on., Lead, coach, and develop a small team of systems administrators and infrastructure engineers, setting priorities and holding the team to consistent operational standards Own the availability, performance, and capacity of the production environments that host the Agilence and IntelliQ product suites Manage the hybrid infrastructure footprint, including Windows Server, Linux, networking equipment, virtualization, and cloud resources in Azure and AWS Administer the production application stack, including web and application servers, databases, integration services, and supporting middleware Administer corporate IT systems and services, including Microsoft 365, identity and access management (Active Directory, Entra ID), endpoint management, and employee onboarding and offboarding Serve as the technical escalation point for Customer Success on production application issues, troubleshooting and diagnosing root cause across application, database, integration, and infrastructure layers Resolve escalated issues within the team’s scope, and escalate confirmed product defects to the development team with clear reproduction steps, logs, and analysis so they can be triaged efficiently Set and enforce escalation standards, response expectations, and hand-off quality between Customer Success, IT Operations, and Engineering; track escalation trends and drive fixes for recurring issues Own the incident management process for infrastructure and platform events, including the on-call rotation, escalation paths, root cause analysis, and corrective action Drive patching, system hardening, vulnerability remediation, and backup operations; own disaster recovery planning, testing, and documented recovery objectives Support SOC 2 and customer security requirements by maintaining controls, access reviews, evidence collection, and audit responses Manage vendor relationships, hardware and software renewals, licensing, and cloud spend; forecast infrastructure budget needs Partner with DevOps and Engineering on release, deployment, and environment requirements Maintain configuration documentation, topology diagrams, runbooks, and SOPs Identify and drive automation and process improvement to reduce manual effort and operational risk Participate in architecture reviews, capacity planning, and infrastructure roadmap planning

Requirements

Bachelor’s degree in Computer Science, Information Technology, or a related field, or equivalent practical experience 7+ years in IT operations, infrastructure, or systems engineering, including 2+ years leading or managing a team Hands-on experience administering Windows Server (2016 and later) and mainstream Linux distributions (Ubuntu, RHEL/CentOS) Experience operating production SaaS or enterprise software environments with uptime commitments Working knowledge of Azure and AWS infrastructure services (VMs, VPCs, storage, IAM, networking) Solid understanding of TCP/IP networking, VLANs, VPNs, and firewall policies Experience with virtualization platforms such as VMware or Hyper-V Experience with Microsoft 365 administration and hybrid identity and SSO (Active Directory, Entra ID, SAML) Competent in scripting and automation using PowerShell and/or Bash Experience with system monitoring, alerting, and log analysis tools Demonstrated ability to troubleshoot production application issues end to end, using application and server logs, SQL queries, and integration traces to isolate root cause Experience operating in an escalation or Tier 3 support model, working between a customer-facing team and a development team, with the judgment to know when an issue is a defect versus a configuration, data, or environment problem Working knowledge of SQL sufficient to investigate data and application behavior directly Demonstrated ownership of backup, disaster recovery, and business continuity practices Excellent troubleshooting, communication, and documentation skills, with the ability to translate technical findings for both Customer Success and Engineering audiences

Preferred Experience Experience supporting a SOC 2 or similarly audited environment, including CIS Benchmarks and access review processes Familiarity with infrastructure-as-code and configuration management tools such as Terraform or Ansible Working knowledge of centralized logging and monitoring platforms (e.g., ELK, Prometheus, Grafana, Nagios) Experience administering Microsoft SQL Server and/or Oracle in production Familiarity with CI/CD tooling and release processes (e.g., Jenkins, TeamCity, Azure DevOps) Experience with a ticketing and case management system used across support and engineering (e.g., Atlassian Jira, Zendesk, Salesforce Service Cloud) Experience managing infrastructure budgets, vendor contracts, and cloud cost optimization Experience supporting distributed teams across multiple regions and time zones

Search Firm Representatives please read: Agilence is not seeking assistance or accepting unsolicited resumes from search firms for this employment opportunity. Regardless of past practice, all resumes submitted by search firms to any employee at Agilence via-email, the Internet or directly to hiring managers at Agilence in any form without a valid written search agreement in place for that position will be deemed the sole property of Agilence, and no fee will be paid in the event the candidate is hired by Agilence as a result of the referral or through other means.

Benefits & conditions

Pulled from the full job description

  • Health insurance
  • Paid time off, We’ve operated almost entirely remote since 2020 and, in 2022, implemented a completely remote workplace culture.

Paid Time Off

As an integral part of our team, you deserve time off to refresh and re-energize.

Medical Insurance

Gain access to various medical plans to fit your and your family’s healthcare needs.

Employee Recognition

Agilence Impact Award Winners are selected every month by their peers.

Flexible 401k Retirement Plans

Plan for the future today by taking advantage of company matching contributions.

About the company

Agilence is the leader in loss prevention analytics, helping prominent retail, restaurant, and grocery companies increase their profit margins by reducing preventable loss.

At Agilence, we specialize in uniting digital and physical transactions to help cutting-edge loss prevention teams expand beyond traditional theft and fraud to tackle preventable loss in all its forms - in the store, online, and at the corporate office.

Every day, Agilence analyzes over 24 million transactions for our customers, transforming data into insights, and insights into actions. Our platform combines data from 200+ sources, including point-of-sale (POS), eCommerce, HR, labor, inventory, product, third-party delivery platforms, alarms, case management, loyalty, access control, video surveillance, and more.

Companies have saved millions of dollars by optimizing operations, identifying sources of margin erosion, and reducing shrink using Agilence. Many have also improved employee and customer safety, identified training opportunities, improved customer experiences, increased promotional success and eliminated productivity gaps.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:02 min

Asking about infrastructure and engineering culture operations

Karol Rogowski · LIVE

1:40 min

Managing containerized infrastructure with Podman Desktop

Cedric Clyburn Cedric Clyburn +1 · World Congress 2025

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · World Congress 2026 Europe

1:41 min

Parallels between cloud and legacy infrastructure lock-ins

Björn Stahl Björn Stahl · World Congress 2024

Videos

See all

Related articles

See all