Server Engineer

True North ITG, Inc.
United States
2 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$91,098.0 - $109,709.0
Working hours
Regular working hours
Job source

Tech stack

Proxmox Link Aggregation (Ethernet) Active Directory Artificial Intelligence Backup Devices Intelligent Platform Management Interface Bash Shell Border Gateway Protocol Health Informatics BIOS Cisco PIX Data Centers
+42 more
Debian Linux Linux RAID Domain Name System (DNS) Network Interface Controllers Trunking Firmware Virtual Private Networks (VPN) Python (Programming Language) Kernel-Based Virtual Machine Network Security Network Layer Linux System Administration Windows Servers Network Diagrams Network Segmentation Packet Analyzer Open Shortest Path First (OSPF) PCI Express Performance Tuning Ansible Virtual Local Area Networks Virtualization Technology Wide Area Networks Zabbix Ceph (Software) Scripting Grafana Caching Firewalls (Computer Science) Backend Bare Metal PfSense Fortinet ZFS File System Hardware Infrastructure Restful APIs Terraform Network Server Cisco Docker Nvme

Job description

Virtualization and storage (primary focus)

  • Design, deploy, upgrade and troubleshoot Proxmox VE clusters end to end, from bare metal through corosync, HA and production cutover
  • Architect and operate Ceph: CRUSH maps, device classes and pools, OSD lifecycle, PG tuning, rebalancing, capacity planning and failure recovery
  • Manage ZFS pools, caching and tuning for local and backup storage
  • Plan and execute live migrations, cross-cluster moves, PCIe/GPU passthrough and version upgrades with minimal downtime
  • Build and maintain standard VM templates, automation and provisioning workflows

Backup and disaster recovery

  • Operate Proxmox Backup Server, including per-customer pools, pool-based jobs, retention, verification and prune/GC schedules
  • Run regular restore tests and maintain documented, tested recovery procedures for each client
  • Monitor backup health and capacity, and resolve failures before they become data-loss events

Backend stack and monitoring

  • Maintain the supporting stack: Active Directory and DNS, monitoring (Zabbix), Docker-based services and internal tooling
  • Respond to alerts, perform root cause analysis and write clear post-incident reports

Hardware and data center operations

  • Travel to data centers for hardware installs, replacements and upgrades, including servers, drives, NICs, memory and power components
  • Diagnose and repair hardware faults across Gigabyte, Supermicro, Dell, HP and other OEM platforms, and manage vendor RMAs and support cases
  • Rack, cable, label and document new deployments to a consistent standard

Networking and edge security

  • Configure and troubleshoot routed and switched networks using OSPF, BGP, VLANs, LACP and SD-WAN
  • Administer edge firewalls on Cisco ASA, Netgate/pfSense and FortiGate, including policy, VPN and segmentation changes
  • Work with network and security teammates on changes that touch cluster and storage traffic

Documentation and collaboration

  • Keep runbooks, network diagrams and asset records accurate and current
  • Participate in change control, capacity planning and on-call rotation
  • Mentor junior staff and support other engineers on escalations, To our Tacoma, WA and Las Vegas, NV data centers as needed for hardware replacements, upgrades and new builds; company-paid

On-call

Shared rotation covering production alerts and after-hours maintenance windows

Maintenance windows

Some planned work is scheduled evenings and weekends to protect client uptime

Physical demands

Lift up to 50 lbs, work in data center environments and rack equipment

Core Competencies

  • Ownership: you treat production as your own, from initial design to the post-incident review.
  • Systems thinking: you understand how a change in one layer, such as a switch config or a Ceph pool setting, affects the rest.
  • Calm under pressure: you troubleshoot methodically during outages and communicate clearly while doing it.
  • Independence: you can take a project from requirements to production with minimal supervision.
  • Documentation discipline: you leave systems better documented than you found them.
  • Client focus: you understand that behind every VM is a business, and in healthcare, often a patient.

Requirements

You do not need to check every box in the lower tiers, but the first tier is essential.

  1. Proxmox, Ceph, ZFS and KVM (essential, very high level) * 5+ years running Proxmox VE and KVM in production, including multi-node clusters, HA and live migration * Deep, hands-on Ceph experience: you have built clusters from scratch, tuned them under load and recovered them from real failures * Strong ZFS knowledge: pool design, ARC/L2ARC/SLOG, replication and performance tuning * Able to design and deploy a complete cluster alone, and to explain how compute, storage, network and backup layers interact * Strong Linux administration (Debian preferred) and Bash or Python scripting

  2. Enterprise networking (strong) * Working knowledge of OSPF and BGP in production environments * VLAN design, trunking, LACP/bonding and MTU/jumbo frame configuration for storage networks * General SD-WAN experience * Ability to troubleshoot Layer 2 and Layer 3 problems with packet captures and switch/router CLI tools * Experience with Cisco, Meraki or similar enterprise platforms

  3. Physical server hardware (strong) * Hands-on experience building, servicing and upgrading rack servers from Gigabyte, Supermicro, Dell, HP and other OEMs * Comfort with IPMI/iDRAC/iLO, firmware and BIOS updates, RAID/HBA and NVMe configuration, and hardware diagnostics * Physical data center skills: racking, cabling, power and labeling

  4. Edge security (solid) * Experience administering Cisco ASA, Netgate/pfSense and FortiGate firewalls * Firewall policy, NAT, site-to-site and remote-access VPN, and network segmentation

General requirements

  • Clear written and verbal communication, including documenting your own work
  • Ability to work independently in a fast-moving, multi-client environment
  • Willingness to travel and to join an on-call rotation, * Experience with Proxmox Backup Server at scale, including pool-based job design and tiered NVMe/HDD backup targets
  • Experience with mixed-media Ceph clusters (NVMe and HDD tiers) and cross-cluster or cross-site migration
  • PCIe and GPU passthrough experience, including GPU nodes for AI workloads
  • Windows Server, Active Directory and Group Policy administration
  • Automation and infrastructure-as-code experience (Ansible, Terraform, Python, REST APIs)
  • Monitoring experience with Zabbix, Grafana or similar tools
  • Prior MSP or multi-tenant hosting experience
  • Basic healthcare IT knowledge, such as HIPAA-aware operations and the uptime and data-integrity demands of clinical systems
  • Relevant certifications, such as CCNA/CCNP, Fortinet NSE, Linux (LPIC, RHCE) or Proxmox certification

Benefits & conditions

As a remote company, we will provide compensation in writing at the time of offer, if extended, and determine it based on work location and other relevant factors, including but not limited to experience, skills, and other job-related factors. Internal equity, market, and organizational factors are also considered.

Equal Opportunity

TrueNorth ITG is an equal opportunity employer. We consider all qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status or any other protected status under applicable law.

Pay: $91,097.92 - $109,709.33 per year, * 401(k)

  • 401(k) matching
  • Dental insurance
  • Health insurance
  • Health savings account
  • Paid time off
  • Professional development assistance
  • Referral program
  • Tuition reimbursement
  • Vision insurance

About the company

TrueNorth ITG is hiring a Server Engineer to own the health, growth and resilience of our Proxmox and Ceph virtualization platform and the backend stack around it. This is a senior, hands-on infrastructure role. You can design and stand up a Proxmox/Ceph cluster from bare metal to production on your own, you understand how every layer fits together, and you will get on a plane when a data center needs hands.

Skills are weighted in this order:

  1. Proxmox / Ceph (most important): a very high level of depth, with ZFS and KVM alongside.

  2. Enterprise networking: routing, switching and SD-WAN at production scale.

  3. Physical server hardware: multi-vendor build, repair and upgrade work.

  4. Edge security: firewall administration across major platforms.

About TrueNorth ITG and the Environment

TrueNorth ITG is a managed services provider and datacenter hosting operator supporting multiple clients, including healthcare organizations where uptime and data integrity are not negotiable. You will join a small, senior infrastructure team that runs production platforms across multiple sites.

What you will work with:

  • Multiple Proxmox VE clusters across our Las Vegas and Tacoma data centers, including all-NVMe and HDD-tier Ceph storage
  • Proxmox Backup Server infrastructure with per-customer backup pools and retention policies
  • A mixed hardware fleet from Gigabyte, Supermicro, Dell, HP and other OEMs, including GPU nodes
  • Cisco, Meraki and FS networking, with FortiGate, Cisco ASA and Netgate at the edge
  • Client workloads that include healthcare organizations

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

41 sec

Massive client data loss and bio-digital storage

Chris Heilmann +1 · LIVE

7:31 min

Essential foundational skills and concepts for infrastructure roles

Megha Kadur · LIVE

2:39 min

Experiencing core Linux capabilities for DevOps administration

Michael Cade · LIVE

6:13 min

Defining cloud proficiency by technical role

Piet Van Dongen · LIVE

Videos

See all

Related articles

See all