Platform Engineer
Role details
Job location
Tech stack
Job description
We are looking for a Platform Engineer to help operate, improve, and scale Tigo Energy's hybrid technology platform. Our on-prem colocation environment is the primary infrastructure environment, supported by Microsoft Azure for secondary cloud workloads.
The role is approximately evenly divided between platform engineering, including Kubernetes, GitLab, GitOps, infrastructure automation, CI/CD, and observability; and hands-on work with Linux systems, networking, storage, server hardware, and colocation operations. This is not a cloud-only position; hands-on infrastructure work is an essential part of the role.
We are open to candidates at different experience levels and will calibrate scope and responsibilities accordingly. We do not expect applicants to have professional experience with every technology in our environment. Strong fundamentals, practical curiosity, structured troubleshooting, and the ability to learn unfamiliar systems are more important than matching a long checklist of tools.
Working environment and on-call
This Bay Area position requires visits to our colocation facility several times per month for planned infrastructure work, maintenance, troubleshooting, or incident response.
Hands-on colocation work is an essential function of the position. With or without reasonable accommodation, the engineer must work safely around server racks, handle cabling, install or replace components, and move equipment using appropriate tools or assistance.
The role includes shared on-call coverage and occasional after-hours maintenance. Scheduling may evolve as coverage needs change; current expectations will be discussed during the interview process., * Operate and improve the hybrid platform anchored in our colocation environment, with supporting workloads and services in Azure.
- Support Linux systems, Kubernetes clusters, networking, storage, databases, and production services across development and operational environments.
- Build and maintain infrastructure-as-code and configuration workflows using tools such as Terraform and Ansible.
- Improve GitLab CI/CD and GitOps workflows, including build pipelines, release promotion, deployment verification, and rollback processes using tools such as Helm, Argo CD, or Flux.
- Install, maintain, and troubleshoot physical infrastructure, including servers, storage, network equipment, cabling, disks, firmware, and out-of-band management systems; coordinate with vendors and remote hands when appropriate.
- Improve monitoring, alerting, logging, dashboards, and operational visibility using platforms such as Prometheus, Grafana, centralized logging, and related observability tools.
- Support data and storage platforms such as PostgreSQL, ClickHouse, Redis, Ceph, NFS, and S3-compatible object storage, including backup, recovery, capacity, and availability considerations.
- Strengthen infrastructure security through access controls, secrets management, patching, vulnerability management, audit logging, backup, and disaster-recovery practices.
- Participate in incident response, root-cause analysis, post-incident improvements, and a shared on-call process.
- Create and maintain clear documentation, including runbooks, architecture and network diagrams, operational procedures, and change records.
- Collaborate with software, hardware, security, and operations teams to improve reliability, deployment quality, and developer experience., * Learn the architecture, deployment workflows, monitoring systems, and operational processes that support our platform.
- Contribute safely to day-to-day Linux, Kubernetes, networking, Azure, storage, and colocation work.
- Improve documentation, runbooks, dashboards, alert quality, and repeatable operational procedures.
- Build trust through clear communication, careful change execution, and reliable follow-through.
Over time
- Own larger platform, infrastructure, reliability, and automation projects appropriate to your level.
- Improve deployment safety, observability, security posture, and recovery readiness.
- Reduce manual operational work and simplify legacy or transitional systems.
- Increase the reliability, maintainability, scalability, and visibility of the hybrid platform.
Requirements
- Education. Bachelor's degree in Computer Science, Computer Engineering, Information Systems, or a closely related technical field.
- Learning and structured troubleshooting. You can learn unfamiliar technologies, form useful hypotheses, and work methodically from evidence toward a root cause.
- Networking fundamentals. You understand concepts such as TCP/IP, DNS, routing, firewalls, VPNs, TLS, VLANs, and load balancing well enough to apply them to real operational problems.
- Cloud or container foundations. You have practical exposure to a cloud platform, containers, Kubernetes, virtualization, or a substantial self-managed environment such as a homelab.
- Hands-on infrastructure aptitude. You are comfortable working with servers, storage, network equipment, cabling, component replacement, and out-of-band management.
- Communication and documentation. You can explain findings clearly, ask useful questions, and create documentation that other people can follow.
- Ownership and operational judgment. You raise risks early, plan changes carefully, verify outcomes, and follow problems through to resolution.
- Collaborative working style. You can work effectively with people across engineering, operations, security, and business teams.
- Relevant experience may come from professional roles, internships, apprenticeships, military service, personal projects, homelabs, or open-source work and will be considered alongside formal education.
Experience That Will Help You Succeed
Experience with any of the following is useful, but no candidate is expected to have worked with every item:
- Linux administration, particularly Ubuntu or Talos Linux, plus Bash, Python, or another scripting language.
- Kubernetes, Helm, operators, ingress, CNI/CSI integrations, RBAC, NetworkPolicies, cluster upgrades, or backup and restore.
- Microsoft Azure services such as virtual networks, network security groups, virtual machines, private DNS, VPN Gateway, Key Vault, Azure Container Registry, Entra ID, Azure Monitor, or cost management.
- GitLab, GitLab CI/CD, runners, container registries, GitOps, Argo CD, Flux, Terraform, or Ansible.
- Prometheus, Grafana, Alertmanager, OpenSearch, ELK, Loki, OpenTelemetry, tracing, service-level indicators, or alert-quality improvements.
- PostgreSQL, ClickHouse, Redis, Ceph, NFS, object storage, replication, high availability, backup and restore testing, or capacity planning.
- Colocation and hardware operations, including rack and stack, cabling, BMC/IPMI/iDRAC/iLO, switch or firewall configuration, inventory, firmware, disk replacement, and RMA processes.
- Security and reliability tooling such as Wazuh, vulnerability or container scanning, secrets management, certificate management, audit logging, patch management, or disaster-recovery testing.