Picogrid's first Site Reliability Engineer

Picogrid, Inc.
United States
3 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
3 years minimum
Compensation
$170,000.0 - $195,000.0
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Encodings Databases Software Debugging Federal Information Processing Standards (FIPS) Identity and Access Management Mesh Networking Reliability Engineering Prometheus System Availability Grafana Git
+3 more
Kubernetes Terraform Devsecops

Job description

As Picogrid’s first Site Reliability Engineer you will own production reliability across cloud and edge, from observability and incident response through node lifecycle, stateful workloads, and a fleet of hardware edge devices in the field. You will help build and define the systems, processes and best practices that ensure Picogrid’s systems can be relied upon by our warfighters in even the toughest battlefield conditions. You will work with engineers to build a strong on-call culture where issues are root caused swiftly, and ensure our alerting and monitoring have exceptional coverage and signal-to-noise ratio.

Security is a shared responsibility across all our DevSecOps roles, and as part of a scrappy startup team you will be expected to help stand up new infrastructure and other related DevSecOps tasks as needed., * Own, define and drive our reliability SLIs and SLOs for cloud deployments

  • Own, define and drive our reliability SLIs and SLOs for our edge devices deployed in remote and sometimes contested areas
  • Own the observability stack: Grafana, Prometheus, Loki, and OpenTelemetry, with dashboards versioned in git and alerting rules checked in alongside the code they watch
  • Participate in on-call and incident response: log-first troubleshooting, blameless postmortems, and follow-up hardening
  • Encode reliability into infrastructure as code

Requirements

  • 3+ years of experience as an SRE or related roles
  • Deep Kubernetes operations experience: node lifecycle, workload scheduling, StatefulSets, graceful drains, and live cluster debugging
  • Experience designing comprehensive observability dashboards and high signal-to-noise ratio alerting rules
  • You are a competent and experienced incident responder practicing methodical evidence-first triage, blameless postmortems, and turning incidents into durable guardrails
  • Production Terraform or OpenTofu experience
  • Fluent in AWS including IAM, networking, multi-account environments, and account and workload hardening
  • Experience managing high availability database deployments
  • IoT or edge fleet operation experience
  • Comfortable operating in scrappy, fast-paced environments, and turning ambiguous requirements into concrete solutions
  • You optimize for providing value early in projects and short iteration cycles, * GovCloud, FIPS, or other regulated or air-gapped environment experience
  • Constrained edge hardware such as NVIDIA Jetson platforms (AGX Thor, Orin Nano), including shared CPU and GPU memory and thermal constraints
  • Overlay or mesh networking operations: Nebula, WireGuard, Tailscale, or similar
  • Standing up SLO and error-budget tooling (sloth, Pyrra, or equivalent) from scratch
  • Active security clearance, To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State.

Benefits & conditions

  • Base salary range: $170,000 - $195,000 per year. Base salary is just one part of your total compensation package at Picogrid.
  • Significant stock options with a high potential upside as an early-stage company
  • 401(k) with employer matching
  • Full health coverage (medical, dental, and vision insurance)
  • Relocation assistance provided (if applicable)
  • Unlimited PTO (two-week minimum) and 11 paid holidays per year
  • Paid parental leave for both parents
  • Lunch provided when working in-office and a fully stocked kitchenette
  • Free EV charging at the HQ
  • Unique office in El Segundo, CA stocked with quality coffee, snacks, and craft beer

About the company

Picogrid is a leading venture-backed defense technology company founded to bridge the decades-long gap between modern technology and the critical demands of national security. Today, we’re building the essential infrastructure to unify sensors, autonomy, and operators with our technology deployed in active operations around the world. Our mission is to deliver an operational advantage to secure the United States and its allies.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

10:40 min

Visualizing Prometheus open metrics using custom Grafana dashboards

Stijn Polfliet · LIVE

3:17 min

Optimizing character encoding with Kim variable byte encoding

Douglas Crockford Douglas Crockford · World Congress 2024

1:34 min

Pivoting careers into specialized platform engineering roles

Xavier Portilla Edo · LIVE

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

3:50 min

Navigating specialized roles and toolsets across engineering teams

Nele Uhlemann · World Congress 2023

Videos

See all

Related articles

See all