> Markdown version of [/jobs/ext/3544873-devops-engineer](https://www.wearedevelopers.com/jobs/ext/3544873-devops-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # DevOps Engineer - **Company:** Southwest Research Institute - **Location:** Boston, MA, United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Amazon Web Services, Amazon S3, Systems Engineering, Unit Testing, Bash Shell, Big Data, C++ (Programming Language), Software as a Service, Cloud Computing, Nvidia CUDA, Continuous Integration, Data Governance, Data Sharing, DevOps, Github, Hardware-In-The-Loop Simulation, Identity and Access Management, Virtual Private Networks (VPN), Python (Programming Language), PostgreSQL, Linux System Administration, Machine Learning, Networking Basics, Simulation Software, Workflow Management Systems, Data Logging, Pulumi, Data Processing, Amazon Virtual Private Cloud (VPC), Kubernetes, AWS Fargate, Machine Learning Operations, Restful APIs, Terraform, Data Pipelines, Docker - **Published:** September 30, 2026 - **Apply:** https://www.juju.com/job/16_e6fc9f251 ## About the Role * 4+ years in DevOps, platform, infrastructure, or SRE roles * AWS in production: ECR, S3, IAM, VPC, EKS, Batch, ECS/Fargate * Infrastructure as code (Terraform, CDK, Pulumi, or equivalent) with a review-and-version discipline * CI/CD design and operation at scale - GitHub Actions strongly preferred, including self-hosted runners * Kubernetes in production, including workload scheduling and resource governance * Container tooling and build optimization: Docker, BuildKit or daemonless alternatives (Buildah, Kaniko), multi-arch builds, remote caching * Workflow orchestration - Apache Airflow or equivalent * Python, plus comfort in Bash and reading C++ * Linux systems administration and networking fundamentals * Observability: logging, metrics, tracing, and alerting you actually built ## Description You will be the first dedicated infrastructure hire inside FAIRI. You will build the CI/CD, data pipelines, and IaC that let a small research team ship reliably - partnering with Field AI's platform, cloud, and data-processing teams rather than rebuilding what they already run well. This is a hands-on ownership role, not a coordination role. What You'll DoCI/CD and Build Infrastructure - 30% * Stand up CI/CD for the humanoid monorepo: containerize, push to ECR, run unit tests, build, and gate on simulation system tests before promotion. * Work with the platform team's self-hosted GitHub Actions runners (ARM, AMD, CUDA, Jetson-class targets) rather than standing up parallel infrastructure. * Cut build times through change detection and remote caching - full builds are currently ~45 minutes uncached. * Build test infrastructure that lets the same test run against simple sim, Isaac Sim, or real hardware, driven over ROS 2 messages or the robot REST API. * Establish per-automation integration tests so shared-library and output-format changes cannot silently break pipelines. Data Pipelines and Orchestration - 30% * Stand up and own FAIRI's Airflow stack for humanoid data processing. * Build the ingest path from robot to usable dataset: rosbag/MCAP capture, episode segmentation, format conversion, and delivery to training. * Implement data lifecycle guardrails - filtering, review-for-deletion, and retention - so idle-robot and failed-run data does not accumulate indefinitely. * Build and operate the dataset and mission registry so every dataset is attributable to a subject, session, robot, and purpose. * Support MoCapDB in production: ECS Fargate services, AWS Batch retargeting workers, RDS Postgres, and S3, integrated with FieldAI Auth. Cloud, IaC, and Security - 25% * Own FAIRI's AWS footprint as code: ECR, S3, IAM roles and cross-account trust policies, VPC and networking, Kubernetes/EKS workloads. * Close the gaps where infrastructure is not yet in code, and bring permissions changes under review. * Own compliance posture for research tooling - SOC 2 constraints on SaaS, experiment tracking, and data-sharing controls - in partnership with IT and Security. * Eliminate person-owned infrastructure: documented owners, runbooks, and access paths for every FAIRI-owned service. * Manage secrets, VPN/Tailscale access paths, and hardware-in-the-loop connectivity to robots on the floor. Enablement and Documentation - 15% * Write and maintain runbooks, onboarding guides, and architecture documentation so a new engineer can test and deploy on day one rather than learning it from a teammate. * Be the interface between FAIRI and Field AI's platform, cloud, and data-processing teams - negotiating what FAIRI reuses versus owns. * Support researchers and systems engineers directly when pipelines, builds, or environments break, including live troubleshooting during demos. * Bring reproducibility discipline to research workflows: versioned configs, pinned environments, traceable runs., * Comfortable being the only infra person in the room. You can take an open-ended ask and run with it without much hand-holding. * Bias toward reuse. You would rather integrate a platform team's runners than build a parallel stack, and you can negotiate that boundary well. * Strong documentation habits - you leave runbooks and processes better than you found them. * Pragmatic about research velocity. You know when to enforce rigor and when it would just slow the team down. What Sets You Apart * Infrastructure experience in robotics, autonomous vehicles, or ML research environments * ROS 2, rosbag/MCAP, or Foxglove familiarity * Large-scale data pipeline work - TB-scale sensor or video data, lifecycle and retention policy design * ML infrastructure: experiment tracking, GPU scheduling, training pipelines, simulation infrastructure (Isaac Sim / Isaac Lab) * Hardware-in-the-loop CI - running tests against physical devices from a pipeline * Compliance and audit experience: SOC 2, access reviews, data governance * Having been the first infrastructure hire on a team before * Interest in humanoid robotics and how machines learn from human movement ## Related Videos - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [Innovating Developer Tools with AI: Insights from GitHub Next](https://www.wearedevelopers.com/videos/1268-innovating-developer-tools-with-ai-insights-from-github-next) - [Why segmenting your infrastructure into tiers makes your infrastructure design better](https://www.wearedevelopers.com/videos/1960-why-segmenting-your-infrastructure-into-tiers-makes-your-infrastructure-design-better) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Bringing AI Model Testing and Prompt Management to Your Codebase with GitHub Models](https://www.wearedevelopers.com/videos/1536-bringing-ai-model-testing-and-prompt-management-to-your-codebase-with-github-models) - [Unleashing Potential Across Teams: The Power of Infrastructure as Code](https://www.wearedevelopers.com/videos/930-unleashing-potential-across-teams-the-power-of-infrastructure-as-code) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence](https://www.wearedevelopers.com/magazine/736-best-us-ai-conferences-for-ctos-in-2026-build-vs-buy-vendor-evaluation-and-peer-intelligence) - [DevOps Engineer Salary [2023]](https://www.wearedevelopers.com/magazine/203-devops-engineer-salary-2023) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)