Senior Software Engineer, MLOps

Rivian
Palo Alto, CA, United States
3 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
2 years minimum
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) Agile Methodology Amazon Web Services Amazon S3 Application Performance Management Cloud Computing Cloud Engineering Configuration Management Continuous Integration Information Engineering Data Infrastructure Software Debugging
+19 more
Distributed Systems Fault Tolerance Python (Programming Language) Linux Kernel Software Tools Prometheus DataOps Software Engineering Datadog AWS Cdk Docker Swarm AWS ECS Gitlab Kubernetes Machine Learning Operations Cloudwatch Terraform Jenkins Microservices

Job description

The Autonomy org at Rivian is seeking a Staff Software Engineer, Data Ops to join the Data team who can provide expertise in cloud and data engineering and collaborate with technical and business users. This candidate needs to have a very good understanding of the AWS Cloud Data Platform and Data Ops processes that helps to build, test, and release complex mission critical infrastructure services for Rivian’s ADAS team on AWS cloud. In this role you will work with the ADAS Cloud, Data, Perception, SIL/HIL, Vehicle integrations & Vehicle Cloud teams, Product Management, and other Technology Partners to leverage best practices and reference architectures highlighting AWS Cloud Platform and Data/Dev/ML Ops practices.

  • Lead, build, test and release complex mission-critical infrastructure services for Rivian’s ADAS team on cloud and/or on-prem.
  • Setup fault tolerant multi-region environments for data operations and data applications.
  • Own CI/CD pipeline for apps and data projects.
  • Define on-call strategy and participate in on-call rotations.
  • Make developers’ lives smooth via automated workflows.
  • Build and optimize highly reliable, scalable, and distributed infra using microservice architecture.
  • Collaborate with the security & privacy team to perform audits and mitigate any findings.
  • Collaborate with cross-functional ADAS teams for development and integrations.
  • Cost optimization in AWS across multiple accounts and services.

Requirements

Do you have experience in S3?, * 5+ years of software engineering or in ML/Dev/Data Ops role.

  • 5+ years of experience authoring, scaling, and managing production infrastructure.
  • 5+ years of experience working with Kubernetes, AWS CI/CD tools, AWS networking stack, S3, Lambda, EKS, ECS, RDS, System Manager, Secrets Manager, CloudTrail, etc.
  • 5+ years Infra as Code and configuration management (Terraform, AWS cloud formation, AWS CDK).
  • 5+ years of experience with monitoring applications in cloud using Datadog, AWS CloudWatch or Prometheus.
  • 5+ years of eperience debugging production systems and performing RCA on incidents.
  • 3+ years of being hands-on with Python, Go or Java and Gitlab for automation.
  • 2+ years of CI/CD and/or GitOps patterns (using Gitlab, Jenkins, Allure etc.)
  • 2+ years of microservice-oriented architectures (using Kubernetes (EKS), AWS ECS, or Docker swarm).
  • 2+ years of knowledge of Agile Development of Accessible Software Tools.
  • Linux internals, networking, and distributed computing are a plus.
  • AWS or Cloud Native certification is a plus.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:14 min

Structuring CI/CD pipelines with integrated security and quality checks

Christoph Ruggenthaler · LIVE

1:02 min

Applying an ETL methodology to infrastructure configuration management

Axel Barbier · WWC 2023

1:52 min

Deploying Celery workers on AWS ECS Fargate containers

Jan Giacomelli · LIVE

2:44 min

Defining core roles and responsibilities in MLOps teams

Bas Geerdink · LIVE

4:54 min

Implementing geographic salary tiers for compensation equity and fairness

Rudi Bauer Rudi Bauer +1 · Cappuccino with HR

5:28 min

Defining MLOps and its role in production systems

Hauke Brammer · WWC 2023

Videos

See all

Related articles

See all