Software Engineer, Machine Learning

Etsy
New York, NY, United States
17 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
5 years minimum
Compensation
$182,000.0 - $246,000.0
Working hours
Regular working hours
Job source

Tech stack

Artificial Neural Networks Cloud Computing Code Review Software Debugging Distributed Systems Python (Programming Language) Machine Learning Language Modeling Open Source Technology Pair Programming Performance Tuning Datadog
+10 more
Data Logging Google Cloud Large Language Models Deep Learning Backend Kubernetes Information Technology Production Code Etsy Machine Learning Operations

Job description

We are looking for a Senior Software Engineer, Machine Learning to help build and operate the infrastructure that trains, hosts, and serves Etsy’s machine learning models - including our internal serving platform, and a growing portfolio of predictive and generative (open source LLM) models.

You won’t just be deploying models; you’ll be architecting the systems that make model serving fast, reliable, and scalable for millions of inferences per second. Your work directly enables Applied Scientists and Engineers across Etsy to train, host, and iterate on models with confidence, from classic ML predictors to large open-source language models.

This is a full-time position in ML Enablement team, reporting to its Engineering Manager.

What’s this team like at Etsy?

  • The Machine Learning Enablement initiative builds the core infrastructure that turns complex ML workflows into seamless, self-service platforms for Etsy’s Applied Scientists and Engineers.
  • The Models Training and Serving Platform team builds and operates the core infrastructure that powers model training and serving across Etsy, including Barista, our internal model-serving platform, along with a range of hosted modelsets and open-source LLMs.
  • We own the path from a trained model to a production-ready, observable, and scalable serving endpoint - spanning infrastructure, deployment tooling, and reliability.
  • We work on significant, complex challenges that intersect multiple critical teams and systems, where you can make a rewarding impact.
  • We are light on process and heavy on collaboration, working with many partner teams within our org and beyond in order to improve our leverage.
  • We are a platform team with a product driven mentality, driving innovation using our Machine Learning systems in effective and creative ways.

What does the day-to-day look like?

  • Write high-quality, production-grade Python code (additional languages a plus); participate in code reviews and pair programming; contribute to and help drive architectural decisions.
  • Build, operate, and improve infrastructure for training and serving ML models on GCP and Kubernetes, with a focus on scalability, reliability, and observability.
  • Design and maintain serving infrastructure for open-source and internally hosted LLMs, including model deployment, resource management, and performance tuning.
  • Apply working knowledge of ML fundamentals - including neural network deep learning as well as latest transformer architectures along with prediction and inference systems - to make sound infrastructure and design decisions.
  • Partner cross-functionally with Applied Scientists to understand model training and serving needs, and translate that understanding into infrastructure that removes friction from their workflows.
  • Contribute to system design discussions, weighing tradeoffs across performance, cost, and reliability for infrastructure serving 100M+ users.
  • Thoughtfully use generative AI and other productivity tools to work efficiently, with a focus on learning and intentional contribution.
  • Of course, this is just a sample of the kinds of work this role will require! You should assume that your role will encompass other tasks, too, and that your job duties and responsibilities may change from time to time at Etsy’s discretion, or otherwise applicable with local law.

Requirements

  • Bachelor’s degree in Computer Science, Applied Statistics, Mathematics, Electrical Engineering, or a related quantitative field, or equivalent professional experience.
  • 5+ years of professional experience building, iterating on, and troubleshooting complex backend and infrastructure systems.
  • Strong software engineering fundamentals, including solid command of algorithms and data structures, with the ability to write production-ready code in Python.
  • Hands-on experience with cloud infrastructure (Google Cloud preferred) and Kubernetes, including deploying and operating production workloads.
  • Familiarity with observability tooling (metrics, logging, tracing) for monitoring and debugging distributed systems.
  • Working knowledge of machine learning fundamentals and concepts, with awareness of LLM serving infrastructure and hosting open-source models.
  • Basic understanding of transformer architectures, predictors, and how model design choices affect serving infrastructure.
  • Comfort with system design for large-scale, high-availability infrastructure.

Benefits & conditions

In addition to salary, you’ll be eligible for an equity package, an annual performance bonus, and our competitive benefits that support you and your family. Base salary is determined by your location, skills, and experience.

About the company

Etsy is the global marketplace for unique and creative goods. We build, power, and evolve the tools and technologies that connect millions of entrepreneurs with millions of buyers around the world. As an Etsy employee, you will tackle unique, meaningful, and large-scale problems alongside passionate coworkers, all the while making a rewarding impact and Keeping Commerce Human.

We believe exceptional companies are built by exceptional teams, and we’re intentional about it. That means hiring great people, setting them up for success from day one, and giving them real reasons to grow their careers here. We invest in development that goes beyond promotions, and we foster the trust and relationships that help people do their best work together.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

58 sec

Real-world application of MLOps architectural patterns at Wayfair

Bas Geerdink · LIVE

1:36 min

Visualizing memory limits and isolating suspicious endpoints

Dina Matveev Dina Matveev · Europe 2026 Virtual

2:32 min

Refactoring bulk frontend operations into scalable backend methods

Noam Honig · LIVE

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

1:08 min

Analyzing error logs and root causes using artificial intelligence

Nishil Patel Nishil Patel · World Congress 2025

1:12 min

Choosing TypeScript for complex backend applications

Maximilian Otto Maximilian Otto · World Congress 2024

Videos

See all

Related articles

See all