Senior ML Engineer

CloudBolt Software
Rockville, MD, United States
9 days ago
Apply on www.indeed.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Experience required
3 years minimum
Compensation
$170,000.0 - $220,000.0
Working hours
Regular working hours
Job source

Tech stack

Algorithm Design Amazon Web Services Amazon S3 Data Transformation Java Virtual Machine (JVM) Python (Programming Language) Machine Learning Regression Testing Prometheus Software Engineering Autoscaling Prophet
+7 more
Deep Learning Numerical Computing Kubernetes Information Technology Production Code Machine Learning Operations Cloud Optimization

Job description

  • Own the recommendation engine end to end: model selection, algorithm design, preprocessing, and the guardrails that keep recommendations safe to apply to live production workloads.
  • Design, evaluate, and productionize time-series forecasting and statistical models (e.g., Prophet, percentile-based estimation) that right-size Kubernetes workloads across CPU, memory, GPU, and JVM heap.
  • Build and maintain the data-quality layer: detecting and filtering anomalies, load-test windows, startup spikes, and autoscaling artifacts from production telemetry before it reaches a model.
  • Define and continuously improve how we measure recommendation quality: regression testing against golden datasets, behavioral validation, and accuracy/safety metrics in production.
  • Investigate and resolve recommendation quality issues reported from customer environments, tracing them through data, preprocessing, and model behavior.
  • Serve as the team’s machine learning authority: guide technical direction on ML questions, make model-vs-heuristic tradeoff calls, and clearly communicate them to platform engineers, product, and leadership.
  • Write production-grade Python for models and pipelines alike and share ownership of the surrounding service (message consumption, metrics ingestion, caching) with the rest of the team.
  • Prototype and validate new optimization capabilities (new resource types, new algorithms, new workload classes) from research through gradual, feature-flagged rollout.
  • Stay current on time-series forecasting and resource optimization techniques, and pragmatically evaluate which are worth adopting.

Requirements

Must have strong experience…?

  • Master’s degree or higher in a quantitative field (Computer Science, Machine Learning, Statistics, Applied Mathematics).
  • 5+ years of software engineering experience, with at least 3 years building and operating machine learning or statistical systems in production.
  • Expert-level Python: you write typed, tested, production-grade code, and you’re fluent in numpy or similar array-based numerical computing.
  • Hands-on experience with time-series analysis and forecasting: seasonality, trend decomposition, anomaly detection, and classical statistical methods (percentiles, distributions, smoothing), not just deep learning.
  • Experience testing ML systems rigorously: regression testing against known-good baselines, behavioral validation, and reasoning about numerical reproducibility.
  • Working knowledge of Kubernetes: resource requests and limits, autoscaling behavior, and what happens to a workload when it’s under-provisioned (OOM kills, CPU throttling).
  • Comfort owning a production service, not just a model: queues, caches, retries, observability, and debugging issues in customer environments from logs and metrics.
  • Clear written and verbal communication: as an ML engineer on the team, so you must be able to explain model behavior and tradeoffs to platform engineers, product managers, and customers.

Experience in the following is beneficial

  • Experience with Prophet or similar forecasting libraries.
  • Prometheus/PromQL and experience working with metrics at scale.
  • Cloud cost optimization, capacity planning, or infrastructure efficiency background
  • AWS (S3, Managed Prometheus).
  • Experience being the ML domain expert on a team of generalists.

Benefits & conditions

Our US benefits package includes:

  • Medical/Dental/Vision coverage
  • 401k with Company Match
  • Health & Dependent Care FSA
  • Unlimited PTO
  • 11 Company Holidays
  • Volunteer/Community Engagement Day
  • Tuition Reimbursement
  • Paid Parental Leave
  • Equity Grants
  • Home internet Reimbursement

Base Salary Range: $170,000 - $220,000 USD Annually, actual compensation will be determined based on job-related factors, including but not limited to relevant experience, skills, geographic location, internal equity, and business needs.

About the company

At CloudBolt we help organizations maximize the value of their cloud investments through greater visibility, governance, and cost optimization across complex cloud environments. Our platform empowers teams to make smarter cloud decisions by turning insights into action, helping businesses improve efficiency, control spending, and accelerate innovation.

As a remote-first global SaaS company, CloudBolt is committed to fostering a collaborative, inclusive, and high-performing culture where employees can do their best work and make a meaningful impact.

Learn more at www.cloudbolt.io.

As a Senior ML Engineer, you’ll be building and owning the recommendation engine at the heart of our Kubernetes resource optimization product, StormForge, cutting our customers’ cloud spend without putting a workload at risk. This is a critical hands-on position at the intersection of applied machine learning and production engineering, where the hardest problems are as much about data quality, guardrails, and knowing when not to recommend as they are about forecasting itself. You’ll need to bring rigor and curiosity in equal measure, designing time-series models and the fallback strategies around them, proving their behavior through repeatable regression testing, and iteratively raising the ceiling on accuracy and safety as we discover and learn more about the workloads our customers run.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.indeed.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · World Congress 2026 Europe

3:43 min

The enduring legacy of the amazon S3 storage API

Chris Heilmann +3 · LIVE

2:28 min

Understanding Kubernetes architecture and core cluster components

Marc Nimmerrichter · World Congress 2022

3:09 min

Balancing data science skillings alongside systems engineering rigor

Nico Schmidt · LIVE

3:44 min

Automating storage savings with S3 intelligent tiering

Sébastien Stormacq · World Congress 2021

4:04 min

Overview of Kubernetes operators and custom resource definitions

Philipp Krenn · World Congress 2022

Videos

See all

Related articles

See all