Software Engineer

Achira Inc.
United States
4 days ago
Apply on startup.jobs
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Amazon Web Services Microsoft Azure Cloud Computing Software Debugging Distributed Computing Environment Distributed Systems Job Scheduling Performance Tuning Tensorflow Google Cloud Pytorch Autoscaling
+8 more
Apache Spark Multi-Cloud Parallel Computation Kubernetes Performance Monitor Slurm Machine Learning Operations Celery

Job description

  • Architect & Build: Design, implement, and optimize distributed compute infrastructure for ML data processing, training, and fine-tuning.
  • Optimize & Monitor: Improve cluster observability, scheduling, and resource utilization (CPU/GPU/TPU).
  • Compute Efficiency: Research and implement cost-efficient compute solutions (spot instances, auto-scaling, multi-cloud strategies).
  • Tooling: Develop tools for monitoring, debugging, and performance tuning of large-scale ML workloads.
  • Collaboration: Collaborate with ML engineers to accelerate training pipelines and reduce bottlenecks.
  • Innovation: Stay current with emerging technologies in distributed computing (e.g., Ray, Kubernetes, Spark, Slurm) and apply them strategically.

Requirements

  • You are excited about and have lots of experience in building or working with distributed computing frameworks (e.g., Ray, Dask, Celery)
  • You have a good grasp of parallel computing, job scheduling, and resource management.
  • You’re comfortable identifying and resolving performance issues in distributed systems (profiling, bottlenecks, network overhead)
  • You’ve implemented solutions using cloud compute platforms (AWS, GCP, Azure) and cluster orchestration (Kubernetes, Slurm)
  • You are familiar with popular ML frameworks (PyTorch, TensorFlow, or JAX) and MLOps best practices such as model deployment and GPU performance monitoring

About the company

Why Achira

  • Join a world-class team of scientists, ML researchers, and engineers working together to reshape the future of drug discovery.
  • Work on cutting edge ML infrastructure at frontier scale: massive compute, massive data, and massive ambition.
  • Own impactful work end-to-end - from ideation to architecture to deployment on large-scale infrastructure.
  • Work in an environment that rewards rigor, speed, and a builder’s mindset., Achira is building best-in-class foundation models to solve the most challenging problems in simulation for drug discovery and beyond. Atomistic Foundation simulation models (FSMs) as world models of the physical microcosm span machine learning interaction potentials (MLIPs), neural network potentials (NNPs), and diverse classes of generative models.

We’re seeking a Software Engineer passionate about distributed computing and its applications in machine learning. You’ll have the opportunity to architect and build from the ground up the infrastructure for our ML data generation pipelines, model training, and fine-tuning workflows across large-scale distributed systems.

Your expertise will ensure our compute clusters are efficient, observable, cost-effective, and reliable-helping us push the boundaries of ML development. If you’re passionate about distributed systems, performance optimization, and cloud cost efficiency, we’d love to hear from you.

You’ll be empowered to eat, breathe, and think about the orchestration of complex workloads on multiple vendors scattered anywhere on the planet. Achira is a company which lives and breaths on computation, facile access at the lowest cost for our uniquely suited workloads is a mission critical endeavor.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on startup.jobs
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:35 min

Preventing remote code execution in PyTorch models

Balázs Kiss · World Congress 2023

1:37 min

Core concepts of Celery and message broker integration

Jan Giacomelli · LIVE

2:22 min

Infrastructure barriers and compliance risks in research

Jeremy Murray Jeremy Murray · World Congress 2026 Europe

4:36 min

Hiring passionate software engineers to tackle unprecedented scaling challenges

Dana Lawson Dana Lawson +1 · World Congress 2026 Europe

1:06 min

Compiling PyTorch environments for advanced time forecasting

Christoph Lohrmann Christoph Lohrmann +1 · World Congress 2026 Europe

2:10 min

Exploiting python celery dependencies for internal container access

Vandana Verma · LIVE

Videos

See all

Related articles

See all