Big Data/Machine Learning Engineer
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
+46 more
Job description
Data Engineering & Big Data Architecture:
- Design, build, and maintain high-throughput streaming and batch data processing pipelines to support training data preparation and feature generation at scale.
- Manage big data storage, transformation, and querying frameworks across multi-terabyte dataset systems.
Machine Learning & Feature Engineering:
- Implement feature engineering processes and feature store architectures ensuring training-serving consistency.
- Convert research-level ML prototypes into optimized, scalable production code.
- Fine-tune, benchmark, and scale traditional ML, Deep Learning, or Generative AI models for real-time and batch inference.
MLOps, Deployment & Infrastructure:
- Build and maintain robust MLOps framework tools, continuous training (CT) pipelines, and CI/CD pipelines for automated deployment.
- Expose ML models via REST/gRPC microservices, ensuring low-latency inference and target availability SLA compliance.
- Implement comprehensive model monitoring tools for data drift, concept drift, system performance, and output quality checks.
Technical Leadership & Governance:
- Mentor junior and mid-level engineers, enforcing software engineering best practices across data and ML codebases.
- Ensure system compliance with enterprise security, data privacy, and AI governance standards.
Requirements
Experience: 5+ years of software engineering, big data engineering, and applied machine learning engineering experience in production environments.Programming Languages: Advanced proficiency in Python and Scala or Java. Big Data Frameworks: Hands-on experience with Apache Spark (PySpark/Scala), Kafka, Flink, Hadoop, Databricks, or Snowflake.Machine Learning Frameworks: Proficiency with frameworks like PyTorch, TensorFlow, scikit-learn, or XGBoost. MLOps & DevOps: Experience with MLflow, Kubeflow, Feature Stores (e.g., Feast, Hopsworks), Docker, Kubernetes, and CI/CD automation.Cloud Architecture: Deep experience with AWS (S3, SageMaker, EMR), Google Cloud Platform (BigQuery, Vertex AI, Dataflow), or Azure (Databricks, Synapse).Databases & Storage: Mastery of SQL, NoSQL databases (Cassandra, MongoDB), and Vector Databases (Pinecone, Milvus, Qdrant). Education: Bachelor’s or Master’s degree in Computer Science, Data Science, Software Engineering, or a related quantitative field.
skills:
S3,experience with AWS,Flink,Hadoop,Kafka,Apache Spark,AI,Big Data,BigQuery,Cassandra,Cloud Architecture,Dataflow,data quality frameworks,data processing pipelines,Databases,Databricks,Deep Learning,DeepSpeed,automated deployment,DevOps,distributed data,Docker,concept drift,feature generation,Feature Engineering,Feature Stores,Generative AI models,Generative AI,Great Expectations,Data Engineering,Computer Science,Java,Kubeflow,Kubernetes,Large Language Models,LLMs,Machine Learning,ML models,applied machine learning,ML infrastructure,MLOps,MLflow,Megatron,microservices,Azure,Milvus,MongoDB,NoSQL databases,Ollama,Pinecone,production code,Programming Languages,PySpark,proficiency in Python,PyTorch,Qdrant,RAG,SQL,scikit-learn,Snowflake,software engineering best practices,software engineering,Machine Learning Frameworks,TensorFlow,training data,XGBoost,resilient,data drift,Architecture,automation,continuous training,enterprise security,prototypes,data privacy,Deequ,Data Science,Vertex AI,Governance,AI governance,Infrastructure,system compliance,Mentor,model training,production systems,quality checks,system implementation,Technical Leadership,Vector Databases
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Loading talks and stories from around this role…