> Markdown version of [/jobs/ext/2198407-data-engineer-ii](https://www.wearedevelopers.com/jobs/ext/2198407-data-engineer-ii). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer II - **Company:** Liminex, Inc. - **Location:** United States (Remote available) - **Experience:** Experienced - **Salary:** $130,000.0 - $150,000.0 - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Airflow, Amazon S3, Data Analysis, Big Data, Code Review, Continuous Integration, Data Architecture, Information Engineering, Extract Transform Load (ETL), Data Warehousing, Distributed Data Store, Python (Programming Language), Machine Learning, Performance Tuning, Software Engineering, SQL Databases, Pandas, Pyspark, Core Data, Information Technology, Data Analytics, Integration Frameworks, Machine Learning Operations, Cloudwatch, Terraform, Software Version Control, Databricks - **Published:** August 23, 2026 - **Apply:** https://www.dice.com/job-detail/57283c3e-6051-4b0f-8971-aea3bc84ee95 ## About the Role The ideal candidate combines strong software engineering and data architecture skills with curiosity about machine learning systems and a drive to automate, optimize, and scale data workflows., * Bachelor's degree in Computer Science, Engineering, or related field. * 2-4 years of experience building and operating large-scale data systems, ideally supporting analytics and ML workloads. * Proficiency in Python and SQL, with experience in PySpark, pandas, or similar data processing frameworks. * Experience with DBT * Experience with modern data warehousing and lakehouse platforms, preferably Databricks. * Hands-on experience with workflow orchestration tools such as Airflow, Dagster, or Prefect. * Strong understanding of data modeling, ETL design, and distributed data systems. * Experience with AWS data and compute services (S3, Lambda, ECS, CloudWatch, etc.) or equivalent cloud platforms. * Familiarity with MLOps concepts (e.g., feature stores, model registries, CI/CD for ML). * Experience using Infrastructure as Code, preferably Terraform. * Excellent problem-solving, collaboration, and communication skills; comfortable working in a dynamic, fast-paced environment. ## Description We're looking for a Data Engineer II to help design, build, and continuously improve the GoGuardian Analytics and AI/ML ecosystem. This position sits on the Data Engineering team, a group responsible for building and maintaining the core data platform that powers analytics, product insights, and machine learning across the company. You'll collaborate closely with Data Science, Business Intelligence, and other teams to enable the next generation of data-driven products and AI capabilities., * Design, build, and optimize ETL pipelines that power analytics, data science, and ML workflows using tools such as Databricks, PySpark, and Airflow. * Develop and maintain labeling and retraining pipelines for machine learning models, ensuring quality, reproducibility, and observability. * Implement and support MLOps practices, including model versioning, CI/CD for ML, and model monitoring in production environments. * Collaborate with data scientists to productionize and scale model training, inference, and evaluation pipelines. * Contribute to the design and evolution of the data lakehouse, including schema design, partitioning strategies, and performance optimization. * Document and communicate data architecture, lineage, and dependencies to ensure transparency and maintainability across teams. * Champion data quality and governance, ensuring that datasets are accurate, well-structured, and compliant with organizational standards. * Leverage infrastructure-as-code and containerization to build reproducible, maintainable environments. * Participate in code reviews and continuous improvement of engineering best practices within the team. ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [Advanced Typing in TypeScript](https://www.wearedevelopers.com/videos/496-advanced-typing-in-typescript) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) ## Related Articles - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [MLops – Deploying, Maintaining And Evolving Machine Learning Models in Production](https://www.wearedevelopers.com/magazine/115-mlops-deploying-maintaining-and-evolving-machine-learning-models-in-production)