> Markdown version of [/jobs/ext/1787926-associate-data-engineer](https://www.wearedevelopers.com/jobs/ext/1787926-associate-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Associate Data Engineer - **Company:** Infoblox - **Location:** Tacoma, WA, United States - **Salary:** $81,500.0 - $117,370.0 - **Contract:** Permanent contract - **Skills:** Adobe InDesign, Artificial Intelligence, Data Analysis, Apple Mac Systems, Automated Storage and Retrieval Systems, Unit Testing, Bash Shell, Microsoft Basic Data Partition, Big Data, Software Quality, Information Engineering, Data Structures, Database Applications, Software Debugging, Linux, Python (Programming Language), Machine Learning, Object-Oriented Software Development, Software Engineering, Enterprise Data Management, Large Language Models, Apache Spark, Parallel Computation, Generative AI, Information Technology, Data Analytics, Data Management, Data Pipelines, Databricks - **Published:** July 17, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=1edd4e9365698451 ## About the Role * Experience with Python3 and Spark; Scala experience is helpful * Exposure to data engineering, data science, and related data-centric fields using large-scale data environments * Proficient in Object Oriented Design and S.O.L.I.D principles * Strong emphasis on unit testing and code quality * Experience with bash shell script on Linux or MacOS * Experience with async, threading, parallel programming * Exposure to machine learning, generative AI, or data platforms that support AI/ML workloads * Familiarity with modern AI technologies such as large language models (LLMs), vector databases, embeddings, or retrieval-augmented generation (RAG) concepts is a plus * BS in Computer Science or a related field, or equivalent work experience required Be Successful - Your Path First 90 Days: Immerse in our culture, connect with mentors (Blox Buddies), and map the systems and meet with key stakeholders that rely on your work. Discuss and create short/long term goals. Six Months: Focus on mastering the tech stack (Python, Spark, Databricks) and understanding data sources, workflows, and pipelines. Contribute to production-ready code by submitting PRs, adding unit tests, and refactoring or porting legacy scripts. Assist data scientists and analysts in curating large datasets and converting prototypes into production-ready solutions. Develop basic data quality monitoring scripts and demonstrate the ability to handle well-scoped tasks independently while actively participating in design discussions and sprint ceremonies. One Year: Own and deliver at least one end-to-end data pipeline, driving improvements in performance, reliability, or cost-efficiency. Implement advanced monitoring, automation, and schema management strategies for critical data sources. Consistently contribute high-quality, tested code while reviewing peers' work and providing constructive feedback. Act as a go-to engineer for specific data domains, mentor new hires, and propose innovative approaches to data modeling or summarization that enhance analytics capabilities and long-term platform scalability. ## Description We have an opportunity for an Associate Data Engineer to join our Data Engineering Team in Tacoma, reporting to the Director, Engineering. In this pivotal role, you will be at the heart of enabling data-driven insights across customer networks. Collaborating closely with data scientists and analysts, you will work to curate and deliver high-quality data that powers key business and product decisions. Be a Contributor - What You'll Do * Curate and manage large-scale data from multiple sources to support data science and threat analytics research and development * Design and implement data source monitoring using summarization, statistical methods, and anomaly detection * Apply computer science algorithms, including probabilistic data structures, to extract insights from large datasets * Productionize prototypes into scalable data engineering solutions using Spark optimizations and modern CI/CD pipelines * Build and maintain a robust research cloud environment with a focus on reliability, monitoring, and architecture improvements * Collaborate with software engineers on the design, implementation, and deployment of data-driven applications * Enable data scientists and analysts by supporting Spark application development, debugging, and deployment * Integrate large language models (LLMs) with existing data platforms and data sources to enable AI-powered analytics, automation, and user experiences * Build and maintain data pipelines, retrieval systems, and metadata services that support generative AI applications, including Retrieval-Augmented Generation (RAG) workflows * Develop and support integrations with Model Context Protocol (MCP) servers and AI tooling to enable secure and scalable access to enterprise data * Collaborate with data scientists, ML engineers, and software engineers to operationalize AI solutions and ensure high-quality data is available for model training, evaluation, and inference ## Related Videos - [Alibaba Big Data and Machine Learning Technology](https://www.wearedevelopers.com/videos/37-alibaba-big-data-and-machine-learning-technology) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Cutting LLM Costs Without Cutting Quality: How to Beat Proprietary LLMs with Fine-Tuned Open Source](https://www.wearedevelopers.com/videos/100151-cutting-llm-costs-without-cutting-quality-how-to-beat-proprietary-llms-with-fine-tuned-open-source) - [Docker network without Docker](https://www.wearedevelopers.com/videos/1418-docker-network-without-docker) - [Modern Data Architectures need Software Engineering](https://www.wearedevelopers.com/videos/1030-modern-data-architectures-need-software-engineering) - [OLTP in the Lakehouse: Redefining Data for AI Workloads](https://www.wearedevelopers.com/videos/2038-oltp-in-the-lakehouse-redefining-data-for-ai-workloads) ## Related Articles - [Navigating the AI Shift](https://www.wearedevelopers.com/magazine/629-navigating-the-ai-shift) - [How to Become an AI Engineer](https://www.wearedevelopers.com/magazine/331-how-to-become-an-ai-engineer) - [MLOps – What’s the deal behind it?](https://www.wearedevelopers.com/magazine/125-mlops-what-s-the-deal-behind-it) - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering)