> Markdown version of [/jobs/ext/2166111-data-engineer](https://www.wearedevelopers.com/jobs/ext/2166111-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** Accenture - **Location:** Tampa, FL, United States - **Salary:** $118,900.0 - $162,300.0 - **Contract:** Permanent contract - **Skills:** Amazon Web Services, Amazon S3, Big Data, Continuous Integration, Information Engineering, Extract Transform Load (ETL), Data Systems, Distributed Data Store, Distributed Systems, Elasticsearch, Monitoring of Systems, Identity and Access Management, Python (Programming Language), Query Optimization, Cloud Services, Data Streaming, Unstructured Data, Data Ingestion, Retrieval-Augmented Generation, Large Language Models, Indexer, Pandas, Pyspark, Kubernetes, Apache Nifi, Software Version Control, Data Pipelines, Docker - **Published:** August 21, 2026 - **Apply:** https://www.techcareers.com/job.asp?id=3361190580&tx=JT9284TYZ&pt=1&aff=0B19D771-A501-4A5E-8338-2A822B784D54&utm_source=Job%20Feed&utm_medium=textkernel&utm_campaign=DE&utm_term=0B19D771-A501-4A5E-8338-2A822B784D54 ## About the Role We are looking for a Data Engineer with strong hands-on experience designing, developing, and managing large-scale data workflows across structured and unstructured datasets. This role focuses heavily on building reliable RAG (Retrieval-Augmented Generation) pipelines, orchestrating ETL/ELT processes, and deploying scalable data systems in AWS., * Solid understanding of ETL/ELT processes and data modeling best practices. * Hands-on experience implementing workflows in Prefect (Prefect 2.0 preferred). In-depth knowledge of ElasticSearch indexing, cluster management, and search optimization. * Proficiency in Python and familiarity with common data libraries (Pandas, PySpark, requests, etc.). Preferred Qualifications * Strong experience with Apache NiFi for data flow management and real-time ingestion. * Experience building or maintaining RAG pipelines (e.g., vector databases, embeddings, document chunking strategies, retrieval optimization). * Proven ability to manage structured and unstructured data pipelines at scale. * Experience with AWS cloud services for data engineering. * Strong version control and CI/CD experience. * Experience with vector databases (OpenSearch, Pinecone, Weaviate, etc.) * Familiarity with containerized workflows (Docker, Kubernetes) * Experience supporting LLM or generative AI production systems * Background in distributed systems, streaming platforms, or data mesh architectures Clearance * An active TS/SCI federal security clearance is required ## Description * Design, build, and maintain RAG pipelines, including document ingestion, indexing, embedding workflows, and model retrieval flows. * Develop and manage structured and unstructured data pipelines supporting analytics, ML, and application workloads. Build and optimize ETL/ELT pipelines in AWS using services such as S3, Lambda, Step Functions, EMR, Glue, ECS/EKS, and IAM best practices. Implement and operate NiFi flows for high-throughput, low-latency data ingestion and transformation. * Develop, orchestrate, and schedule workflows using Prefect, ensuring reliability, observability, and proper error handling. Implement indexing, search, and retrieval patterns using ElasticSearch, including schema design, cluster management, and query optimization. * Collaborate closely with architecture, ML, and application teams to support scalable data solutions. * Ensure data quality, lineage, governance, and security across all pipelines. * Monitor system performance and troubleshoot issues across distributed data systems. ## Related Videos - [PySpark - Combining Machine Learning & Big Data](https://www.wearedevelopers.com/videos/44-pyspark-combining-machine-learning-big-data) - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Advanced Typing in TypeScript](https://www.wearedevelopers.com/videos/496-advanced-typing-in-typescript) - [Docker Compose: Rediscovered](https://www.wearedevelopers.com/videos/1978-docker-compose-rediscovered) - [Data Science in Retail](https://www.wearedevelopers.com/videos/586-data-science-in-retail) - [Empowering Retail Through Applied Machine Learning](https://www.wearedevelopers.com/videos/976-empowering-retail-through-applied-machine-learning) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Data Engineer Salary UK](https://www.wearedevelopers.com/magazine/253-data-engineer-salary-uk) - [Top Big Data Technologies That You Need to Know](https://www.wearedevelopers.com/magazine/108-top-big-data-technologies-that-you-need-to-know) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Coffee with Developers - Maria Apazoglou - Making AI understandable for all in production](https://www.wearedevelopers.com/magazine/475-coffee-with-developers-maria-apazoglou-making-ai-understandable-for-all-in-production) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline)