> Markdown version of [/jobs/ext/3409134-data-engineer](https://www.wearedevelopers.com/jobs/ext/3409134-data-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Data Engineer - **Company:** SimilarWeb LTD - **Location:** United States - **Experience:** Experienced - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Amazon Web Services, Big Data, Cloud Computing, Python (Programming Language), Machine Learning, Retrieval-Augmented Generation, Large Language Models, Prompt Engineering, Apache Spark, Generative AI, Pyspark, Information Technology, HuggingFace, Document Classification, Databricks - **Published:** September 30, 2026 - **Apply:** https://job-boards.greenhouse.io/similarweb/jobs/8239378 ## About the Role 1. Holds a B.Sc. or M.Sc. in Computer Science, Data Science, Mathematics or another relevant field 2. Has 4+ years of hands-on experience as a data engineer, ML engineer or data scientist, with solutions running in production 3. Has strong Python skills and writes production-quality code 4. Has hands-on experience building LLM-based applications in production (prompt engineering, structured outputs, RAG, embeddings, evaluation) 5. Has worked with the modern LLM stack: LLM provider APIs (OpenAI, Anthropic, etc.), LangGraph or LangChain, Hugging Face and vector stores 6. Has a solid grasp of text classification and NLP methods, both classical and modern, and knows when to use each 7. Has experience processing large-scale data with Spark/PySpark, Databricks or similar, on AWS or another cloud 8. Understands evaluation and data quality well: precision/recall trade-offs, building ground truth, error analysis 9. Is pragmatic and delivery-focused, comfortable with ambiguity, and communicates clearly with Product and business stakeholders 10. Has experience with taxonomies, entity resolution or product/e-commerce data (advantage) 11. Has experience with fine-tuning or deploying open-source models (advantage) At Similarweb, collaborating with our colleagues in-office creates a more connected, unified culture. Our best work is a product of our face-to-face collaboration, with the ability to work partially from home. ## Description Your daily responsibilities may include: * Designing and building LLM-powered and ML-based pipelines that classify, normalize, structure and match product, brand and category data * Building agentic workflows (LangGraph or similar) that automate complex data tasks end to end * Choosing the right approach for each problem (LLMs, embeddings, fine-tuned models, classical classifiers or rules), balancing accuracy, cost and latency * Scaling solutions to run efficiently over billions of records, using Spark, Databricks and our cloud infrastructure * Building evaluation frameworks: ground-truth datasets, labeling processes, quality metrics and ongoing monitoring * Taking solutions from POC to production, and owning them after launch * Working closely with Product to define requirements and shape the roadmap * Collaborating with data engineers, data scientists and other R&D teams on infrastructure and best practices