Associate Data Engineer 2027 - AI & Data Analytics

IBM
San Francisco, CA, United States
3 days ago
Apply on www.juju.com
Prepare application

Role details

Contract type
Internship / Graduate position
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Java (Programming Language) JavaScript (Programming Language) Application Programming Interfaces (APIs) Artificial Intelligence Airflow Amazon Web Services Data Analysis IBM System I Automated Storage and Retrieval Systems Microsoft Azure Cloud Computing Cloud Engineering
+46 more
Computer Programming Databases Continuous Integration Data as a Services Data Architecture Data Cleansing Information Engineering Data Governance Data Infrastructure Data Integrity Extract Transform Load (ETL) Data Transformation Query Languages Linux Distributed Computing Environment Apache Hadoop IBM Cloud Computing Information Retrieval Python (Programming Language) Machine Learning Scala (Programming Language) Search Technologies Software Engineering SQL Databases Unstructured Data Google Cloud Real Time Systems Large Language Models Snowflake Multi-Agent Systems Prompt Engineering Apache Spark Model Validation Generative AI Git Data Lakes AI Platforms Kubernetes Information Technology Data Analytics Apache Kafka Data Management Virtual Agents Cloud Optimization Data Pipelines Databricks

Job description

Ready to think boldly, work with some of the world’s most recognized brands, and kick-start your career? Welcome to IBM’s Associate Program for university hires. From day one, you will collaborate with global clients and IBM teams on projects that help organizations solve tough challenges across digital transformation, cloud strategy, AI adoption, process redesign, analytics, modern data platforms, and agentic AI-enabled data transformation.

As an Associate, you will work alongside a global cohort of diverse, ambitious peers and have access to industry-recognized certifications, digital badges, and a minimum of 40 hours of structured learning per year on IBM’s AI-driven learning platform, supported by coaches, mentors, and professional communities across practices. IBM’s culture of internal mobility means you can explore new technologies, industries, and career paths as your interests evolve.

Bring your curiosity. Grow your skills. Build what’s next - like an IBMer. This role is a strong fit for builders who like turning messy data into reliable systems and who want to help make analytics, generative AI, and AI agents useful, reliable, and safe through governed data, trustworthy context, observable pipelines, and secure interfaces to enterprise systems.

To give yourself the best opportunity for success, we advise applying only to roles that align with your skills and experience, rather than applying broadly across all entry-level positions. You’ll receive a status update email for each application, so be sure to check your IBM Careers account regularly - it’s the best way to get a centralized view of which roles you have active applications against.

Your role and responsibilities

As an Associate Data Engineer, you will help design, build, and improve data platforms, data products, and services that support analytics, machine learning, generative AI, and agentic AI solutions. You will work across data gathering, ingestion, transformation, storage, batch and real-time processing, semantic enrichment, retrieval, APIs, visualization, data quality, observability, and governance.

Collaborating closely with diverse teams, you will play an important role in selecting suitable data management systems and identifying the critical data needed for insightful analysis. As a Data Engineer, you will help tackle challenges related to database integration, data quality, and complex structured and unstructured datasets.

Key Responsibilities may include:

  • Assisting designing and implementing scalable data architecture, data products, and management systems for modern cloud environments, analytics, and AI-enabled use cases.

  • Work on optimizing existing data pipelines, retrieval indexes, and data services for improved performance, reliability, data quality, and freshness of AI-ready context.

  • Collect, prepare, and analyze structured, semi-structured, and unstructured data to identify trends, providing clients with actionable insights that enhance marketing, operational, and business practices.

  • Participate in troubleshooting data-related issues, working to solve data quality challenges, retrieval-quality gaps, processing failures, and inconsistencies affecting analytics, generative AI, or agentic workflows.

  • Create visually compelling and user-friendly dashboards, reports, and observability views to communicate findings, pipeline health, and AI system insights to both technical and non-technical stakeholders.

  • Ensure data integrity, accuracy, reliability, lineage, and access control through rigorous data cleaning, validation, preprocessing, cataloging, and governance practices.

  • Work with project teams to prioritize and translate client requirements into current and future operational scenarios, processes, models, use cases, data products, APIs, tool interfaces, plans, and solutions; collaborate with clients, architects, and AI engineers.

  • Present analytical findings, data quality insights, and recommendations clearly and concisely, demonstrating the value of data-driven and AI-enabled decision-making to clients.

  • Work with cross-functional teams to tackle complex business problems, utilizing data expertise across cloud platforms, RAG, vector search, APIs, responsible AI, and agent orchestration patterns while staying current on modern data stack trends.

Requirements

  • High School Diploma/GED.

  • Familiarity with one or more programming or query languages such as Python, SQL, Java, Scala, or JavaScript.

  • Foundational understanding of data engineering concepts, including data pipelines, databases, APIs, distributed processing, ETL/ELT, data modeling, or data products.

  • Basic understanding of cloud computing environments such as AWS, Azure, Google Cloud, IBM Cloud, or similar platforms.

  • Ability to apply foundational statistical, machine learning, or information retrieval concepts to data preparation, analysis, search, or model-support work.

  • Interest in AI/ML, generative AI, agentic AI, intelligent automation, RAG, LLM-powered systems, or enterprise AI platforms.

  • Strong analytical thinking, problem solving, adaptability, teamwork, and communication skills.

  • Willingness to travel up to 100%, based on project requirements.

Preferred technical and professional experience

  • Bachelor’s degree in a related field such as Computer Science, Data Science, Statistics, Mathematics, MIS, Engineering, AI/ML, or another quantitative field.

  • Coursework, projects, internship experience, or portfolio work involving data engineering, software engineering, cloud platforms, analytics, AI/ML, or modern data products.

  • Experience with technologies such as Spark, Hadoop, Kafka, Airflow, dbt, Databricks, Snowflake, Delta Lake, Linux, or similar platforms.

  • Familiarity with LLMs, embeddings, vector databases such as Pinecone, Weaviate, orpgvector, retrieval systems, prompt workflows, RAG, and AI agent architectures.

  • Exposure to orchestration and agentic AI frameworks such asLangChain,LangGraph,LlamaIndex, Semantic Kernel, AutoGen, MCP-based tooling, or similar tools.

  • Understanding of Git, containers, APIs, testing frameworks, CI/CD, Kubernetes, observability, data governance, privacy, security, or responsible AI practices.

  • Preferred certifications, coursework, or credentials such asSnowProCore, Google Associate Cloud Engineer, Google Data Engineer coursework, or related cloud/data credentials.

  • Familiarity with AI coding tools and GenAI concepts, including prompt engineering, RAG, fine-tuning, and model evaluation.

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.juju.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

52 sec

Running persistent Linux environments directly on Windows

Ben Breard Ben Breard · World Congress 2025

2:15 min

Empowering domain teams with an open data platform

Sandhya Menon Sandhya Menon · World Congress 2026 Europe

3:03 min

Career evolution in data engineering and AI platforms

Maria Apazoglou · Coffee With Developers

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

3:14 min

Structuring career paths and localized data architectures

Ulrich Wurstbauer +1 · LIVE

Videos

See all

Related articles

See all