Senior Engineer, AI Data Management

Tata Consultancy Services Limited
Seattle, WA, United States
2 months ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Expert
Compensation
$124,300.0 - $168,100.0
Working hours
Regular working hours
Job source

Tech stack

Microsoft Excel Application Programming Interfaces (APIs) Agile Methodology Artificial Intelligence Airflow Amazon S3 Data Analysis JIRA Big Data Cloud Computing Cloud Storage Collaborative Software
+57 more
Computer Programming Databases Continuous Integration Data Architecture Data Validation Data Transmissions Information Engineering Data Integration Extract Transform Load (ETL) Data Transformation Serialization Data Stores Data Visualization Relational Databases Software Debugging DevOps Distributed Computing Environment Apache Hadoop Hadoop Distributed File System Monitoring of Systems JSON Python (Programming Language) Network Security Unix Shell MongoDB NoSQL Parsing Raw Data Standard Sql Shell Script Simple Data Format SQL Databases Data Streaming Web Services Extensible Markup Language (XML) Jupyter Notebook Data Processing Scripting Data Ingestion Azure Data Factory Informatica Powercenter Delivery Pipeline Apache Spark Git Documentation System Pytest Cassandra Real Time Data Apache Kafka Data Management Virtual Agents Api Design Restful APIs Stream Processing Software Version Control Data Pipelines Web Api

Job description

The Agentic AI Data Engineer is a hands-on role focused on building and maintaining the data pipelines and infrastructure that fuel AI agent systems. Within TCS’s AI & Data group (Americas), you will be the builder who turns data architecture plans into reality, ensuring that AI models and agents have continuous access to high-quality, timely data. This client-facing consulting role involves hybrid work from client site as needed for deployment. You’ll work across wide array of business functions within Retail. By combining expertise in data ingestion, transformation, and integration with knowledge of AI data needs, you will play a critical part in enabling AI agents to perform reliably and accurately in production. What You Would Be Doing: Build Data Ingestion Pipelines: Develop robust pipelines to extract data from various sources (databases, APIs, flat files, streaming sources) relevant to the AI solution. Data Transformation & Processing: Implement transformation and cleaning steps on raw data to make it usable for AI, ensuring efficiency and scalability. Loading Data to Storage/Indices: Set up processes to load processed data into target storage systems that AI agents or models will use. Real-Time Data Feeds: Implement streaming or incremental update pipelines when AI systems require real-time or frequently updated data. Pipeline Automation & Scheduling: Use orchestrators or schedulers to automate the data workflows. Data Integration & API Development: Develop and maintain integration components for real-time data fetching. Collaborate on RAG/Knowledge Base Updates: Work closely with AI Data Architects on implementing RAG updates. Testing and Validation of Data Pipelines: Develop tests and monitoring for your data pipelines. Optimize Pipeline Performance: Profile and optimize data pipelines for speed and resource usage. Documentation and Handover: Document pipeline processes, configurations, and dependencies clearly. Industry-Specific Data Handling: Adapt data engineering to specific domain needs. Collaboration & Agile Implementation: Work as part of an agile product team, collaborating with data architects, AI engineers, and others. Maintain and Evolve Pipelines: Monitor pipelines and handle maintenance post go-live.

Requirements

Do you have experience in Quality issues?, Programming & Scripting: Strong programming skills, especially in Python, and experience with other languages like SQL. Data Pipeline Development: Practical experience building data pipelines end-to-end. Database and SQL Skills: Proficiency in writing and optimizing SQL queries. Big Data & Distributed Processing: Experience with big data technologies like Apache Spark. Streaming Data Experience: Familiarity with streaming frameworks and tools like Kafka. API Integration and Web Services: Ability to interact with web APIs for data ingestion or extraction. Data Formats and Parsing: Strong understanding of data formats and ability to parse JSON, XML, or custom text formats. DevOps for Data Pipelines: Basic DevOps skills, including using Git for version control and CI/CD pipelines. Problem Solving & Debugging: Strong ability to troubleshoot data issues. Data Quality Focus: Attentiveness to data quality and skills in implementing checks and validating outputs. Collaboration & Commun ication: Good communication skills to work with the team and clients. Time Management & Flexibility: Ability to handle multiple tasks and prioritize effectively. Domain Data Understanding: Aptitude to learn domain context from data. Security & Privacy Business Units: Understanding of handling sensitive data securely in pipelines. Continuous Learning: Willingness to learn new tools or frameworks as needed. Key Technology Capabilities: ETL / Data Integration Tools: Experience with tools such as Apache Airflow, Informatica PowerCenter, or cloud-based ones like Azure Data Factory. Big Data Processing: Proficiency in Apache Spark and knowledge of Hadoop HDFS. SQL & Databases: Strong practical SQL skills and familiarity with relational database systems. NoSQL and Other Data Stores: Knowledge of specific systems like MongoDB or Cassandra. Stream Processing: Hands-on usage of Apache Kafka and understanding of consumer group mechanics. Cloud Storage & Compute: Familiarity with cloud storage services like Amazon S3 and cloud compute for ETL. APIs & Web Services: Experience building or using connectors to RESTful APIs. File Formats & Data Serialization: Understanding of various file formats and ability to convert between them. Operating Systems & Scripting: Comfortable with Linux shell and basic shell scripting. Version Control & CI/CD: Using Git for source control and setting up CI pipelines for data projects. Monitoring & Logging Tools: Utilizing monitoring tools for data workflows. Data Visualization/Verification: Basics of tools like Excel or Python’s Jupyter notebooks for data sanity checks. Security & Networking: Understanding network configurations for data transfer. Testing Frameworks: Familiarity with PyTest or unittest for writing tests for data transformations. Collaboration Tools: Experience with tools like JIRA and documentation tools.

Benefits & conditions

(part of Tata group) 3.93.9 out of 5 stars Seattle, WA Hybrid work $124,300 - $168,100 a year

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on indeed.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

2:57 min

Core technical practices for robust data engineering

Sandhya Menon Sandhya Menon · WWC Europe 2026

6:21 min

Investigating push inefficiencies with upstream Git experts

Jonathan Creamer · Coffee With Developers

3:47 min

Exploring JSON, CBOR, and JOSE for data serialization

Aaron Russell · LIVE

3:05 min

Integrating an assistant application with Jira software

Felix Augenstein · LIVE

6:08 min

Applying software engineering environments and testing to data pipelines

Matthias Niehoff Matthias Niehoff · WWC 2024

56 sec

Favorite git commands and the importance of patch commits

Eileen Uchitelle Eileen Uchitelle +1 · Coffee With Developers

Videos

See all

Related articles

See all