INTL India - LLM Engineer
Role details
Job location
Tech stack
Job description
We are seeking a Data Engineer with strong AI expertise to design and maintain scalable data pipelines while integrating advanced language model applications. The role requires proficiency in Python, PySpark, Pandas, and web frameworks like FastAPI or Flask, along with hands-on experience in LangChain and building Retrieval-Augmented Generation (RAG) pipelines. You will be responsible for leveraging vector stores and retrievers to enable real-time document processing and querying, while deploying robust, scalable solutions on AWS to solve complex business challenges at the intersection of data engineering and applied AI., 1. Python Development
a. Design, develop, and maintain scalable web services using frameworks such as FastAPI.
b. Build and manage containerization solutions with Docker.
c. Develop and maintain CI/CD workflows leveraging tools like Jenkins, GitHub Actions, and ArgoCD.
d. Apply Test-Driven Development (TDD) practices, including unit and integration testing, to ensure code quality.
- Big Data Development
a. Design, develop, deploy, and maintain scalable data pipelines using Apache Spark/PySpark and Databricks
- LangChain & Generative AI Frameworks
a. Write efficient, reusable, and modular Python code to support API-driven LLM applications.
b. Implement LangChain to build custom pipelines for document indexing, retrieval, and summarization.
c. Integrate LangChain's Retrieval-Augmented Generation (RAG) capabilities with vector stores and retrievers to enable real-time querying and document processing.
- AWS Cloud Deployment:
a. Deploy and manage applications on AWS, leveraging services such as Lambda, EC2, S3, EKS, and RDS.
b. Ensure the scalability, availability, and reliability of deployed applications.
- Dashboards and Monitoring (Optional):
a. Create monitoring dashboards using tools like Grafana or Power BI for real-time system monitoring, analytics, and performance insights.
Requirements
- Proficiency in Python, with practical experience using web frameworks such as FastAPI.
- Advanced expertise in data technologies, including SQL, PySpark, and Databricks.
- Hands-on experience with Retrieval-Augmented Generation (RAG) pipelines, LangChain, and related Generative AI libraries and APIs.
- Experience deploying and managing containerized applications (Docker/Kubernetes) using CI/CD pipelines, adhering to best practices for automation, versioning, security, and reliability.
- Strong understanding of AWS cloud services, including EKS, EC2, RDS, and other core offerings.
Soft Skills
- Problem-solving: Ability to handle complex and dynamic challenges with AI solutions.
- Collaboration: Experience working in multidisciplinary teams (data scientists, DevOps, etc.).
- Adaptability: Eagerness and passion to keep up with the latest AI advancements and incorporate them into solutions.
- Communication: Excellent verbal and written communication skills to convey technical information to both technical and non-technical stakeholders.
This role is ideal for engineers who are passionate about pushing the boundaries of data driven applications with generative AI and have the technical skills to create cutting-edge, deployable solutions.