GenAI Engineer

QTech US, Inc
New York, NY, United States
21 days ago

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Working hours
Regular working hours
Job source

Tech stack

Adobe Flash Application Programming Interfaces (APIs) Artificial Intelligence BigQuery Cloud Computing Cloud Storage Computer Programming Software Debugging Distributed Systems Identity and Access Management Python (Programming Language) Machine Learning
+21 more
NoSQL Standard Sql Search Technologies Google Cloud Enterprise Software Applications ReactJS Flask (Web Framework) Large Language Models Multi-Agent Systems Generative AI Fastapi Build Management AI Platforms HuggingFace Google Cloud Functions Machine Learning Operations Front End Software Development Virtual Agents Restful APIs Data Pipelines Automation Anywhere

Job description

We are looking for an experienced GenAI Engineer to design and build next-generation AI applications using Google Gemini, Vertex AI, and the Google Cloud Platform (Google Cloud Platform) ecosystem. The ideal candidate will have strong expertise in LangChain, LangGraph, Retrieval-Augmented Generation (RAG), agentic AI workflows, and scalable cloud-native architectures. This role involves building production-grade AI solutions, integrating LLMs into enterprise applications, and developing intelligent multi-agent systems., Design, develop, and deploy Generative AI applications powered by Google Gemini (Pro, Flash, Ultra) and Vertex AI. Build advanced prompt pipelines, RAG applications, and AI workflows using LangChain. Design and implement stateful, multi-agent AI systems using LangGraph. Develop scalable AI solutions utilizing Google Cloud services including Vertex AI Search, BigQuery, Cloud Run, Cloud Storage, and IAM. Build robust data ingestion pipelines supporting multiple document formats. Implement vector search architectures using Vertex AI Vector Search or vector databases such as Chroma, Milvus, Pinecone, Weaviate, or Qdrant. Optimize LLM performance using prompt engineering, few-shot learning, and PEFT techniques. Establish evaluation metrics for LLM accuracy, latency, hallucination detection, and model performance. Implement LLMOps best practices including observability, scalability, monitoring, and security. Develop REST APIs using FastAPI or Flask to expose AI services. Collaborate with Product Managers, Data Engineers, and Front-End Developers to integrate AI capabilities into enterprise applications., Vertex AI Google Cloud Platform (Google Cloud Platform) Langchain LangGraph FastAPI Flask RAG Preferred Skills: Llama Index Hugging Face React TypeScript PEFT LLMOps Agentic AI Vertex AI Vector Search

Requirements

Strong programming experience in Python. Experience building REST APIs using FastAPI or Flask. Hands-on experience with Google Gemini APIs, Vertex AI, and other enterprise LLM platforms. Strong expertise with Langchain and LangGraph. Experience implementing RAG architecture. Strong knowledge of Google Cloud Platform (Google Cloud Platform). Experience with Vertex AI, IAM, Cloud Run, BigQuery, and Google Cloud Storage. Experience working with Vector Databases including Pinecone, Weaviate, Qdrant, Chroma, or Milvus. Strong SQL and NoSQL database experience. Experience debugging complex AI pipelines and distributed applications. Strong problem-solving and communication skills. Preferred Qualifications: Google Cloud Professional Machine Learning Engineer Certification. Google Cloud Professional Cloud Architect Certification. Experience with Llama Index. Required Skills: Python

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.dice.com

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

4:04 min

Building practical AI agents using Google Gemini

Philipp Schmid Philipp Schmid · WWC 2025

2:37 min

Comparing traditional SQL tables versus NoSQL non-tabular databases

Stanimira Vlaeva · JS Congress

1:21 min

Exploring the target application for front end tests

Anna Mcdougall · JS Congress

3:33 min

Connecting frontends via a FastAPI proxy backend layer

Saoussen Chaabnia Saoussen Chaabnia · Europe 2026 Virtual

3:58 min

Launching a ChatGPT driver's license for HR professionals

Rudi Bauer Rudi Bauer +1 · Cappuccino with HR

3:16 min

Terminology differences between relational and NoSQL databases

Tim Faulkes · LIVE

Videos

See all

Related articles

See all