> Markdown version of [/jobs/ext/2593458-gen-ai-engineer](https://www.wearedevelopers.com/jobs/ext/2593458-gen-ai-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Gen AI Engineer - **Company:** Alchemy Software Solutions LLC - **Location:** Mountain View, CA, United States - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Java (Programming Language), Test Suite, Application Programming Interfaces (APIs), Artificial Intelligence, Amazon Web Services, Automation of Tests, Continuous Integration, Python (Programming Language), Open Source Technology, Regression Testing, Statistical Process Control (SPC), Test Execution Engine, Management of Software Versions, WebSocket, Large Language Models, Caching, Generative AI, Backend, Kubernetes, Production Code, Api Management - **Published:** August 4, 2026 - **Apply:** https://www.dice.com/job-detail/d931cd8e-1802-416a-8a1d-88ddcc782b78 ## About the Role GenOS is Client''s Generative AI Operating System - the platform every GenAI experience at Client is built, deployed, and governed on, including the customer-facing Client Assist. The stack: · GenStudio - LLM sandbox and extensible model catalog; new models onboarded in days. · AI Workbench - versioned Prompt Management and the LLM Leaderboard for benchmarking. · GenRuntime - GenOrchestrator (planner, executor, memory, retrieval) plus agents and tools grounding LLMs in Client domain knowledge. · GenUX - 140+ AI UX components consumed by product teams. · GenSRF - security, risk, and fraud guardrails as a platform feature, not an afterthought. · Multi-LLM catalog: Anthropic Claude via AWS Bedrock, Gemini, Llama, and Mistral., · 7+ years backend/platform engineering (Java, Python, or Go) with deep distributed-systems fundamentals: async processing, caching, idempotency, failure handling. · 2-3+ years building LLM-powered systems in production, not prototypes: agent/orchestration frameworks (LangChain / LlamaIndex / homegrown), structured tool calling, prompt versioning and evaluation. · Structured-output engineering at production grade - generated tests are code artifacts: schema enforcement, output validation, and repair loops are the daily job. · Demonstrated QE domain depth: has built or owned test automation architecture - framework design, CI quality gates, flaky-test economics - enough to encode that judgment into agent behavior. · A platform engineer who has never owned a test suite will build QE agents that generate garbage confidently. · One major LLM provider at scale - AWS Bedrock strongly preferred - with real operational scars: rate limits, latency variance, provider failover, version drift. · AWS + Kubernetes deployment depth (services run on Client Kubernetes Service). · Fluent in AI-assisted development workflows - Copilot-class tooling is the expected daily working mode. Nice-to-Have · LLM evaluation engineering: golden datasets, LLM-as-judge calibration, mutation testing or fault injection to validate generated-test quality. · MCP tool integrations; SSE/WebSocket streaming for agentic responses. · Vector stores and retrieval in production (OpenSearch / pgvector / Pinecone or equivalent). · Multi-tenant enterprise platforms under strict security/compliance; fintech background. · Open-source / inner-source contribution record. ## Description · Build QE-generation agents on GenRuntime: agents that consume artifacts the GenOS dev flow already produces - requirements, code diffs, API specs, GenUX component usage - and emit functional, API, and regression test assets as validated, structured output. · Wire QE enablement into the paved road: an app scaffolded through GenOS gets QE agents attached by default - no opt-in ceremony, self-serve onboarding, zero hand-holding. · Build the agent toolbelt as reusable GenRuntime tools: test execution, self-healing selectors and API contracts, failure triage, defect summarization - including agent-to-agent patterns where test agents interrogate the application''s own agents. · Extend AI Workbench eval primitives (LLM Leaderboard, prompt evaluation) into release gates for GenAI applications: golden datasets, LLM-as-judge scoring, statistical quality thresholds enforced in paved-road CI/CD - not just model selection. · Embed GenSRF coverage into generated tests: safety, privacy, and moderation regressions become executable test cases, not audit findings. · Integrate with Feature Management so QE-agent rollout is flagged, measured, and adoption-tracked per product team. · Write well-tested, production-grade code; participate in reviews and design discussions - PR merge velocity and AI-assisted code in PRs are tracked org KPIs. · Participate in the production support/on-call rotation for the QE capability surface (rotation shape and compensation treatment per Open Items). · Contribute self-serve onboarding docs and inner-source repos - inner-source contribution is a tracked PDX metric. ## Related Videos - [Beyond GPT: Building Unified GenAI Platforms for the Enterprise of Tomorrow](https://www.wearedevelopers.com/videos/1525-beyond-gpt-building-unified-genai-platforms-for-the-enterprise-of-tomorrow) - [Developing the Backend with Stefan Lingler, CTO at Shpock](https://www.wearedevelopers.com/videos/100360-developing-the-backend-with-stefan-lingler-cto-at-shpock) - [HTTP headers that make your website go faster](https://www.wearedevelopers.com/videos/1676-http-headers-that-make-your-website-go-faster) - [Minimal infrastructure for Real‑Time Phone Agents: transcripts in, responses out](https://www.wearedevelopers.com/videos/1736-minimal-infrastructure-for-real-time-phone-agents-transcripts-in-responses-out) - [How to Avoid LLM Pitfalls - Mete Atamel and Guillaume Laforge](https://www.wearedevelopers.com/videos/1328-how-to-avoid-llm-pitfalls-mete-atamel-and-guillaume-laforge) - [Nest.js - TypeScript in the backend can also be clean](https://www.wearedevelopers.com/videos/1033-nest-js-typescript-in-the-backend-can-also-be-clean) ## Related Articles - [Graph and AI Trends 2026: Why Is AI Running but Not Yet Delivering?](https://www.wearedevelopers.com/magazine/680-graph-and-ai-trends-2026-why-is-ai-running-but-not-yet-delivering) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Got AI ideas but no money? Here are 10 free ways to level up your AI skills with Google Cloud](https://www.wearedevelopers.com/magazine/600-got-ai-ideas-but-no-money-here-are-10-free-ways-to-level-up-your-ai-skills-with-google-cloud) - [From Prototype to Production: Build AI Agents with This Free 4-Course Learning Path](https://www.wearedevelopers.com/magazine/655-from-prototype-to-production-build-ai-agents-with-this-free-4-course-learning-path) - [MLOps And AI Driven Development](https://www.wearedevelopers.com/magazine/82-mlops-and-ai-driven-development) - [Everything a Developer Needs to Know About MCP with Neo4j](https://www.wearedevelopers.com/magazine/604-everything-a-developer-needs-to-know-about-mcp-with-neo4j)