Artificial Intelligence Specialist

HCLTech
Greater London, UK
1 day ago
Apply on www.collegerecruiter.com
Prepare application

Role details

Contract type
Permanent contract
Employment type
Full-time (> 32 hours)
Experience level
Experienced
Experience required
2 years minimum
Working hours
Regular working hours

Tech stack

A/B Testing Application Programming Interfaces (APIs) Artificial Intelligence Amazon Web Services Audit Trail Microsoft Azure Cloud Engineering Databases Data Architecture Fault Tolerance Identity and Access Management Key Management
+16 more
Knowledge-Based Systems Metadata Performance Tuning Search Technologies Management of Software Versions Data Logging Data Ingestion Large Language Models Multi-Agent Systems Caching Rate Limiting Event Driven Architecture Machine Learning Operations Virtual Agents Software Coding Grpc

Job description

1) Define reference architectures for GenAI systems: RAG, agentic orchestration, tool/function calling, multi-step reasoning workflows, memory patterns, and context strategies.

  • Define reference architectures for GenAI systems including RAG, agentic orchestration, tool/function calling, multi-step reasoning workflows, memory patterns, and context strategies.
  • Design multi-tenant and enterprise-scale GenAI platforms with clear separation of concerns: UI, orchestration, retrieval, inference, evaluation, and observability.
  • Select model strategies: hosted LLMs, open-weight models, fine-tuning vs. prompt/RAG, latency and cost tradeoffs, and deployment patterns.

2) Agentic AI Orchestration & Tooling

  • Architect agent systems (single/multi-agent) including:
  • Tool use patterns (APIs, databases, search, workflow engines)
  • Guardrails to prevent unsafe tool actions and hallucinated commands
  • Build reliable flows for “human-in-the-loop” decision points and approvals (e.g., procurement, customer comms, incident triage).

3) Retrieval, Knowledge Systems & Data Design

  • Lead design of knowledge ingestion pipelines:
  • document parsing, chunking strategies, embeddings, metadata, lineage, freshness SLAs
  • Architect vector search and hybrid retrieval:
  • semantic + keyword, reranking, filtering, ACL-aware retrieval
  • Ensure retrieval respects access control, PII handling, data residency, and auditability.

4) Production Engineering, Reliability & Cost

  • Set non-functional requirements for GenAI workloads:
  • SLOs, latency budgets, fallback models, caching, rate limiting
  • Design cost controls: prompt/token optimization, model routing, batching, and usage governance.
  • Implement resiliency patterns: circuit breakers, retries, queue-based orchestration, idempotency.
  • Establish AI security posture:
  • Define policies and controls for:
  • sensitive data, logging, redaction, encryption, secret management, and auditing
  • Collaborate with risk/compliance to drive:
  • model governance, content safety, bias/quality monitoring, and regulatory alignment

6) Evaluation, Observability & Continuous Improvement

  • offline evals (golden sets), automated regression, and scenario-based testing
  • Instrument systems for observability:
  • traces, prompt/versioning, retrieval diagnostics, tool-call logs, and outcome metrics
  • Run A/B tests and iterate on prompts, retrieval, and agent policies based on measurable outcomes.

7) Leadership & Stakeholder Management

  • Partner with product leaders to identify high-value use cases and define roadmap.
  • Mentor engineers and data scientists on best practices for LLM apps.
  • Produce architecture artifacts: ADRs, threat models, system diagrams, runbooks.

Required Skills & Experience

Core Technical Skills (Must Have)

  • 8+ years in software/solution architecture with 2+ years delivering GenAI/LLM solutions in production (adjust as needed).
  • Strong knowledge of LLMs: prompting patterns, context windows, tool/function calling, model limitations, and safety risks.
  • orchestrators, workflows, multi-step reasoning, tool usage, HITL patterns
  • RAG expertise:
  • Cloud architecture (Azure/AWS/GCP) with production engineering rigor:
  • Solid programming skills (one or more):
  • Experience with APIs and integration patterns:
  • REST/gRPC, event-driven systems, queues, workflow engines

Security & Governance (Must Have)

  • Understanding of GenAI-specific threats:
  • Familiarity with enterprise controls:
  • IAM, key management, encryption, network isolation, audit logging
  • Responsible AI practices:
  • evaluation, content moderation, privacy, and compliance-by-design

Architecture & Systems Skills (Must Have)

  • scalability, fault tolerance, caching, performance tuning
  • Observability:
  • logging/metrics/tracing, prompt/version tracking, monitoring SLIs/SLOs
  • Cost management and performance optimization:
  • model selection/routing, token reduction, caching, batching

Preferred / Nice-to-Have Skills

  • Fine-tuning approaches:
  • LoRA/QLoRA, instruction tuning, adapters, distillation (when appropriate)
  • Experience with:
  • Advanced evaluation:
  • LLM-as-judge with safeguards, rubric scoring, adversarial testing
  • MLOps/LLMOps toolchains:
  • customer support automation, developer productivity copilots, IT ops agents, finance or healthcare compliance
  • Experience building platforms

Requirements

  • 8+ years in software/solution architecture with 2+ years delivering GenAI/LLM solutions in production (adjust as needed).
  • Strong knowledge of LLMs: prompting patterns, context windows, tool/function calling, model limitations, and safety risks.
  • orchestrators, workflows, multi-step reasoning, tool usage, HITL patterns
  • RAG expertise:
  • Cloud architecture (Azure/AWS/GCP) with production engineering rigor:
  • Solid programming skills (one or more):
  • Experience with APIs and integration patterns:
  • REST/gRPC, event-driven systems, queues, workflow engines

Security & Governance (Must Have)

  • Understanding of GenAI-specific threats:
  • Familiarity with enterprise controls:
  • IAM, key management, encryption, network isolation, audit logging
  • Responsible AI practices:
  • evaluation, content moderation, privacy, and compliance-by-design

Architecture & Systems Skills (Must Have)

  • scalability, fault tolerance, caching, performance tuning
  • Observability:
  • logging/metrics/tracing, prompt/version tracking, monitoring SLIs/SLOs
  • Cost management and performance optimization:
  • model selection/routing, token reduction, caching, batching

Apply for this position

This job is hosted externally. Click below to view the full posting and apply.

Apply on www.collegerecruiter.com
Prepare application

Good distractions

Talks and stories from around this role — technically off-topic, practically not.

5:30 min

Building components of a real-world LLM lifecycle

Maxim Salnikov Maxim Salnikov · LIVE

4:56 min

Establishing internal service communication with gRPC

Florian Bader Florian Bader · World Congress 2026 Europe

3:15 min

Reversing the caching model for artifact delivery

Thijs Feryn Thijs Feryn · World Congress 2026 Europe

1:47 min

Comparing Egeria to alternative open metadata solutions

Ferd Scheepers · World Congress 2022

2:12 min

Navigating technical clarity as a global black belt

Chris Heilmann +2 · LIVE

1:11 min

Evaluating architectural trade-offs between REST and gRPC

Sakshi Nasha Sakshi Nasha · Europe 2026 Virtual

Videos

See all

Related articles

See all