Skip to content

Session

Who Tests the AI? Building Trustworthy AI Systems at Enterprise Scale

with Him Raj Singh

About This Session

As organizations rapidly adopt Generative AI and autonomous AI agents, ensuring the reliability, safety, and trustworthiness of AI-powered systems has become a critical business and engineering challenge. Unlike traditional software, AI systems introduce unique risks such as hallucinations, bias, model drift, prompt injection attacks, unpredictable behavior, and compliance concerns that cannot be addressed through conventional testing approaches alone. This session explores the emerging discipline of AI Assurance and the evolving role of quality engineering in validating AI systems at enterprise scale. Attendees will gain insights into modern strategies for testing and evaluating AI applications, including large language models (LLMs), Retrieval-Augmented Generation (RAG) systems, and AI agents. The discussion will cover key areas such as AI evaluation frameworks, safety and security testing, continuous model validation, observability, governance, and responsible AI practices. Through real-world examples and practical lessons learned, participants will discover how leading organizations are building confidence in AI solutions while balancing innovation, regulatory requirements, and customer trust. The session will also examine how quality engineering teams are evolving from traditional test execution toward becoming stewards of AI reliability, transparency, and accountability. Whether you are a quality engineer, software developer, architect, engineering leader, or AI practitioner, this session will provide actionable insights and a practical framework for building, testing, and governing trustworthy AI systems in production environments.

Topics

  • AI Models
  • AI Standards
  • Anthropic
  • Agents
  • Agentic AI
  • Automation Testing
  • E2E Testing