LLM & Agentic AI Application Evaluation
How do we decide whether an AI application is actually working well?
Research direction
Exploring evaluation approaches for agentic and multi-agent applications, including how evidence can be captured, compared and surfaced for technical and governance decisions.
My contribution
Conceived the problem space, researched the gap, shaped the approach and proof of concept, and drove implementation into reusable evaluation and governance capabilities.