AgentEval Studio
An evaluation and observability workbench that compares AI prompt / RAG / agent variants on quality, cost, latency, and failure modes, and recommends a release gate. Demoed as the ship/hold gate for an airline support copilot.
Five core AI PM skills: PRDs, evals, launch plans, risk and safety, and metrics. Each grounded in a real product from a company you know, and turned into a decision you actually make. Unofficial concept work, everything in mock mode.
Evaluation, RAG, multimodal, and safe agents. Each one is a working demo, then reads two ways: an engineering case study for how it's built, and a product case study for why. Every demo runs in mock mode, so it works with no API keys.
An evaluation and observability workbench that compares AI prompt / RAG / agent variants on quality, cost, latency, and failure modes, and recommends a release gate. Demoed as the ship/hold gate for an airline support copilot.
An AI product-intelligence workspace that ingests user feedback, clusters pain points, finds evidence, and generates PRDs, roadmap bets, and experiment plans, every claim cited.
A multimodal UX/product QA tool that reviews UI screenshots for accessibility, friction, copy clarity, and visual hierarchy, and returns prioritised, severity-scored recommendations. Ships with a worked review of a flawed SaaS checkout, each finding pinned to the exact spot on screen.
A safe-agent demo that turns a business goal into a proposed multi-step workflow, runs only human-approved tool calls, and records every action in a timestamped audit trail. Demoed on a live incident: status email, ticket, and CRM update, all behind approvals.