In this model, evals operate as executable specifications that encode desired behavior, edge cases, and quality standards in machine‑checkable form. Product teams design test suites that translate functional requirements into concrete inputs, expected outputs, and scoring functions, often using large language models as judges to capture nuanced acceptance criteria. Across AI‑first organizations, these suites increasingly function as evals-as-PRD, supplanting traditional requirement documents as the primary definition of product behavior.
Security and red‑teaming scenarios are folded directly into these eval pipelines, replacing separate safety sections and checklists that once lived inside long PRDs. Automated checks act as blueprints that enforce safety and core design goals. This approach aligns with the need for continuous safety evaluation, ensuring that evaluations evolve alongside the deployment of AI systems.
Expedia Group has begun encoding this philosophy into its governance processes for AI agents, establishing “agent release toll gates” linked to each system’s risk level. Low‑risk use cases face lighter evaluation gates, while higher‑risk, traveler‑facing agents must pass more stringent red‑teaming, safety, and compliance checks before broader deployment.
Risk-based agent release tollgates align evaluation rigor to traveler impact, safety obligations, and deployment scope.
These tollgates are tied to corporate AI principles that assign explicit ownership for standards, evaluations, and monitoring, ensuring that agents cannot scale without documented accountability. Release pathways require high‑impact models to clear extra, risk‑aligned safety thresholds.
Industry observers describe these evaluation suites as living artifacts derived from real user data rather than static documents frozen at launch. As prompts, inputs, and data distributions shift, evals continuously test whether systems still meet required behavior, quality, and safety thresholds.
Requirement information increasingly resides in benchmark datasets, rubrics, and judge prompts that evolve alongside models, replacing one‑time specification exercises. This shift moves evaluation from a late quality gate to a front‑loaded mechanism guiding agent design and ongoing performance.
For product managers, prompts and rubrics used in LLM‑as‑a‑judge evaluations function as a new kind of PRD, spelling out behavior expectations and pass‑fail criteria in executable form. Instead of debating prose about what agents “should” do, teams agree on measurable signals, target scores, and edge‑case tests that define success.
Expedia’s emphasis on risk‑based tollgates illustrates how these eval suites become both specification and governance instruments, tying release decisions directly to performance under stress, adversarial prompts, and safety scenarios.
Traditional documents do not disappear entirely; they increasingly serve as human‑readable alignment artifacts that capture strategy, business context, and non‑measurable intent alongside the evals. However, for AI‑first organizations such as Expedia, the decisive question of whether an agent is ready to ship is increasingly answered by numbers and thresholds inside evaluation suites rather than paragraphs in a PRD.
Living evals become the ongoing contract that defines “good,” catches regressions, and keeps AI behavior aligned with user needs closely.








