Evaluate and Observe Gemini Workflows
Measure Gemini quality, safety, factuality, multimodal behavior, tool trajectories, latency, cost, and drift with versioned cases, calibrated graders, trace evidence, and release gates.
By the end
You will be able to
- Define multidimensional success, critical slices, and release thresholds before testing.
- Build representative, multimodal, adversarial, incident, and holdout evaluation cases.
- Combine deterministic, human, and calibrated model grading.
- Connect ADK traces and privacy-reviewed production evidence to regression and rollback.
Define the evaluation contract
Name the workflow, users, observable outcome, required evidence, forbidden behavior, quality, factuality, safety and permission criteria, latency and cost budgets, important modalities and slices, and release decision before running a candidate.
Declare must-not-regress gates. An average improvement can conceal a critical injection, permission, accessibility, language, modality, factuality, or tool-side-effect failure.
Build versioned cases from real work
Cover common tasks, important user groups, ambiguous and missing context, text and approved modality fixtures, safety blocks, prompt injection, tool failures, prior incidents, and legitimate sensitive cases. Preserve a holdout set.
Version inputs, modality provenance, reference evidence, expected output properties, allowed and forbidden trajectories, data approval, and case owner. Review synthetic cases for product realism.
Choose and calibrate graders
Use deterministic checks for schemas, calculations, required evidence, citations, permissions, postconditions, latency, and cost. Use blinded human review for consequential or subjective judgment and calibrated model graders for scalable rubric signals.
A model grader is not ground truth. Version it, test labeled positive and negative examples, measure disagreement and bias by slice and modality, and escalate uncertain or high-impact cases to people.
Evaluate output and agent trajectories
For tool and ADK workflows, evaluate tool choice, arguments, order, retries, callbacks, approvals, intermediate evidence, side effects, stopping, and final output. Trace evidence should locate failures, not grant authorization.
Compare baseline and candidate under identical cases, prompts, model configuration, tools, budgets, environment, and graders. Report per-case and per-slice results, critical failures, uncertainty, disagreement, latency, usage, and cost.
Close the production feedback loop
Observe allowlisted request and trace identifiers, application, prompt, model and tool versions, latency, usage, finish outcome, safety and validation events, tool results, user feedback, and terminal state. Apply redaction, access control, sampling, and retention.
Convert reproducible incidents and failures into regression cases with provenance. Track drift by version, slice, and modality; investigate threshold breaches and preserve a tested rollback with accountable owners.
Practice activity
Build a Gemini release evaluation
- Define at least six quality, factuality, safety, permission, trajectory, latency, and cost criteria with ship, hold, and rollback thresholds.
- Create twelve no-paid-call fixtures covering representative, multimodal, edge, adversarial, tool-failure, incident, user-slice, false-block, and holdout cases.
- Assign deterministic, blinded human, or calibrated model graders; validate model graders on labeled examples and measure disagreement.
- Compare baseline and candidate, grade outputs and traces, report slice evidence, and make a signed release decision with rollback triggers.
What to produce
- A versioned twelve-case suite, modality provenance, rubric, grader calibration, thresholds, and data-use approval.
- A reproducible comparison report with output and trajectory evidence, slice metrics, critical failures, disagreement, latency, cost, and decision.
Reflect before continuing
Which aggregate result looked acceptable until a critical modality, user slice, or tool trajectory was examined?
Evidence
Sources and verification
- Evaluate model and system for safetyGoogle · verified 2026-07-25
- ADK evaluationGoogle · verified 2026-07-25
- ADK observabilityGoogle · verified 2026-07-25
- Troubleshooting guideGoogle · verified 2026-07-25
Knowledge check
Make it stick.
Choose the strongest answer for each question. Your attempts become part of your device-local transcript.