Evaluate the Exact Serving Build
Gate a self-hosted endpoint on representative quality, safety, security, performance, and recovery evidence from the exact deployable build.
By the end
You will be able to
- Turn intended behavior, prohibited behavior, service objectives, and recovery requirements into versioned testable claims.
- Build representative, boundary, adversarial, failure, incident, accessibility, and holdout cases with approved data.
- Combine deterministic checks, calibrated human judgment, and bounded model-assisted grading without treating one score as truth.
- Compare baseline and candidate under identical conditions and issue an evidence-linked release disposition.
Write claims and gates before running the candidate
List supported tasks, users, languages, modalities, contexts, tools, safety boundaries, latency and throughput objectives, cost envelope, privacy rules, and recovery behavior. Convert each into observable cases, scoring rules, critical slices, must-not-regress conditions, and a release threshold.
Separate model behavior from application, retrieval, template, tokenizer, adapter, runtime, quantization, hardware, and policy behavior. The evaluation target is the exact immutable serving build, not a family benchmark or unpinned model card result.
Build versioned and governed cases
Cover common tasks, important user and data slices, long and short contexts, multilingual and accessibility needs, edge cases, known failures, adversarial inputs, unsafe requests, prompt injection, malformed protocol input, timeouts, overload, dependency failure, restart, rollback, and incident reproductions. Keep holdout cases separate from prompt and configuration tuning.
Record case owner, purpose, source, permission, sensitivity, expected properties, prohibited effects, allowed tool trajectory, rubric, and expiry. Minimize personal or restricted data and use synthetic or approved fixtures when possible. Review generated cases for realism and hidden leakage.
Combine deterministic checks and calibrated judgment
Use exact checks for schemas, citations, policy flags, tool authorization, side effects, timeouts, resource limits, and recovery. Use blinded human review for usefulness, nuance, accessibility, and high-consequence quality. A model grader can add a versioned signal after calibration against labeled examples and disagreement analysis.
Do not let the candidate grade itself as the sole judge, treat an uncalibrated grader as ground truth, or average away a critical failure. Preserve raw outputs, traces, grader versions, rationales where appropriate, human overrides, and adjudication.
Compare baseline and candidate by slice and distribution
Run the same cases with the same seeds or sampling policy where meaningful, request limits, runtime configuration, hardware allocation, and load conditions. Report pass rates, severity, uncertainty, disagreement, latency distributions, resource use, and failures by important slice rather than one aggregate score.
Investigate every critical regression and unexpected improvement. A faster quantized build may change quality; a safer refusal policy may create unacceptable false blocks; a higher average score may hide harm to a smaller group or rare critical workflow.
Issue a release disposition and preserve regression evidence
Choose pass, bounded pilot, hold for evidence, reject, or roll back. Link every claim to cases and results, document exceptions and compensating controls, name reviewers and authority, record expiry and monitoring, and retain the baseline and candidate identities.
Convert privacy-reviewed production failures and incidents into reproducible regression cases. Re-run gates after artifact, template, tokenizer, runtime, quantization, hardware, policy, data, adapter, traffic, or grader changes. Publication and promotion remain human decisions.
Practice activity
Create an exact-build release gate
- Define one exact synthetic serving build and write five supported claims, three prohibited effects, two service objectives, and one recovery claim.
- Create twelve cases spanning representative, slice, boundary, adversarial, accessibility, failure, incident, and holdout coverage.
- Assign deterministic, human, or calibrated model-assisted graders and define disagreement and critical-failure handling.
- Compare synthetic baseline and candidate results under identical conditions and report by slice and distribution.
- Issue a pass, pilot, hold, reject, or rollback disposition with evidence links, reviewers, monitoring, expiry, and regression additions.
What to produce
- A versioned claim, case, rubric, threshold, and data-governance package.
- A grader calibration and disagreement record plus baseline-candidate slice results.
- A signed-paper release disposition with exact build identities, critical findings, monitoring, expiry, and rollback criteria.
Reflect before continuing
Which aggregate result could conceal a critical slice failure, and how does your gate prevent that?
Evidence
Sources and verification
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence ProfileNIST · verified 2026-07-27
- MLPerf Inference: DatacenterMLCommons · verified 2026-07-27
- Define success criteria and build evaluationsAnthropic · verified 2026-07-27
- Evaluation best practicesOpenAI · verified 2026-07-27
Knowledge check
Make it stick.
Choose the strongest answer for each question. Your attempts become part of your device-local transcript.