Module 4 of 5 · 45 min

Compare Exact Product Experiences

Use dated first-party evidence to compare model endpoints, assistants, developer agents, frameworks, and managed agent services.

Core concept

By the end

You will be able to

  • Compare exact experiences without ranking entire providers.
  • Use documented, tested, inference, and unknown evidence correctly.
  • Separate model capability from harness behavior within a provider family.
  • Apply observation and review dates to volatile claims.
01

The comparison unit is the configured experience

Compare the surface a learner can actually use: model endpoint, consumer product, developer tool, framework application, managed service, or composed system. Record the version, date, tools, permissions, state, and runtime.

The same provider can supply both a model and an agentic product. Those are not contradictory classifications because they describe different boundaries.

02

Keep evidence kinds separate

Documented means a current first-party source directly describes a contract or behavior. Tested means a controlled run of the exact version and configuration produced retained evidence. Inference is a cited conclusion. Unknown is unresolved.

Documentation about a feature is not proof that a particular account, mode, or run used it. A live test does not establish future behavior after versions or configuration change.

03

Use model-versus-harness pairs

Pairs make boundaries visible: Claude API and Claude Code; an OpenAI model and Codex; MAI model endpoints and Foundry Agent Service; Gemini API and an ADK application; Kimi or DeepSeek tool-call APIs and the client loops around them.

For xAI and consumer assistants, qualify the exact mode and enabled connectors. For open-weight deployments, separate access to model artifacts from the agent framework and executor.

04

Volatile claims expire

Every product claim needs an observation date, review date, source owner, affected content, and a rule for downgrade or removal. Recheck sooner after a rename, runtime release, permission change, incident, or documentation revision.

When a source becomes stale or contradictory, block publication, preserve the old evidence for audit, and mark the current conclusion unknown until reviewed.

Practice activity

Build a dated cross-provider comparison

  1. Select at least eight exact experiences spanning the eight provider families in the evidence matrix.
  2. Include at least four within-family model-versus-harness comparisons.
  3. Record observable behavior, classification, claim evidence kind, source, observation date, review date, boundaries, and unknowns.
  4. Have another learner challenge one conclusion and revise or defend it using evidence.

What to produce

  • A dated eight-family comparison with no provider-wide classification.
  • A retained challenge and disposition showing how contradictory or missing evidence was handled.

Reflect before continuing

Which classification changed when you narrowed the unit from provider to exact experience?

Evidence

Sources and verification

Knowledge check

Make it stick.

Pass at 80%

Choose the strongest answer for each question. Your attempts become part of your device-local transcript.

01What is the correct unit for an agentic-AI product comparison?
02What does documented evidence prove?
03Why compare a model endpoint with an agent product from the same provider?
04What should stale evidence do?
05How should an unresolved product mode be classified?