Compare Exact Product Experiences
Use dated first-party evidence to compare model endpoints, assistants, developer agents, frameworks, and managed agent services.
By the end
You will be able to
- Compare exact experiences without ranking entire providers.
- Use documented, tested, inference, and unknown evidence correctly.
- Separate model capability from harness behavior within a provider family.
- Apply observation and review dates to volatile claims.
The comparison unit is the configured experience
Compare the surface a learner can actually use: model endpoint, consumer product, developer tool, framework application, managed service, or composed system. Record the version, date, tools, permissions, state, and runtime.
The same provider can supply both a model and an agentic product. Those are not contradictory classifications because they describe different boundaries.
Keep evidence kinds separate
Documented means a current first-party source directly describes a contract or behavior. Tested means a controlled run of the exact version and configuration produced retained evidence. Inference is a cited conclusion. Unknown is unresolved.
Documentation about a feature is not proof that a particular account, mode, or run used it. A live test does not establish future behavior after versions or configuration change.
Use model-versus-harness pairs
Pairs make boundaries visible: Claude API and Claude Code; an OpenAI model and Codex; MAI model endpoints and Foundry Agent Service; Gemini API and an ADK application; Kimi or DeepSeek tool-call APIs and the client loops around them.
For xAI and consumer assistants, qualify the exact mode and enabled connectors. For open-weight deployments, separate access to model artifacts from the agent framework and executor.
Volatile claims expire
Every product claim needs an observation date, review date, source owner, affected content, and a rule for downgrade or removal. Recheck sooner after a rename, runtime release, permission change, incident, or documentation revision.
When a source becomes stale or contradictory, block publication, preserve the old evidence for audit, and mark the current conclusion unknown until reviewed.
Practice activity
Build a dated cross-provider comparison
- Select at least eight exact experiences spanning the eight provider families in the evidence matrix.
- Include at least four within-family model-versus-harness comparisons.
- Record observable behavior, classification, claim evidence kind, source, observation date, review date, boundaries, and unknowns.
- Have another learner challenge one conclusion and revise or defend it using evidence.
What to produce
- A dated eight-family comparison with no provider-wide classification.
- A retained challenge and disposition showing how contradictory or missing evidence was handled.
Reflect before continuing
Which classification changed when you narrowed the unit from provider to exact experience?
Evidence
Sources and verification
- What is Microsoft Foundry Agent Service?Microsoft · verified 2026-07-27
- xAI tools overviewxAI · verified 2026-07-27
- smolagentsHugging Face · verified 2026-07-27
Knowledge check
Make it stick.
Choose the strongest answer for each question. Your attempts become part of your device-local transcript.