Domain 3 — Product and Model Selection
CCA Associate Foundations course · Page 4 of 10 · ← Back to all courses · Weight: 12%. Mark this page complete at the bottom to advance your course progress.
3.1 — Match the Entry Point to the User, Not the Capability · Core
| Entry point | Who it's for |
|---|---|
| Claude.ai | Individuals, ad-hoc drafting/research/analysis, no integration needed |
| Claude.ai Projects | Non-technical teams with a repeated workflow needing shared, persistent instructions |
| Claude API | Engineering teams building an integrated product or automation |
| Claude Code | Developers doing agentic coding and codebase work |
| Claude in Chrome / for Excel | Users who need Claude in-context inside a browser or spreadsheet, without switching tabs |
The decision rule: name the user before naming the entry point. "Same model underneath" is not a valid reason to recommend the API to a non-technical team — Claude.ai Projects is simpler, cheaper, and immediately usable for a team with no engineering resources. Recommending the API "for everything because it's the most powerful and flexible" is a named exam anti-pattern: it optimizes for raw capability instead of matching the actual user and job.
3.2 — The Three Model Tiers and Their Trade-Off · Core
| Tier | Optimizes for | Pick it when |
|---|---|---|
| Top (Opus class) | Hardest reasoning, long-horizon work | Complex, low-volume, failure is expensive |
| Balanced (Sonnet class) | Capability + cost + latency together | The default for most production workloads |
| Fast (Haiku class) | Lowest cost/latency | High-volume, well-bounded tasks (classification, routing, extraction) |
Default to Sonnet. Move to Haiku once measurement shows the cheaper tier holds accuracy; move to Opus only once evals show Sonnet actually falls short. "Best available model everywhere" is not a strategy — a team that ran Opus on every step of a 5-step pipeline (including simple intent classification) saw 7× cost and 2.3s vs. 800ms target latency, with no change in customer satisfaction.
Routing beats a single model at scale: a Haiku classifier triages requests and escalates only the hard slice to a capable tier — cutting cost with no quality cliff on the easy majority.
⚠ Often-missed — Prompts Don't Transfer 1-to-1 Across Tiers · Gap
A prompt carefully tuned for Opus (heavy scaffolding, many few-shot examples, explicit CoT) is not a finished artifact for Sonnet — a more capable model needs less scaffolding, a less capable one needs more. Every model swap is a release: build a test set with known-good outputs, define a grading function, set the acceptance threshold before running evals, and pin the exact model version in config — never point production at a rolling "latest" alias.
3.3 — Read the Business Constraint First · Core
| Constraint named in the scenario | It implies |
|---|---|
| Cost | Prefer Haiku; use the Batches API; avoid Opus at high volume |
| Latency / SLA | Prefer Haiku + streaming |
| Quality / accuracy floor | Prefer Sonnet; escalate to Opus only if evals show a gap |
When two architectures both pass the quality bar and one is cheaper, the cost-constrained scenario always favors the cheaper one — read the named constraint before comparing technical merits.
Extended thinking: available on current Sonnet/Opus tiers via the effort parameter; billed as output tokens and adds latency. Run evals without it first — enable only if accuracy still falls short after prompt improvements. Enabling it "just in case" on every call (including simple routing steps) is a silent cost/latency tax with no benefit.
Batches API: the correct pattern whenever there's no per-request latency requirement — overnight or bulk jobs where cost is the primary concern, regardless of which tier is chosen.
3.4 — Multimodal: What Claude Can and Cannot Do with Images · Core
Claude can read photographs, handwritten forms, charts, screenshots, and scanned documents directly — no separate OCR step needed. It cannot generate images, cannot fetch an image from a URL without a retrieval tool configured, does not provide engineering-grade color measurement (no Pantone-from-photo), and is not a forensic authentication tool. For consequential visual tasks (medical imaging, content moderation), Claude is a first-pass assistant — human review stays in the loop.
3.5 — The Three Parameters the Exam Tests · Core
| Symptom | Parameter to adjust |
|---|---|
| Output cut off mid-sentence | max_tokens too low — raise it |
| Same input, inconsistent results needed | Lower temperature (toward 0) for repeatable/classification tasks |
| Outputs too similar, need variety | Raise temperature for brainstorming/creative tasks |
The context window is one shared budget for input + output tokens together — a 200K-token input leaves zero room for output in a 200K window, regardless of max_tokens. And temperature zero minimizes but does not guarantee identical outputs run to run; design for validated quality, not bit-for-bit reproducibility.
Exam reflexes for Domain 3
- "Non-technical team, repeated weekly workflow" → Claude.ai Projects, not the API.
- "Use the API because it's more powerful/flexible" → wrong; match the user, not raw capability.
- "Cost-constrained scenario, two options both pass quality bar" → pick the cheaper one, always.
- "Opus everywhere, cost/latency exploded, CSAT unchanged" → route by task, don't default to the top tier.
- "Switching model tiers, reusing the old prompt unchanged" → wrong; re-evaluate, prompts don't transfer 1:1.
- "Overnight bulk job, no latency requirement" → Batches API.
- "Output cut off mid-sentence" → raise
max_tokens. - "Same input needs the same output every time" → lower
temperature.
Test yourself on this domain. Take the Domain 3 practice quiz — 38 questions, instant scoring, an explanation for every answer.