AI
    September 25, 20269 min read

    I Passed the Claude Certified Developer – Foundations Exam: Here's What Actually Helped

    A personal account of passing the Claude Certified Developer – Foundations exam on my first attempt — how I prepared, what the exam is actually like, what caught me off guard, and a full study plan for anyone else sitting it.

    Share

    I sat the Claude Certified Developer – Foundations exam on September 18, 2026, and passed on my first attempt. It's the fourth Claude certification I've completed, after the Associate – Foundations and both Architect exams. Out of all four, this is the one I'd recommend first to anyone who actually writes code against the Claude API — it's the most directly useful to your day job, and the one where the preparation itself changes how you build.

    Here's what the exam is actually like, how I prepared, and what I'd tell anyone sitting it after me.


    What the Exam Is

    The Claude Certified Developer – Foundations (CCDV-F) is Anthropic's practitioner-level certification for people who build production software with Claude. Unlike the Associate exam, which is deliberately non-technical, this one assumes you write code — it tests the API, the SDK, Claude Code, MCP, and the operational realities of running Claude in production.

    Format: 53 questions, 120 minutes, scenario-based multiple choice (some single-answer, some "select two"). Like all four Claude certifications, scoring is on a scaled 100–1000 range with a passing score of 720 — so a raw percentage isn't quite the right mental model for how close you are to passing.

    Domains covered (8 total):

    1. Agents and Workflows
    2. Applications and Integration
    3. Claude Code
    4. Eval, Testing, and Debugging
    5. Model Selection and Optimization
    6. Prompt and Context Engineering
    7. Security and Safety
    8. Tools and MCPs

    It is not a coding exam in the sense of writing and submitting code — there's no IDE. But it reads like one was designed by people who've actually shipped Claude integrations. The scenarios are specific: production incident descriptions, cost/latency tradeoffs, a snippet of behavior that needs diagnosing. Memorizing API parameter names won't get you through it; you need to have actually reasoned about why a given design choice is correct.


    My Study Approach

    I gave myself about six weeks, studying in evenings and weekends around a full-time role. Two things anchored the prep, plus a third I leaned on because I'd already built it:

    Anthropic's own documentation — read as a syllabus, not a reference. I didn't start with the API reference. I started with the conceptual docs: agent architecture patterns, prompt caching mechanics, context window management, the Claude Code permission model, MCP's tool/resource/prompt primitives. If you've been treating the docs as something to Ctrl+F when you're stuck, that's exactly the gap this exam finds.

    Hands-on work with Claude Code and the API — not just reading about the failure modes, hitting them. This mattered more than any single resource. Reading that a stop_reason of "tool_use" means "keep looping" is different from building an agent loop, watching it hang because you checked content[0].type === "text" instead, and fixing it. Same with prompt caching: reading that a dynamic prefix invalidates the cache is one thing; watching your own cache hit rate stay at zero because you put the user's question before the system prompt is another. The exam rewards the second kind of knowledge specifically — several questions are constructed exactly around "this looks like it should work, why doesn't it."

    My own site's material, because I'd built it while preparing. I ended up writing a full CCA Developer prep course (five modules plus a scenarios/drills lesson and a cheatsheet) as part of studying — turning my own notes into something structured forced me to be precise about the parts I was hand-waving. If you're not building your own version of this, working through someone else's structured material (mine or otherwise) gets you most of the same benefit.


    What the Exam Is Actually Like

    Duration: 120 minutes for 53 questions — roughly 2.25 minutes per question if you pace evenly, though scenario questions take longer to read than they do to answer once you've parsed them.

    Passing mark: 720 out of a 100–1000 scaled score. You see your result, broken down by domain, immediately after submitting.

    Difficulty: Fair, if you've actually built something. The exam isn't trying to trick you with obscure trivia — it's testing whether you understand the consequence of a design choice, not just its name. If you've only read about agent loops, tool design, and prompt caching without hitting their failure modes yourself, several questions will feel harder than they should.

    What surprised me most: how much the exam rewards precision over general competence. Two options can both look defensible; the difference is often one word. "Which approach guarantees the step runs first" and "which approach usually runs the step first" have different correct answers — a programmatic prerequisite gate guarantees it, a system-prompt instruction does not. That distinction runs through the entire Security and Safety domain, and it shows up again in Eval/Testing questions about what a retry can and can't fix.


    Things That Caught Me Off Guard

    1. How much weight production concerns carry, not just correctness.

    I expected the exam to mostly test "does this code work." It spends real weight on "does this scale, does this fail safely, and does this cost what you think it costs." Questions about prompt-caching prefix order, Batches API vs. synchronous calls for bulk workloads, and horizontal scaling under concurrent load all assume you've thought about a system beyond a single successful request.

    2. Security questions test the enforcement layer, not the intent.

    Several scenarios describe a system prompt instruction meant to stop something bad from happening — sending an email without approval, following an instruction hidden in a fetched web page — and ask what's wrong with relying on it. The correct answer is consistently structural: a PreToolUse hook, a scoped credential, a human-in-the-loop gate. An instruction is a request the model can fail to follow; a hook is enforced outside the model's judgment entirely. Once you see this pattern once, you start spotting it everywhere in the exam.

    3. The Claude Code and MCP domain is deeper than "I use it daily" prepares you for.

    Using Claude Code well and knowing its configuration hierarchy precisely are different skills. Questions on CLAUDE.md scope (project vs. user vs. directory-level, and that they accumulate rather than override each other), hook exit codes, and MCP transport choices (stdio for local processes vs. HTTP for shared remote servers) go deeper than casual daily use requires. I use Claude Code constantly and still had to go back and read the configuration docs properly.


    What I'd Do Differently

    Build the thing that breaks, don't just read that it breaks. Every concept I found "obvious" in the docs but actually struggled with on practice questions was one I hadn't personally debugged. If I were starting over, I'd front-load small, deliberately-broken exercises — an agent loop with the wrong termination check, a cache that never hits, a tool with an ambiguous description next to a similar one — before reading the fix.

    Take the domain-by-domain practice seriously, not just the full mock. A full practice exam tells you your overall readiness; it doesn't tell you which of the eight domains is actually weak, because a good score in seven domains hides a bad one in the eighth. Drilling one domain at a time is slower but far more diagnostic.

    Don't skip Security and Safety because it "isn't your job." I nearly did — it's the domain furthest from day-to-day feature work for a lot of developers. It's also where the exam's sharpest precision questions live, and the underlying discipline (never trust model-layer instructions as your only control) is the single idea that transfers most directly back into real production systems.


    Who Should Take This Exam

    The Developer – Foundations exam is right for you if:

    • You write code against the Claude API, Claude Agent SDK, or Claude Code, even occasionally
    • You want to know precisely where your production Claude usage has a gap — cost control, security boundary, eval coverage — rather than finding out from an incident
    • You're evaluating candidates or reviewing PRs that touch Claude integrations and want a shared, validated vocabulary for what "correct" looks like
    • You've already got the Associate certification and want the technical-depth counterpart to it

    If you don't write code day-to-day, the Associate – Foundations is the better starting point — this exam assumes you do.


    Why This Certification Is Actually Useful

    It's the certification most directly tied to shipping. The Associate exam validates that you understand Claude conceptually. This one validates that you'd build it correctly — caching set up right, tool design that doesn't misroute, security enforced structurally instead of hoped for. That's a different, more expensive kind of knowledge to fake, which is exactly why it's worth having verified.

    It surfaces the gap between "it works in my demo" and "it survives production." A huge share of the exam's difficulty is in the distance between a working prototype and a system that handles concurrency, cost, and adversarial input correctly. Studying for it is one of the more efficient ways to close that gap deliberately instead of learning it the hard way, in production, at 2am.

    It's a credible, specific signal. A badge issued by Anthropic — verifiable on Credly, tied to a real scored exam — says something that "I've used the Claude API" on a resume doesn't. In a market full of self-declared AI expertise, that specificity is the whole value.

    It completes the picture if you're building toward the Architect track. The Architect exams assume this level of hands-on fluency and test design judgment on top of it. Sitting this one first — or even after, as I did — makes the gaps in either direction visible rather than assumed.


    Resources

    • Anthropic's own documentation — the conceptual guides on agents, prompt caching, and MCP specifically, not just the API reference
    • Hands-on practice — build something small enough to break on purpose; the exam rewards having personally hit a failure mode, not just read about it
    • My CCA Developer learning track on this site — five modules, an exam-scenarios lesson, a cheatsheet, plus a full practice exam and domain-by-domain drilling with a second mock exam if you want a fresh run before the real thing

    The badge: Claude Certified Developer – Foundations

    If you're preparing for this exam and get stuck on something specific, I'm reachable via the contact page. Good luck — build something real while you study, and the exam takes care of itself.

    Ask about this article

    Get answers grounded in this post. AI-generated — based on this article, and may be imperfect.

    Was this helpful?
    AY
    Avaneesh Yadav

    I build enterprise AI systems — Spring AI, RAG, and agents — and write about shipping LLMs to production. I also run advisory and workshops for engineering teams.

    Scaled AI Weekly

    Enjoyed this? Get more like it every Monday.

    Real architecture decisions, LLMOps patterns that survive production, and engineering leadership advice — from 12+ years of building at enterprise scale. Free. No spam. Unsubscribe anytime.

    Join engineers building production AI systems

    Free: LLM Production Readiness Checklist (PDF)

    50 checks across observability, rate limiting, cost optimization, failure handling, and security — for teams shipping AI features to production.

    No spam. Unsubscribe any time.

    Comments