Claude Best Practices: A Shopify Plus Agency Guide

How Ambaum uses Claude for Shopify Plus QA, audits, and code reviews. Prompt structure, XML tags, examples, and hallucination control from real agency work.

AI is no longer a side experiment for agencies. At Ambaum, Claude is embedded in how we run QA cycles, technical audits, code reviews, theme research, and merchant deliverables across our Shopify Plus practice. Output quality scales directly with prompt quality, which is why we treat prompt engineering as a core operating skill, not a soft skill.
‍

This is our internal playbook for working with Claude, refined against real Shopify Plus work and informed by Anthropic's own published guidance.
‍

How We Use Claude at Ambaum
‍

Before the mechanics, here is where Claude actually shows up in our day-to-day:

  • Code review and refactors for Liquid, JavaScript, and Hydrogen components
  • Technical audits of theme performance, metafield architecture, and checkout extensibility
  • QA scripting for regression testing across launches
  • Research on Shopify Plus features, API changes, and merchant industry context
  • Documentation generation for client handoffs and internal SOPs
    ‍

Each of these has different prompt requirements. A QA script needs deterministic output. A research brief needs synthesized reasoning. The structure of your prompt is what tells Claude which mode to operate in.
‍

The Anatomy of a Strong Claude Prompt
‍

Strong prompts follow a consistent five-part structure.
‍

1. Role and Task (1-2 sentences)

Establish who Claude is and what the job is. Example: "You are a senior Shopify Plus developer reviewing a Liquid template for performance issues."
‍

2. Dynamic or Retrieved Content

Drop in the actual artifacts: the code block, the metafield schema, the GraphQL response, the client brief. This is the raw material Claude operates on.
‍

3. Detailed Task Instructions

Spell out the work. What to look for, what to ignore, what format the output should take, what edge cases matter.
‍

4. Examples

Show, don't just tell. (Full section below.)
‍

5. Repeated Critical Instructions

For large or complex prompts, restate the most important constraints at the end. Claude weights recency, and a 2,000-token prompt can drift if a critical instruction is buried at the top.
‍

This structure isn't optional polish. It's the difference between Claude producing a usable audit and producing a confident-sounding generic response.
‍

Why Claude Loves XML Structure
‍

Claude is trained to follow structural delimiters, and it prefers XML-style tags over Markdown headers, dashes, or plain whitespace.
‍

Why XML wins:

  • Clear boundaries: angle-bracketed tags leave no ambiguity about where content starts and ends
  • Token efficiency: XML tags are short and predictable
  • Nestability: You can nest examples inside instructions without losing structure
  • Self-documenting: Tag names communicate intent like liquid_template, merchant_brief, and audit_criteria
    ‍

The practical effect is significant: the same prompt rewritten with XML tags will outperform a flat prose version on long, multi-input tasks. Section titles and headers help humans skim; XML helps Claude parse.
‍

The Power of Examples (Few-Shot Prompting)
‍

If you only adopt one habit from this playbook, make it this: give Claude examples.
‍

Examples act as concrete templates. They are especially valuable when the task requires:
‍

  • Consistent formatting (audit reports, code comments, ticket descriptions)
  • Specific jargon (Shopify-native terminology, Plus-only features, Liquid filters)
  • Adherence to industry standards (accessibility, performance budgets, Shopify theme conventions)
    ‍

It is often faster to show Claude a sample output with brief annotations than to describe the output in prose. Pair the example with a short note on what makes it correct.
‍

Best practices for examples:
‍

  • Relevance: examples should match the task domain
  • Diversity: include edge cases, not just happy paths
  • Quantity: 3 to 5 examples is the sweet spot
    ‍

Below three, Claude has to generalize too aggressively. Above five, you are usually burning tokens without meaningful gain.
‍

Controlling Hallucinations in Production Work
‍

For agency work, hallucination control is non-negotiable. A confidently wrong answer about a Shopify Plus feature, a GraphQL field, or a checkout extension capability can cost a client real money.
‍

Four techniques work reliably.
‍

1. Give Claude permission to say "I don't know"

Most hallucinations come from prompts that implicitly require an answer. Explicitly instruct Claude: "If you do not have enough information to answer confidently, say so."
‍

2. Set a confidence threshold

"Only answer if you are very confident in your response. Otherwise, flag the uncertainty and explain what additional context would help."
‍

3. Force chain-of-thought reasoning

Ask Claude to think before answering. For technical questions, "Walk through your reasoning step by step before stating a conclusion" produces dramatically more accurate output than direct-answer prompts.
‍

4. Anchor responses in source documents

For long documents (API references, client briefs, technical specs), instruct Claude to find relevant quotes first, then answer using those quotes. This grounds the response in verifiable source material instead of synthesized memory.
‍

A Reusable Shopify Plus Prompt Template
‍

Putting it all together, here is a skeleton you can adapt:
‍


You are a senior Shopify Plus developer reviewing [artifact type] for [merchant context].



[Merchant brief, technical environment, constraints]



[The actual code, schema, or content under review]



1. [Specific task]
2. [Output format]
3. [Edge cases or exclusions]



  [Sample correct output 1]
  [Sample correct output 2]
  [Sample correct output 3]



- If you are not confident, say so
- Anchor recommendations in the artifact above
- Use the format shown in examples

Adapt the tags to the use case. The structure stays consistent.
‍

The Bottom Line
‍

Prompt engineering is becoming a core agency skill, not a side hobby for engineers. The teams that systematize their Claude usage now will produce more accurate QA, faster audits, and higher-quality client deliverables than teams still treating AI as a novelty.
‍

This is how Ambaum works with Claude. Steal it, adapt it, ship better Shopify Plus work.
‍

Frequently Asked Questions
What is Claude best used for in a Shopify Plus agency?
Claude excels at code review, theme audits, QA scripting, technical research, and documentation across Shopify Plus work. At Ambaum we use it for Liquid refactors, metafield architecture reviews, GraphQL query generation, and merchant industry research. Its strength is structured technical reasoning, especially when prompts include code artifacts and explicit instructions.
How should you structure a prompt for Claude?
A strong Claude prompt follows five parts: a one or two sentence role and task description, any dynamic or retrieved content, detailed task instructions, examples of desired output, and a repeat of critical instructions at the end. This structure consistently outperforms loose prose prompts for technical work.
Why does Claude prefer XML tags over Markdown?
Claude is trained to recognize XML-style delimiters because their boundaries are unambiguous and token-efficient. Tags like and separate content cleanly and nest without ambiguity, which helps Claude parse complex prompts. Markdown headers work for humans, but XML produces more reliable output on long, multi-input agency prompts.
How many examples should you give Claude in a prompt?
Three to five examples is optimal for most tasks. Fewer than three forces Claude to generalize too aggressively, while more than five usually wastes tokens without improving output. Prioritize relevance to your task, diversity across edge cases, and inclusion of both standard outputs and tricky scenarios you want handled correctly.
How do you stop Claude from hallucinating?
Four techniques work reliably: tell Claude it can say "I don't know," require high confidence before answering, force step-by-step reasoning before conclusions, and for long documents, instruct Claude to pull relevant quotes first and answer using those quotes. Combined, these dramatically reduce confident-but-wrong responses in production work.
No pitch. No pressure. Just perspective.

Get the strategic input you’ve been missing