AI is no longer a side experiment for agencies. At Ambaum, Claude is embedded in how we run QA cycles, technical audits, code reviews, theme research, and merchant deliverables across our Shopify Plus practice. Output quality scales directly with prompt quality, which is why we treat prompt engineering as a core operating skill, not a soft skill.
This is our internal playbook for working with Claude, refined against real Shopify Plus work and informed by Anthropic's own published guidance.
How We Use Claude at Ambaum
Before the mechanics, here is where Claude actually shows up in our day-to-day:
- Code review and refactors for Liquid, JavaScript, and Hydrogen components
- Technical audits of theme performance, metafield architecture, and checkout extensibility
- QA scripting for regression testing across launches
- Research on Shopify Plus features, API changes, and merchant industry context
- Documentation generation for client handoffs and internal SOPs
Each of these has different prompt requirements. A QA script needs deterministic output. A research brief needs synthesized reasoning. The structure of your prompt is what tells Claude which mode to operate in.
The Anatomy of a Strong Claude Prompt
Strong prompts follow a consistent five-part structure.
1. Role and Task (1-2 sentences)
Establish who Claude is and what the job is. Example: "You are a senior Shopify Plus developer reviewing a Liquid template for performance issues."
2. Dynamic or Retrieved Content
Drop in the actual artifacts: the code block, the metafield schema, the GraphQL response, the client brief. This is the raw material Claude operates on.
3. Detailed Task Instructions
Spell out the work. What to look for, what to ignore, what format the output should take, what edge cases matter.
4. Examples
Show, don't just tell. (Full section below.)
5. Repeated Critical Instructions
For large or complex prompts, restate the most important constraints at the end. Claude weights recency, and a 2,000-token prompt can drift if a critical instruction is buried at the top.
This structure isn't optional polish. It's the difference between Claude producing a usable audit and producing a confident-sounding generic response.
Why Claude Loves XML Structure
Claude is trained to follow structural delimiters, and it prefers XML-style tags over Markdown headers, dashes, or plain whitespace.
Why XML wins:
- Clear boundaries: angle-bracketed tags leave no ambiguity about where content starts and ends
- Token efficiency: XML tags are short and predictable
- Nestability: You can nest examples inside instructions without losing structure
- Self-documenting: Tag names communicate intent like liquid_template, merchant_brief, and audit_criteria
The practical effect is significant: the same prompt rewritten with XML tags will outperform a flat prose version on long, multi-input tasks. Section titles and headers help humans skim; XML helps Claude parse.
The Power of Examples (Few-Shot Prompting)
If you only adopt one habit from this playbook, make it this: give Claude examples.
Examples act as concrete templates. They are especially valuable when the task requires:
- Consistent formatting (audit reports, code comments, ticket descriptions)
- Specific jargon (Shopify-native terminology, Plus-only features, Liquid filters)
- Adherence to industry standards (accessibility, performance budgets, Shopify theme conventions)
It is often faster to show Claude a sample output with brief annotations than to describe the output in prose. Pair the example with a short note on what makes it correct.
Best practices for examples:
- Relevance: examples should match the task domain
- Diversity: include edge cases, not just happy paths
- Quantity: 3 to 5 examples is the sweet spot
Below three, Claude has to generalize too aggressively. Above five, you are usually burning tokens without meaningful gain.
Controlling Hallucinations in Production Work
For agency work, hallucination control is non-negotiable. A confidently wrong answer about a Shopify Plus feature, a GraphQL field, or a checkout extension capability can cost a client real money.
Four techniques work reliably.
1. Give Claude permission to say "I don't know"
Most hallucinations come from prompts that implicitly require an answer. Explicitly instruct Claude: "If you do not have enough information to answer confidently, say so."
2. Set a confidence threshold
"Only answer if you are very confident in your response. Otherwise, flag the uncertainty and explain what additional context would help."
3. Force chain-of-thought reasoning
Ask Claude to think before answering. For technical questions, "Walk through your reasoning step by step before stating a conclusion" produces dramatically more accurate output than direct-answer prompts.
4. Anchor responses in source documents
For long documents (API references, client briefs, technical specs), instruct Claude to find relevant quotes first, then answer using those quotes. This grounds the response in verifiable source material instead of synthesized memory.
A Reusable Shopify Plus Prompt Template
Putting it all together, here is a skeleton you can adapt:
Adapt the tags to the use case. The structure stays consistent.
The Bottom Line
Prompt engineering is becoming a core agency skill, not a side hobby for engineers. The teams that systematize their Claude usage now will produce more accurate QA, faster audits, and higher-quality client deliverables than teams still treating AI as a novelty.
This is how Ambaum works with Claude. Steal it, adapt it, ship better Shopify Plus work.




