Large language models can answer questions, summarise documents, write code, and interact with external systems. But building a reliable AI application requires more than sending a prompt and displaying the response.
A production-ready application must manage conversation history, provide relevant context, use tools safely, handle different response types, and evaluate whether the generated output is useful.
In this tutorial, we’ll build ShopHelper, a customer-support assistant for an imaginary online shop. By the end, ShopHelper will be able to:
- Answer general questions in a consistent tone
- Remember what a customer said earlier
- Look up order statuses by calling a function in your code
- Handle Claude’s multi-block responses safely
- Process support tickets using workflows
- Evaluate whether prompt changes improve results
Each section adds one piece, so you can follow along in your own editor.
Prerequisites
You should have:
- Basic Python knowledge
- Python 3.9 or later
- An Anthropic API key
- Familiarity with functions and JSON
How to Set Up the Project and Keep Your API Key Secure
Create a virtual environment and install the Anthropic Python SDK:
python -m venv .venv
source .venv/bin/activate
pip install anthropic python-dotenv
On Windows:
.venv\Scripts\activate
Create a .env file:
ANTHROPIC_API_KEY=your_api_key_here
An API key is a secret credential. Never place it in browser JavaScript, mobile-app code, or client-side configuration. Never commit it to a repository:
echo ".env" >> .gitignore
If you add a web interface later, keep the key on your backend:
Browser → Your backend → Claude API
Create app.py:
import os
from anthropic import Anthropic
from dotenv import load_dotenv
load_dotenv()
MODEL = "claude-sonnet-5"
client = Anthropic(
api_key=os.environ["ANTHROPIC_API_KEY"]
)
load_dotenv() loads the value from .env. The MODEL constant means you only need to change the model name in one place. Confirm that the model identifier is available to your account before running the example.
How to Make Your First Request
response = client.messages.create(
model=MODEL,
max_tokens=500,
messages=[
{
"role": "user",
"content": "Explain what an API is in simple terms."
}
],
)
answer = "".join(
block.text
for block in response.content
if block.type == "text"
)
print(answer)
A request contains three important parts:
modelselects the Claude model that handles the request. Models can differ in capability, speed, and cost.max_tokenslimits the maximum amount of text Claude can generate. A smaller value can reduce latency, but Claude may stop before completing its answer.messagescontains the conversation. Each message has aroleandcontent. The role is usuallyuserorassistant.
For example, a one-off request contains one user message. A multi-turn conversation contains earlier user and assistant messages.
Claude returns response.content, which is a list of typed content blocks. Common blocks include:
| Block type | Meaning |
|---|---|
text | Generated text |
tool_use | A request for your application to call a tool |
thinking | Reasoning content when enabled |
The example collects text blocks instead of assuming response.content[0] is always text.
You can inspect usage information for monitoring:
print(response.usage.input_tokens)
print(response.usage.output_tokens)
How to Manage Conversation History
Claude doesn’t automatically remember separate API requests. Send relevant history with every request:
messages = [
{
"role": "user",
"content": "What is your returns policy?"
},
{
"role": "assistant",
"content": "Items can be returned within 30 days."
},
{
"role": "user",
"content": "How long do I have?"
},
]
response = client.messages.create(
model=MODEL,
max_tokens=300,
messages=messages,
)
The assistant message records Claude’s earlier answer, allowing the final question to be interpreted in context.
A simple chat function can maintain the history:
def chat(history, user_text):
history.append({
"role": "user",
"content": user_text,
})
response = client.messages.create(
model=MODEL,
max_tokens=500,
messages=history,
)
reply = "".join(
block.text
for block in response.content
if block.type == "text"
)
history.append({
"role": "assistant",
"content": reply,
})
return reply
history = []
print(chat(history, "What is your returns policy?"))
print(chat(history, "How long do I have?"))
Each call adds the new user message, sends the complete history, and stores Claude’s response for the next turn. In production, store histories by customer or session ID.
How to Manage History as it Grows
Unlimited history increases input size and may make it harder for Claude to focus. One option is to retain only recent messages:
def trim_history(history, max_messages=10):
trimmed = history[-max_messages:]
while trimmed and trimmed[0]["role"] != "user":
trimmed.pop(0)
return trimmed
Another option is to summarise older turns while keeping recent messages:
def summarise_history(history, keep_last=6):
old = history[:-keep_last]
recent = history[-keep_last:]
transcript = "\n".join(
f"{message['role']}: {message['content']}"
for message in old
)
response = client.messages.create(
model=MODEL,
max_tokens=250,
messages=[{
"role": "user",
"content": (
"Summarise this conversation in under 100 words. "
"Keep order numbers and unresolved issues.\n\n"
f"<conversation>{transcript}</conversation>"
),
}],
)
summary = "".join(
block.text
for block in response.content
if block.type == "text"
)
return summary, recent
Keep the summary as separate application state and include it as context in the next request. Don’t insert it as an additional user message before recent, because that can create invalid consecutive user messages.
Sensitive information should also be redacted before storage or transmission:
import re
def redact(text):
return re.sub(
r"\b(?:\d[ -]?){13,16}\b",
"[REDACTED CARD]",
text,
)
How to Structure Prompts with Clear Boundaries
XML-style tags are ordinary text, not special API commands. They make each part of a prompt explicit:
<customer_reviews>
The product is comfortable, but the available colours are limited.
Customers also describe it as durable.
</customer_reviews>
<sales_data>
January: 120 units
February: 150 units
March: 98 units
</sales_data>
<task>
Compare the reviews with the sales data.
Identify possible relationships and state uncertainty.
</task>
Here, <customer_reviews> identifies reference material, <sales_data> identifies the data, and <task> identifies the instruction. Use similar boundaries for policies, user-generated content, examples, and output requirements.
How to Use a System Prompt
A system prompt defines ShopHelper’s general behaviour:
system_prompt = """
You are ShopHelper, a friendly customer-support assistant.
Keep answers concise and clear.
Do not invent prices, policies, or order details.
If information is missing, ask for it.
"""
Pass it separately from the conversation:
response = client.messages.create(
model=MODEL,
max_tokens=500,
system=system_prompt,
messages=[
{"role": "user", "content": "Where is my order?"}
],
)
Because the customer didn’t provide an order number, ShopHelper should ask for one instead of guessing.
How to Add Tools
Tools let Claude request data from your application at runtime. Define each tool as a JSON schema:
tools = [
{
"name": "get_order_status",
"description": (
"Returns the current status of an order. "
"Call this whenever the customer asks about an order."
),
"input_schema": {
"type": "object",
"properties": {
"order_id": {
"type": "string",
"description": "The order ID, for example ORD-12345.",
},
},
"required": ["order_id"],
},
}
]
Pass the tools list to each request:
response = client.messages.create(
model=MODEL,
max_tokens=500,
system=system_prompt,
tools=tools,
messages=[
{"role": "user", "content": "Where is order ORD-12345?"}
],
)
Claude reads the tool descriptions and decides whether to call a tool or reply directly.
How to Handle a Tool-Use Response
When Claude wants to call a tool, response.stop_reason is "tool_use" and the content includes a tool_use block. Your application must call the function, then send the result back to Claude:
import json
# Simulated order database
orders = {
"ORD-12345": {"status": "Shipped", "eta": "2 days"},
"ORD-67890": {"status": "Processing", "eta": "5 days"},
}
def get_order_status(order_id):
return orders.get(order_id, {"error": "Order not found"})
def run_tool(tool_name, tool_input):
if tool_name == "get_order_status":
return get_order_status(tool_input["order_id"])
return {"error": f"Unknown tool: {tool_name}"}
def process_response(response, messages):
while response.stop_reason == "tool_use":
tool_results = []
for block in response.content:
if block.type == "tool_use":
result = run_tool(block.name, block.input)
tool_results.append({
"type": "tool_result",
"tool_use_id": block.id,
"content": json.dumps(result),
})
messages.append({
"role": "assistant",
"content": response.content,
})
messages.append({
"role": "user",
"content": tool_results,
})
response = client.messages.create(
model=MODEL,
max_tokens=500,
system=system_prompt,
tools=tools,
messages=messages,
)
return "".join(
block.text
for block in response.content
if block.type == "text"
)
The loop continues until Claude stops requesting tools and returns a final text response.
Claude Responses Can Contain Multiple Blocks
A single response may contain a thinking block, a text block, and one or more tool_use blocks. Always iterate over response.content and check block.type rather than accessing a fixed index:
for block in response.content:
if block.type == "text":
print("Text:", block.text)
elif block.type == "tool_use":
print("Tool:", block.name, block.input)
elif block.type == "thinking":
print("Thinking:", block.thinking)
When you add an assistant turn back to the message history, pass the full response.content list, not just the extracted text. This preserves all blocks, including tool_use blocks that the API requires when a subsequent tool_result references them.
Workflows vs Agents
An agent gives Claude tools and lets it decide what to do next. This is flexible but less predictable.
A workflow is a sequence of steps you control. Claude performs each step, but your code decides the order. Workflows are easier to test, debug, and monitor in production.
For ShopHelper, use a workflow to process a support ticket:
- Classify the ticket (refund, shipping, technical, other)
- Extract the order ID if present
- Look up the order status if an ID was found
- Draft a reply
def process_ticket(ticket_text):
# Step 1: classify
classification_response = client.messages.create(
model=MODEL,
max_tokens=50,
messages=[{
"role": "user",
"content": (
"Classify this support ticket into one category: "
"refund, shipping, technical, or other. "
"Reply with the category name only.\n\n"
f"<ticket>{ticket_text}</ticket>"
),
}],
)
category = "".join(
block.text
for block in classification_response.content
if block.type == "text"
).strip().lower()
# Step 2: extract order ID
extraction_response = client.messages.create(
model=MODEL,
max_tokens=50,
messages=[{
"role": "user",
"content": (
"Extract the order ID from this ticket. "
"Reply with the ID only, or 'none' if absent.\n\n"
f"<ticket>{ticket_text}</ticket>"
),
}],
)
order_id = "".join(
block.text
for block in extraction_response.content
if block.type == "text"
).strip()
# Step 3: look up order status
order_info = ""
if order_id.lower() != "none":
status = get_order_status(order_id)
order_info = f"Order status: {json.dumps(status)}"
# Step 4: draft reply
context = f"Category: {category}\n{order_info}"
reply_response = client.messages.create(
model=MODEL,
max_tokens=300,
system=system_prompt,
messages=[{
"role": "user",
"content": (
f"<context>{context}</context>\n\n"
f"<ticket>{ticket_text}</ticket>\n\n"
"Write a helpful reply to this support ticket."
),
}],
)
reply = "".join(
block.text
for block in reply_response.content
if block.type == "text"
)
return {
"category": category,
"order_id": order_id,
"order_info": order_info,
"reply": reply,
}
Each step is a separate API call with a focused prompt. This makes it straightforward to log, test, or replace individual steps without changing the rest of the pipeline.
Chaining, Parallelisation, Routing, and Evaluator-Optimizer
Chaining
Chaining passes the output of one step as the input to the next. The ticket workflow above is an example. Use chaining when later steps depend on earlier results.
Parallelisation
When steps are independent, run them at the same time:
import concurrent.futures
def analyse_review(review):
sentiment_response = client.messages.create(
model=MODEL,
max_tokens=10,
messages=[{
"role": "user",
"content": f"Sentiment of this review (positive/negative/neutral): {review}",
}],
)
sentiment = "".join(
block.text for block in sentiment_response.content
if block.type == "text"
).strip()
topic_response = client.messages.create(
model=MODEL,
max_tokens=20,
messages=[{
"role": "user",
"content": f"Main topic of this review in 3 words: {review}",
}],
)
topic = "".join(
block.text for block in topic_response.content
if block.type == "text"
).strip()
return {"sentiment": sentiment, "topic": topic}
reviews = [
"Great product, fast shipping!",
"The colour faded after one wash.",
"Good value for the price.",
]
with concurrent.futures.ThreadPoolExecutor() as executor:
results = list(executor.map(analyse_review, reviews))
for review, result in zip(reviews, results):
print(f"{review[:40]!r}: {result}")
Routing
Route each request to a specialised handler based on its category:
def handle_refund(ticket):
response = client.messages.create(
model=MODEL,
max_tokens=200,
system="You are a refund specialist. Be empathetic and clear about the refund process.",
messages=[{"role": "user", "content": ticket}],
)
return "".join(
block.text for block in response.content
if block.type == "text"
)
def handle_shipping(ticket):
response = client.messages.create(
model=MODEL,
max_tokens=200,
system="You are a shipping specialist. Provide tracking information and delivery estimates.",
messages=[{"role": "user", "content": ticket}],
)
return "".join(
block.text for block in response.content
if block.type == "text"
)
def handle_general(ticket):
response = client.messages.create(
model=MODEL,
max_tokens=200,
system=system_prompt,
messages=[{"role": "user", "content": ticket}],
)
return "".join(
block.text for block in response.content
if block.type == "text"
)
handlers = {
"refund": handle_refund,
"shipping": handle_shipping,
}
def route_ticket(ticket_text):
result = process_ticket(ticket_text)
category = result["category"]
handler = handlers.get(category, handle_general)
return handler(ticket_text)
Evaluator-Optimizer
An evaluator-optimizer loop generates a response, scores it, and regenerates if the score is too low:
def evaluate_response(ticket, response_text):
eval_response = client.messages.create(
model=MODEL,
max_tokens=100,
messages=[{
"role": "user",
"content": (
"Score this support response from 1–10. "
"Consider accuracy, tone, and completeness. "
"Reply with a number only.\n\n"
f"<ticket>{ticket}</ticket>\n"
f"<response>{response_text}</response>"
),
}],
)
score_text = "".join(
block.text for block in eval_response.content
if block.type == "text"
).strip()
try:
return int(score_text)
except ValueError:
return 5
def generate_with_quality_check(ticket, min_score=7, max_attempts=3):
for attempt in range(max_attempts):
response = client.messages.create(
model=MODEL,
max_tokens=300,
system=system_prompt,
messages=[{"role": "user", "content": ticket}],
)
response_text = "".join(
block.text for block in response.content
if block.type == "text"
)
score = evaluate_response(ticket, response_text)
print(f"Attempt {attempt + 1}: score {score}")
if score >= min_score:
return response_text
return response_text # return best attempt after max tries
How to Evaluate Prompt Quality
Testing prompt changes systematically prevents regressions. Define test cases with expected keywords, then compare two prompt versions:
test_cases = [
{
"input": "Where is my order ORD-12345?",
"expected_keywords": ["shipped", "2 days"],
},
{
"input": "I want a refund for my broken item",
"expected_keywords": ["return", "refund", "30 days"],
},
{
"input": "Do you ship internationally?",
"expected_keywords": ["international", "ship"],
},
]
def evaluate_prompt(system_prompt, test_cases):
scores = []
for test in test_cases:
messages = [{"role": "user", "content": test["input"]}]
if "ORD-" in test["input"]:
reply = process_response(
client.messages.create(
model=MODEL,
max_tokens=300,
system=system_prompt,
tools=tools,
messages=messages,
),
messages,
)
else:
response = client.messages.create(
model=MODEL,
max_tokens=300,
system=system_prompt,
messages=messages,
)
reply = "".join(
block.text for block in response.content
if block.type == "text"
)
keywords_found = sum(
1 for kw in test["expected_keywords"]
if kw.lower() in reply.lower()
)
score = keywords_found / len(test["expected_keywords"])
scores.append(score)
print(f"Input: {test['input'][:50]}")
print(f"Score: {score:.0%} ({keywords_found}/{len(test['expected_keywords'])} keywords)\n")
return sum(scores) / len(scores)
prompt_v1 = """
You are ShopHelper, a friendly customer-support assistant.
Keep answers concise and clear.
Do not invent prices, policies, or order details.
If information is missing, ask for it.
"""
prompt_v2 = """
You are ShopHelper, a friendly customer-support assistant for an online shop.
Guidelines:
- Always greet the customer warmly
- Be specific about timeframes (e.g., "within 30 days" not "soon")
- For order issues, always mention the order ID in your response
- End with an offer to help further
- Do not invent prices, policies, or order details
"""
print("Evaluating prompt v1:")
score_v1 = evaluate_prompt(prompt_v1, test_cases)
print(f"Overall score: {score_v1:.0%}\n")
print("Evaluating prompt v2:")
score_v2 = evaluate_prompt(prompt_v2, test_cases)
print(f"Overall score: {score_v2:.0%}\n")
print(f"Winner: {'v2' if score_v2 > score_v1 else 'v1'}")
Conclusion
You’ve now built ShopHelper from a single API call into a structured customer-support assistant. The application manages multi-turn conversation history, uses XML-style tags to keep prompts unambiguous, calls external tools and processes their results, routes tickets through specialised handlers, and measures prompt quality with automated test cases.
These patterns — history management, tool use, workflow chaining, parallelisation, routing, and evaluation — apply to any production AI application, not just customer support. Start with the simplest approach that meets your requirements, then add complexity only where it’s needed.