If you've been searching for a working, no-fluff guide on how to build AI agents, you're in the right place. Most tutorials show you a toy example and stop before anything gets interesting. This one builds a real autonomous agent using the Claude API and Python — one that can call tools, reason through multi-step problems, and keep looping until it has an actual answer.
What You'll Build
By the end of this tutorial you'll have a working AI agent that uses Claude's tool-calling feature to answer questions it couldn't answer on its own. The agent can check the current date and time, do math calculations, and look up weather data — all triggered automatically based on what you ask it. You'll run it from the command line in under ten minutes.
Prerequisites
- Python 3.9 or higher installed on your machine
- An Anthropic API key (get one at console.anthropic.com)
- Basic familiarity with Python — you don't need to be an expert
pipavailable to install the Anthropic SDK- A terminal or command prompt you're comfortable using
All the code in this tutorial fits together into one complete, runnable project. I'll build it up section by section below so every line makes sense by the time you run it. Copy each snippet in order, or assemble the final file from the steps — either way works.
Step 1: Set Up Your Claude API Environment
First, install the Anthropic SDK. One command and you're done.
terminalpip install anthropic
Next, set your API key as an environment variable. Never hardcode it in your source file — that's how keys get leaked.
terminal (Mac/Linux)export ANTHROPIC_API_KEY="sk-ant-your-key-here"
$env:ANTHROPIC_API_KEY="sk-ant-your-key-here"
Now create a new file called agent.py. Start with a quick sanity check to make sure your key is loading correctly before writing any agent logic.
import os
import anthropic
# Quick connection test before building anything else
client = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY"))
response = client.messages.create(
model="claude-sonnet-4-6",
max_tokens=64,
messages=[{"role": "user", "content": "Say: API connection successful"}]
)
print(response.content[0].text)Run it. If you see "API connection successful" printed back, you're good to go. If you get an authentication error, double-check that the environment variable is actually set in the same terminal session you're running Python from.
Step 2: Define Your Agent's Tools and Functions
Tools are what turn Claude from a chatbot into an agent. You define them as JSON schemas, and Claude decides when and how to call them based on the user's question. Think of it like giving Claude a toolbox — it picks up the right tool without you telling it which one to use.
Here are three tools we'll give our agent: get the current datetime, run a math calculation, and fetch weather data. In a real production system, the weather tool would hit a live API. Here I'm keeping it simple with a mock so you can run this without any extra API keys.
agent.py — tool definitionsimport os
import json
import math
import datetime
import anthropic
client = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY"))
# Tool schemas tell Claude what each tool does and what inputs it expects
TOOLS = [
{
"name": "get_current_datetime",
"description": "Returns the current date and time in a human-readable format.",
"input_schema": {
"type": "object",
"properties": {},
"required": []
}
},
{
"name": "calculate",
"description": (
"Evaluates a mathematical expression and returns the result. "
"Supports standard arithmetic, powers, and math functions like sqrt, log, sin, cos."
),
"input_schema": {
"type": "object",
"properties": {
"expression": {
"type": "string",
"description": "A valid Python math expression, e.g. '2 ** 10' or 'math.sqrt(144)'"
}
},
"required": ["expression"]
}
},
{
"name": "get_weather",
"description": "Returns current weather conditions for a given city.",
"input_schema": {
"type": "object",
"properties": {
"city": {
"type": "string",
"description": "The name of the city, e.g. 'Naples, FL'"
}
},
"required": ["city"]
}
}
]
def get_current_datetime() -> str:
now = datetime.datetime.now()
return now.strftime("Today is %A, %B %d, %Y. The current time is %I:%M %p.")
def calculate(expression: str) -> str:
try:
# Provide math module in eval context so Claude can use math.sqrt() etc.
result = eval(expression, {"__builtins__": {}}, {"math": math})
return f"Result: {result}"
except Exception as e:
return f"Calculation error: {str(e)}"
def get_weather(city: str) -> str:
# Mock response — swap this for a real weather API call in production
mock_data = {
"naples": "Sunny, 88°F, humidity 72%, light southwest wind.",
"miami": "Partly cloudy, 85°F, humidity 80%, calm wind.",
"new york": "Overcast, 74°F, humidity 60%, northeast wind 12 mph.",
}
key = city.lower().replace(",", "").split()[0]
weather = mock_data.get(key, f"Weather data not available for {city}.")
return f"Weather in {city}: {weather}"
def run_tool(tool_name: str, tool_input: dict) -> str:
"""Dispatch table — maps tool names to their Python functions."""
if tool_name == "get_current_datetime":
return get_current_datetime()
elif tool_name == "calculate":
return calculate(tool_input["expression"])
elif tool_name == "get_weather":
return get_weather(tool_input["city"])
else:
return f"Unknown tool: {tool_name}"eval()The
calculate tool uses eval() with a restricted namespace so Claude can't accidentally execute dangerous system calls. In a production agent, consider a proper expression parser like asteval or simpleeval instead.
Step 3: Implement the Agentic Loop
This is the core of the whole thing. An agentic loop is just a while loop that keeps sending messages to Claude, executing any tools it asks for, and feeding those results back — until Claude decides it has enough information to give a final answer.
Claude signals it's done by returning a stop_reason of "end_turn". If it wants to call a tool first, the stop_reason is "tool_use". The loop handles both.
class AgentLoop:
"""
Wraps the Claude API agentic loop.
Keeps sending messages until Claude produces a final text response.
"""
def __init__(self, system_prompt: str = None):
self.client = client
self.model = "claude-sonnet-4-6"
self.tools = TOOLS
self.system_prompt = system_prompt or (
"You are a helpful AI assistant with access to tools. "
"Use them whenever they would help you answer accurately. "
"Always be concise and direct."
)
def run(self, user_message: str) -> str:
"""Run the agentic loop for a single user query. Returns the final answer."""
messages = [{"role": "user", "content": user_message}]
print(f"\n🤔 User: {user_message}")
while True:
response = self.client.messages.create(
model=self.model,
max_tokens=1024,
system=self.system_prompt,
tools=self.tools,
messages=messages
)
# If Claude is done, extract and return the text response
if response.stop_reason == "end_turn":
final_text = ""
for block in response.content:
if hasattr(block, "text"):
final_text += block.text
print(f"\n✅ Agent: {final_text}")
return final_text
# Claude wants to call one or more tools
if response.stop_reason == "tool_use":
# Add Claude's response (including tool_use blocks) to message history
messages.append({"role": "assistant", "content": response.content})
# Collect results for all tool calls in this response
tool_results = []
for block in response.content:
if block.type == "tool_use":
tool_name = block.name
tool_input = block.input
print(f"\n🔧 Calling tool: {tool_name} | Input: {tool_input}")
result = run_tool(tool_name, tool_input)
print(f" ↳ Result: {result}")
tool_results.append({
"type": "tool_result",
"tool_use_id": block.id,
"content": result
})
# Send all tool results back to Claude in one user turn
messages.append({"role": "user", "content": tool_results})
else:
# Unexpected stop reason — break to avoid infinite loop
print(f"Unexpected stop_reason: {response.stop_reason}")
break
return "Agent loop ended without a final response."Step 4: Test Your Agent with Real Queries
Now let's wire it all together and run some real queries. The main() function below creates one agent instance and fires off three different test questions — one for each tool — so you can see the full loop in action.
def main():
agent = AgentLoop()
test_queries = [
"What day of the week is it today, and what time is it?",
"What is the square root of 1764, and what is 2 to the power of 16?",
"What's the weather like in Naples, FL right now?"
]
for query in test_queries:
agent.run(query)
print("\n" + "─" * 60)
if __name__ == "__main__":
main()Run it with python agent.py. Here's what the output looks like:
🤔 User: What day of the week is it today, and what time is it?
🔧 Calling tool: get_current_datetime | Input: {}
↳ Result: Today is Wednesday, August 13, 2026. The current time is 10:24 AM.
✅ Agent: Today is Wednesday, August 13, 2026, and the current time is 10:24 AM.
────────────────────────────────────────────────────────────
🤔 User: What is the square root of 1764, and what is 2 to the power of 16?
🔧 Calling tool: calculate | Input: {'expression': 'math.sqrt(1764)'}
↳ Result: Result: 42.0
🔧 Calling tool: calculate | Input: {'expression': '2 ** 16'}
↳ Result: Result: 65536
✅ Agent: The square root of 1764 is 42, and 2 to the power of 16 is 65,536.
────────────────────────────────────────────────────────────
🤔 User: What's the weather like in Naples, FL right now?
🔧 Calling tool: get_weather | Input: {'city': 'Naples, FL'}
↳ Result: Weather in Naples, FL: Sunny, 88°F, humidity 72%, light southwest wind.
✅ Agent: It's currently sunny in Naples, FL with a temperature of 88°F,
humidity at 72%, and a light southwest wind — a classic Southwest Florida day.Notice how Claude made two separate calculate calls in the second query without you telling it to. That's the agentic behavior — it figured out it needed two tool calls and ran both before answering.
How It Works
Here's what's actually happening under the hood. When you call agent.run(), the loop sends your message to Claude along with the list of tool schemas. Claude reads the schemas and decides whether it needs to use a tool to answer accurately.
If Claude wants a tool, it returns a tool_use content block with the tool name and arguments filled in. Your Python code runs the actual function, then sends the result back to Claude as a tool_result. Claude reads that result and either asks for another tool or produces its final text response.
The message history grows with each loop iteration — Claude sees the full conversation including every tool result. This is what lets it reason across multiple steps instead of treating each message as isolated. It's a simple pattern, but it's the foundation of basically every serious AI agent in production today.
Common Errors and Fixes
Error 1: AuthenticationError on startup
anthropic.AuthenticationError: Error code: 401 - {'type': 'error', 'error': {'type': 'authentication_error', 'message': 'invalid x-api-key'}}Fix: Your API key isn't loading from the environment. Run echo $ANTHROPIC_API_KEY (Mac/Linux) or echo $env:ANTHROPIC_API_KEY (PowerShell) to confirm it's set. If the output is blank, you need to export it again in the current terminal session — environment variables don't persist across sessions unless you add them to your shell profile.
Error 2: Infinite loop — agent never stops
# The loop keeps printing tool calls but never prints "✅ Agent:" # No exception is raised — it just runs forever
Fix: This usually means your tool_result messages aren't formatted correctly. Claude keeps calling tools because it never receives usable results. Double-check that each result in the tool_results list has the correct tool_use_id matching the block.id from Claude's response. A mismatch causes Claude to retry indefinitely. Adding a max_iterations counter to the loop is good practice for production code.
Error 3: ValueError from eval() in the calculate tool
Calculation error: name 'sqrt' is not defined
Fix: Claude sometimes writes sqrt(144) instead of math.sqrt(144). The restricted eval namespace only exposes the math module, so bare function names fail. You can fix this by adding the math module's functions directly to the eval namespace: replace {"math": math} with {**vars(math), "math": math}. That makes both sqrt(144) and math.sqrt(144) work.
Next Steps
You've got a working agent — here's where to take it next.
- Add memory: Right now each call to
agent.run()starts fresh. Store themessageslist between calls to give your agent conversation memory that persists across questions. - Connect real APIs: Swap the mock weather function for a real call to the OpenWeatherMap API or any REST service your business uses. The tool pattern stays exactly the same.
- Add more tools: Give your agent the ability to search the web, read files, write to a database, or send emails. Each new capability is just a new tool schema plus a Python function.
- Build a web interface: Wrap the
AgentLoopclass in a FastAPI endpoint and connect it to a simple front-end. You now have a deployable AI assistant your team can actually use.
Frequently Asked Questions
What is an agentic loop in Claude API?
An agentic loop is a while loop that repeatedly sends messages to Claude, executes any tools it requests, and feeds those results back until Claude returns a final answer with stop_reason: "end_turn". It's the core pattern behind autonomous AI agents — the loop is what lets the model take multiple steps instead of just responding once.
How does Claude tool calling work with the Anthropic SDK?
You pass a list of tool schemas (JSON objects describing each tool's name, purpose, and expected inputs) to the messages.create() call. When Claude decides to use a tool, it returns a tool_use content block. Your code runs the actual function, then sends a tool_result message back. Claude reads the result and continues reasoning from there.
Which Claude model should I use for building AI agents in 2026?
claude-sonnet-4-6 is the right default for most agent use cases. It's fast, cost-effective, and handles multi-step tool calling reliably. If your agent needs to work through very complex, long-horizon reasoning tasks, claude-opus-4-5 gives you more reasoning power at higher cost per token.
Can I build AI agents with Claude API without LangChain?
Yes — and for many use cases, doing it directly with the Anthropic SDK is simpler and easier to debug. LangChain adds abstraction layers that can make troubleshooting harder when something goes wrong. The pattern in this tutorial gives you full visibility into every message Claude sends and receives, which is exactly what you want when you're building and testing.
How do I give my Claude agent persistent memory between conversations?
The simplest approach is to store the messages list to a JSON file or a database after each conversation, then load it back in at the start of the next one. For a production app, you'd typically store messages in PostgreSQL or a similar database keyed to a user or session ID, and pass the last N messages to each API call to avoid exceeding the context window.
Conclusion
You now have a fully working AI agent built on the Claude API — one that reasons through multi-step problems, calls real tools, and loops until it has an answer worth giving. This same pattern is what powers the custom AI agents we build at Naples AI for businesses across Southwest Florida. From automating real estate listing workflows to intelligent process automation for manufacturers and restaurants, the agentic loop you just wrote is the foundation underneath all of it. If you want to take what you've built here and turn it into something your business actually runs on, book a free 30-minute call with Chris and let's figure out what that looks like for you.