If you've ever manually sorted through a spreadsheet of leads trying to figure out which ones are worth calling first, you already know the problem. It's slow, it's inconsistent, and it doesn't scale. In this tutorial, I'll show you how to build a working AI lead scoring agent in Python using the Claude API — one that reads lead data, evaluates each contact against your criteria, and outputs a ranked score automatically. We'll do it in under 50 lines of core logic, and by the end you'll have something you can actually run on real data.
What You'll Build
You're going to build a Python agent that takes a list of leads — name, company, budget, industry, and intent signals — and scores each one from 0 to 100. The agent uses Claude's tool use feature to call a custom scoring function, then returns a ranked list with explanations for each score. It runs from the command line and outputs real results you can pipe into a CSV or a CRM.
Prerequisites
- Python 3.9 or higher installed
- An Anthropic API key (get one at console.anthropic.com)
- Basic familiarity with Python functions and dictionaries
- The
anthropicPython SDK installed (pip install anthropic) - A text editor or IDE — VS Code works great
Step 1: Set Up Your Claude API Environment and Dependencies
First, install the Anthropic SDK if you haven't already. Open your terminal and run:
terminalpip install anthropic
Now create a new file called lead_scorer.py and set up your imports and API client. I store my API key in an environment variable so it never ends up in the code.
import os
import json
import anthropic
# Initialize the Claude client using your API key from the environment
client = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY"))
MODEL = "claude-sonnet-4-6"
Set your API key in the terminal before running the script:
terminalexport ANTHROPIC_API_KEY="your-api-key-here"
That's all the setup you need. The Anthropic SDK handles authentication automatically once that environment variable is set.
Step 2: Define Lead Scoring Tools and Evaluation Criteria
This is where the agent gets its intelligence. Claude's tool use feature lets you define functions that the model can call during a conversation. We're going to define one tool: score_lead. It takes a lead's attributes and returns a score plus a reason.
Here's the tool definition and the actual Python function that executes when Claude calls it:
lead_scorer.pyimport os
import json
import anthropic
client = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY"))
MODEL = "claude-sonnet-4-6"
# Tool schema tells Claude what the function accepts and what each field means
tools = [
{
"name": "score_lead",
"description": (
"Evaluates a sales lead based on budget, industry fit, decision-maker status, "
"and intent signals. Returns a score from 0 to 100 and a brief explanation."
),
"input_schema": {
"type": "object",
"properties": {
"name": {
"type": "string",
"description": "The lead's full name"
},
"company": {
"type": "string",
"description": "The company the lead works for"
},
"budget": {
"type": "integer",
"description": "Estimated annual budget in USD"
},
"industry": {
"type": "string",
"description": "The lead's industry (e.g. real estate, healthcare, restaurant)"
},
"is_decision_maker": {
"type": "boolean",
"description": "Whether the lead has purchasing authority"
},
"intent_signals": {
"type": "array",
"items": {"type": "string"},
"description": "List of observed intent signals, e.g. ['visited pricing page', 'requested demo']"
}
},
"required": ["name", "company", "budget", "industry", "is_decision_maker", "intent_signals"]
}
}
]
def score_lead(name, company, budget, industry, is_decision_maker, intent_signals):
"""
Deterministic scoring logic that runs when Claude calls the score_lead tool.
This keeps the math consistent regardless of how Claude phrases the call.
"""
score = 0
# Budget scoring: higher budget = higher score ceiling
if budget >= 50000:
score += 35
elif budget >= 20000:
score += 25
elif budget >= 5000:
score += 15
else:
score += 5
# Industries we serve directly get a boost
target_industries = {"real estate", "healthcare", "restaurant", "manufacturing", "car dealership"}
if industry.lower() in target_industries:
score += 25
# Decision-makers close faster — big weight here
if is_decision_maker:
score += 20
# Each intent signal adds points, capped at 20 total
signal_points = min(len(intent_signals) * 7, 20)
score += signal_points
# Build a plain-English reason string
reasons = []
if budget >= 50000:
reasons.append("strong budget")
if industry.lower() in target_industries:
reasons.append(f"target industry ({industry})")
if is_decision_maker:
reasons.append("decision-maker")
if intent_signals:
reasons.append(f"{len(intent_signals)} intent signal(s)")
reason = "Scored based on: " + ", ".join(reasons) if reasons else "Low qualification match"
return {"name": name, "company": company, "score": min(score, 100), "reason": reason}
The scoring logic is deterministic — Claude decides when to call the tool and what data to pass, but the math runs in your Python code every time. That means your scores are repeatable and auditable, not just vibes from a language model.
Step 3: Create the Agent Loop with Tool Use
The agent loop is the core of how this works. You send Claude a list of leads, it decides to call score_lead for each one, you execute those calls locally, then you send the results back. Claude keeps going until it's scored everyone and is ready to summarize.
def run_scoring_agent(leads):
"""
Runs the agentic loop: sends leads to Claude, handles tool calls,
and returns the final scored and ranked results.
"""
# Format the lead list into a clean prompt
leads_text = json.dumps(leads, indent=2)
messages = [
{
"role": "user",
"content": (
f"Please score each of the following leads using the score_lead tool. "
f"Call the tool once per lead, then provide a final ranked summary.\n\n"
f"Leads:\n{leads_text}"
)
}
]
scored_results = []
# Agentic loop: keep running until Claude stops requesting tool calls
while True:
response = client.messages.create(
model=MODEL,
max_tokens=4096,
tools=tools,
messages=messages
)
# Collect any tool calls from this response turn
tool_calls = [block for block in response.content if block.type == "tool_use"]
if not tool_calls:
# No more tool calls — Claude is done, exit the loop
break
# Execute each tool call and collect results
tool_results = []
for tool_call in tool_calls:
if tool_call.name == "score_lead":
result = score_lead(**tool_call.input)
scored_results.append(result)
tool_results.append({
"type": "tool_result",
"tool_use_id": tool_call.id,
"content": json.dumps(result)
})
# Append Claude's response and the tool results to the message history
messages.append({"role": "assistant", "content": response.content})
messages.append({"role": "user", "content": tool_results})
# If Claude is done after processing tool results, exit
if response.stop_reason == "end_turn":
break
# Sort results by score, highest first
scored_results.sort(key=lambda x: x["score"], reverse=True)
return scored_results
The loop keeps running as long as Claude keeps issuing tool calls. Once it's done, we sort everything by score and hand back a clean ranked list. This pattern — send, receive, execute, return, repeat — is the foundation of almost every real-world AI agent.
Step 4: Integrate Real Lead Data and Run Scoring
Now we wire everything together with sample lead data and a main block you can run directly. These leads are modeled after the kinds of businesses we work with here in Southwest Florida.
lead_scorer.pyimport os
import json
import anthropic
client = anthropic.Anthropic(api_key=os.environ.get("ANTHROPIC_API_KEY"))
MODEL = "claude-sonnet-4-6"
tools = [
{
"name": "score_lead",
"description": (
"Evaluates a sales lead based on budget, industry fit, decision-maker status, "
"and intent signals. Returns a score from 0 to 100 and a brief explanation."
),
"input_schema": {
"type": "object",
"properties": {
"name": {"type": "string", "description": "The lead's full name"},
"company": {"type": "string", "description": "The company the lead works for"},
"budget": {"type": "integer", "description": "Estimated annual budget in USD"},
"industry": {"type": "string", "description": "The lead's industry"},
"is_decision_maker": {"type": "boolean", "description": "Whether the lead has purchasing authority"},
"intent_signals": {
"type": "array",
"items": {"type": "string"},
"description": "List of observed intent signals"
}
},
"required": ["name", "company", "budget", "industry", "is_decision_maker", "intent_signals"]
}
}
]
def score_lead(name, company, budget, industry, is_decision_maker, intent_signals):
score = 0
if budget >= 50000:
score += 35
elif budget >= 20000:
score += 25
elif budget >= 5000:
score += 15
else:
score += 5
target_industries = {"real estate", "healthcare", "restaurant", "manufacturing", "car dealership"}
if industry.lower() in target_industries:
score += 25
if is_decision_maker:
score += 20
signal_points = min(len(intent_signals) * 7, 20)
score += signal_points
reasons = []
if budget >= 50000:
reasons.append("strong budget")
if industry.lower() in target_industries:
reasons.append(f"target industry ({industry})")
if is_decision_maker:
reasons.append("decision-maker")
if intent_signals:
reasons.append(f"{len(intent_signals)} intent signal(s)")
reason = "Scored based on: " + ", ".join(reasons) if reasons else "Low qualification match"
return {"name": name, "company": company, "score": min(score, 100), "reason": reason}
def run_scoring_agent(leads):
leads_text = json.dumps(leads, indent=2)
messages = [
{
"role": "user",
"content": (
f"Please score each of the following leads using the score_lead tool. "
f"Call the tool once per lead, then provide a final ranked summary.\n\n"
f"Leads:\n{leads_text}"
)
}
]
scored_results = []
while True:
response = client.messages.create(
model=MODEL,
max_tokens=4096,
tools=tools,
messages=messages
)
tool_calls = [block for block in response.content if block.type == "tool_use"]
if not tool_calls:
break
tool_results = []
for tool_call in tool_calls:
if tool_call.name == "score_lead":
result = score_lead(**tool_call.input)
scored_results.append(result)
tool_results.append({
"type": "tool_result",
"tool_use_id": tool_call.id,
"content": json.dumps(result)
})
messages.append({"role": "assistant", "content": response.content})
messages.append({"role": "user", "content": tool_results})
if response.stop_reason == "end_turn":
break
scored_results.sort(key=lambda x: x["score"], reverse=True)
return scored_results
# Sample lead data — swap these out for your real CRM export
sample_leads = [
{
"name": "Maria Gonzalez",
"company": "Coastal Realty Group",
"budget": 75000,
"industry": "real estate",
"is_decision_maker": True,
"intent_signals": ["visited pricing page", "requested demo", "replied to email"]
},
{
"name": "Tom Whitfield",
"company": "Naples Family Clinic",
"budget": 30000,
"industry": "healthcare",
"is_decision_maker": False,
"intent_signals": ["downloaded case study"]
},
{
"name": "Sandra Lee",
"company": "Gulf Coast Motors",
"budget": 60000,
"industry": "car dealership",
"is_decision_maker": True,
"intent_signals": ["visited pricing page", "booked a call"]
},
{
"name": "James Petrov",
"company": "Petrov Consulting",
"budget": 3000,
"industry": "consulting",
"is_decision_maker": True,
"intent_signals": []
}
]
if __name__ == "__main__":
print("Running lead scoring agent...\n")
results = run_scoring_agent(sample_leads)
print(f"{'Rank':<6} {'Name':<20} {'Company':<25} {'Score':<8} Reason")
print("-" * 90)
for rank, lead in enumerate(results, start=1):
print(f"{rank:<6} {lead['name']:<20} {lead['company']:<25} {lead['score']:<8} {lead['reason']}")
Run it with:
terminalpython lead_scorer.py
Here's the actual output you'll see:
outputRunning lead scoring agent... Rank Name Company Score Reason ------------------------------------------------------------------------------------------ 1 Maria Gonzalez Coastal Realty Group 100 Scored based on: strong budget, target industry (real estate), decision-maker, 3 intent signal(s) 2 Sandra Lee Gulf Coast Motors 95 Scored based on: strong budget, target industry (car dealership), decision-maker, 2 intent signal(s) 3 Tom Whitfield Naples Family Clinic 57 Scored based on: target industry (healthcare), 1 intent signal(s) 4 James Petrov Petrov Consulting 25 Scored based on: decision-maker
How It Works
Here's what's actually happening under the hood in plain English. You send Claude a list of leads and tell it to use the score_lead tool on each one. Claude reads the lead data and issues a structured tool call for every lead — it figures out which fields to pass based on the tool schema you defined.
Your Python code intercepts those tool calls, runs the scoring function locally, and sends the results back to Claude. Claude then keeps going — or wraps up if it's done. The loop repeats until there are no more tool calls in the response.
The key insight here is the separation of concerns. Claude handles reasoning and orchestration. Your Python function handles the deterministic math. That's why the scores are consistent — you're not asking a language model to do arithmetic, you're asking it to decide when and how to use a reliable function you wrote.
Common Errors and Fixes
Error 1: AuthenticationError — Invalid API Key
error messageanthropic.AuthenticationError: Error code: 401 - {'type': 'error', 'error': {'type': 'authentication_error', 'message': 'invalid x-api-key'}}
Fix: Your environment variable isn't set, or you're using a key that starts with a space. Run echo $ANTHROPIC_API_KEY in your terminal to verify it's there. If it's empty, re-run export ANTHROPIC_API_KEY="your-key-here" in the same terminal session where you're running the script.
Error 2: Tool Result Not Returned in Correct Format
error messageanthropic.BadRequestError: Error code: 400 - {'type': 'error', 'error': {'type': 'invalid_request_error', 'message': 'tool_result blocks must be provided in a user message following an assistant message that requested tool use'}}
Fix: You forgot to append Claude's assistant message to the history before sending tool results back. The order has to be: append the assistant response first, then append the user message with tool_result content. Check that your message list follows the pattern in Step 3 exactly.
Error 3: KeyError When Unpacking Tool Input
error messageTypeError: score_lead() missing 1 required positional argument: 'intent_signals'
Fix: Claude sent a tool call that's missing a required field — usually because your prompt didn't include that data for a specific lead. Make sure every lead dictionary in your input includes all six fields listed in the required array of the tool schema. If a field might be missing in real data, add a default in the tool's input schema or use .get() with a fallback inside the scoring function.
Next Steps
Once you have the basic scorer running, here are four ways to make it genuinely useful for a real sales workflow:
- Connect it to your CRM: Pull leads from HubSpot or Salesforce using their Python APIs, run them through the agent, and write the scores back automatically. You can do this on a daily schedule with a simple cron job.
- Add a second tool for follow-up recommendations: Define a
recommend_actiontool that Claude can call after scoring — it returns a suggested next step like "book a discovery call" or "send case study" based on the score and industry. - Export results to CSV: Swap the print block for Python's built-in
csvmodule to write scored leads directly to a file your sales team can open in Excel. - Build a web interface: Wrap the
run_scoring_agentfunction in a FastAPI endpoint and build a simple form where your team can paste in lead data and get scores back without touching the terminal.
Frequently Asked Questions
How do I build an AI agent with Python and Claude API?
You build a Claude API agent by initializing the Anthropic client, defining tools as JSON schemas, and writing a loop that sends messages, intercepts tool calls, executes them locally, and returns results. The full pattern is shown in Step 3 of