If you're spending hours manually reviewing leads, copying data into spreadsheets, and still guessing which prospects are worth your time — you're leaving money on the table. An AI lead qualification chatbot can do that triage automatically, score every lead against your criteria, and hand you a ranked list with reasoning attached. That's exactly what we're building today.
In this tutorial you'll build a production-ready lead qualifier agent in Python using the Claude API. It handles multi-turn conversations, extracts structured contact data, scores leads on a weighted rubric, and outputs a clear qualify/disqualify decision with a confidence score — all in under 200 lines of code.
What You'll Build
By the end of this tutorial you'll have a working Python agent that conducts a natural sales discovery conversation with any prospect. It extracts key qualification data automatically using tool calls, then scores the lead across four dimensions and returns a final decision your sales team can act on immediately. The whole thing runs in the terminal and is ready to drop behind a web endpoint or webhook.
Prerequisites
- Python 3.10 or higher installed
- An Anthropic API key (get one at console.anthropic.com)
- Basic familiarity with Python classes and dictionaries
- The
anthropicPython SDK (pip install anthropic) - Optional but helpful: understanding of how tool use / function calling works in LLM APIs
The complete, working code for this agent is built step-by-step in the sections below. Every snippet connects to the next — by Step 6 you'll have the entire file assembled and ready to run. Copy each block in order, or scroll to the bottom of Step 6 for the fully assembled version.
Step 1: Set Up Claude API and Your Python Environment
First, install the Anthropic SDK and set your API key. I keep mine in a .env file so it never touches the source code.
pip install anthropic python-dotenv
Create a .env file in your project root:
ANTHROPIC_API_KEY=sk-ant-your-key-here
Now create the main file and verify the SDK connects correctly before writing anything else. This quick sanity check saves a lot of debugging time later.
test_connection.pyimport os
from dotenv import load_dotenv
import anthropic
load_dotenv()
client = anthropic.Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"])
response = client.messages.create(
model="claude-sonnet-4-5",
max_tokens=64,
messages=[{"role": "user", "content": "Say: API connected."}]
)
print(response.content[0].text)If you see API connected. in your terminal, you're good to go. If you see an authentication error, double-check your key has no extra spaces in the .env file.
Step 2: Define Lead Qualification Criteria and Scoring Rubric
Before writing any agent code, nail down what a "qualified lead" actually means for your business. For this tutorial I'm using a B2B SaaS rubric, but the weights and criteria are easy to swap for real estate, healthcare, or any other industry Naples AI works in.
Our scoring covers four dimensions, each worth up to 25 points for a maximum score of 100:
- Budget fit (0–25): Does their budget align with your pricing?
- Authority (0–25): Are they a decision-maker or just a researcher?
- Need (0–25): Do they have a specific, urgent problem you solve?
- Timeline (0–25): Are they ready to move in a reasonable window?
This is the classic BANT framework, and it maps perfectly to tool-based extraction. Threshold: 65+ = Qualified, 40–64 = Nurture, below 40 = Disqualify.
lead_qualifier.pyimport os
import json
from dotenv import load_dotenv
import anthropic
load_dotenv()
# Scoring thresholds — adjust these for your sales process
QUALIFIED_THRESHOLD = 65
NURTURE_THRESHOLD = 40
SCORING_RUBRIC = {
"budget": {
"description": "Does prospect budget align with our pricing?",
"max_points": 25,
"tiers": {
"high": {"range": "50k+", "points": 25},
"medium": {"range": "10k-50k", "points": 15},
"low": {"range": "under 10k", "points": 5},
"unknown": {"range": "not disclosed", "points": 0}
}
},
"authority": {
"description": "Is prospect a decision maker?",
"max_points": 25,
"tiers": {
"decision_maker": {"role": "owner/C-suite/VP", "points": 25},
"influencer": {"role": "manager/director", "points": 15},
"researcher": {"role": "individual contributor", "points": 5},
"unknown": {"role": "not identified", "points": 0}
}
},
"need": {
"description": "Urgency and clarity of the problem they need solved",
"max_points": 25,
"tiers": {
"urgent": {"description": "active pain point, needs solution now", "points": 25},
"moderate": {"description": "has a problem, exploring options", "points": 15},
"low": {"description": "curious but no clear problem", "points": 5},
"unknown": {"description": "need not identified", "points": 0}
}
},
"timeline": {
"description": "How soon are they ready to make a decision?",
"max_points": 25,
"tiers": {
"immediate": {"window": "within 30 days", "points": 25},
"near_term": {"window": "1-3 months", "points": 15},
"long_term": {"window": "3-6 months", "points": 5},
"unknown": {"window": "not specified", "points": 0}
}
}
}Step 3: Create Tool Definitions for Lead Data Extraction
This is where the magic happens. Instead of parsing free-form text with regex, we give Claude three tools and let it decide when to call them during the conversation. The model will call extract_contact_info once it knows who it's talking to, assess_fit as qualification signals emerge, and score_lead when it has enough data to make a final decision.
Tool definitions in the Anthropic SDK are just Python dictionaries with a specific schema. Keep the descriptions precise — they directly affect how reliably the model calls the right tool at the right time.
lead_qualifier.py (continued)# Tool definitions passed to Claude on every API call
TOOLS = [
{
"name": "extract_contact_info",
"description": (
"Extract and store structured contact information from the conversation. "
"Call this as soon as you have identified the prospect's name, company, "
"email, phone, or job title. Partial information is fine."
),
"input_schema": {
"type": "object",
"properties": {
"name": {
"type": "string",
"description": "Prospect's full name"
},
"company": {
"type": "string",
"description": "Company or organization name"
},
"email": {
"type": "string",
"description": "Email address"
},
"phone": {
"type": "string",
"description": "Phone number"
},
"job_title": {
"type": "string",
"description": "Prospect's job title or role"
}
},
"required": []
}
},
{
"name": "assess_fit",
"description": (
"Assess one or more BANT qualification dimensions based on what the "
"prospect has shared. Call this whenever the prospect reveals information "
"about their budget, authority, need, or timeline. You can call it multiple "
"times as new information emerges."
),
"input_schema": {
"type": "object",
"properties": {
"budget_tier": {
"type": "string",
"enum": ["high", "medium", "low", "unknown"],
"description": "Budget alignment tier based on SCORING_RUBRIC"
},
"authority_tier": {
"type": "string",
"enum": ["decision_maker", "influencer", "researcher", "unknown"],
"description": "Decision-making authority level"
},
"need_tier": {
"type": "string",
"enum": ["urgent", "moderate", "low", "unknown"],
"description": "Urgency and clarity of the prospect's need"
},
"timeline_tier": {
"type": "string",
"enum": ["immediate", "near_term", "long_term", "unknown"],
"description": "Decision timeline"
},
"notes": {
"type": "string",
"description": "Brief notes on what the prospect said that informed this assessment"
}
},
"required": ["notes"]
}
},
{
"name": "score_lead",
"description": (
"Calculate the final lead score and produce a qualification decision. "
"Call this ONLY when you have gathered sufficient information across all "
"four BANT dimensions, or when the prospect indicates they are done. "
"This ends the qualification conversation."
),
"input_schema": {
"type": "object",
"properties": {
"budget_tier": {
"type": "string",
"enum": ["high", "medium", "low", "unknown"]
},
"authority_tier": {
"type": "string",
"enum": ["decision_maker", "influencer", "researcher", "unknown"]
},
"need_tier": {
"type": "string",
"enum": ["urgent", "moderate", "low", "unknown"]
},
"timeline_tier": {
"type": "string",
"enum": ["immediate", "near_term", "long_term", "unknown"]
},
"summary": {
"type": "string",
"description": "One or two sentence summary of the prospect and their situation"
}
},
"required": [
"budget_tier",
"authority_tier",
"need_tier",
"timeline_tier",
"summary"
]
}
}
]Claude decides when to call a tool based entirely on its description. Vague descriptions lead to missed calls or calls at the wrong time. Be specific about the trigger condition — "Call this as soon as you have identified…" works better than "Use this to collect contact info."
Step 4: Build the Main Agent Loop with Multi-Turn Conversation
Now we build the LeadQualifierAgent class. The core of it is a loop that sends messages to Claude, handles tool calls when Claude decides to use them, and keeps the conversation going until the score_lead tool fires. The system prompt is where you define the agent's persona and behavior — I keep mine tight and direct.
class LeadQualifierAgent:
"""
Conversational AI agent that qualifies sales leads using BANT scoring.
Maintains full conversation history and lead context across turns.
"""
def __init__(self):
self.client = anthropic.Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"])
self.model = "claude-sonnet-4-5"
self.conversation_history = []
self.lead_data = {
"contact": {},
"bant": {
"budget_tier": "unknown",
"authority_tier": "unknown",
"need_tier": "unknown",
"timeline_tier": "unknown"
},
"notes": [],
"qualified": None,
"score": 0,
"confidence": 0.0,
"summary": ""
}
self.qualification_complete = False
self.system_prompt = """You are Alex, a friendly and professional sales development representative
for Naples AI, a custom AI solutions agency in Southwest Florida. Your job is to have a
natural discovery conversation with prospects and qualify them as leads.
Your goals during the conversation:
1. Make the prospect feel heard and comfortable — this is a conversation, not an interrogation.
2. Uncover their budget, decision-making authority, specific business problem, and buying timeline.
3. Use the available tools to record contact information and assess fit as you learn things.
4. Ask one question at a time. Never ask multiple questions in the same message.
5. When you have enough information across all four BANT dimensions (or after 8-10 exchanges),
call the score_lead tool to finalize qualification.
Naples AI services include: custom AI development, intelligent process automation, AI chatbots,
predictive analytics, AI-powered SEO, computer vision, AI knowledge bases, and real estate
listing automation. We serve real estate, healthcare, restaurants, car dealerships, and
manufacturing businesses in Southwest Florida and beyond.
Keep your responses conversational and under 3 sentences. Do not reveal the scoring rubric
or that you are evaluating them — just have a genuine conversation."""
def _call_claude(self):
"""Send current conversation history to Claude and return the response."""
response = self.client.messages.create(
model=self.model,
max_tokens=1024,
system=self.system_prompt,
tools=TOOLS,
messages=self.conversation_history
)
return response
def _handle_tool_call(self, tool_name, tool_input):
"""
Execute the appropriate tool and update lead_data.
Returns a result string that gets sent back to Claude as the tool result.
"""
if tool_name == "extract_contact_info":
# Merge new contact data — don't overwrite existing values with blanks
for key, value in tool_input.items():
if value:
self.lead_data["contact"][key] = value
return json.dumps({"status": "contact_info_saved", "data": self.lead_data["contact"]})
elif tool_name == "assess_fit":
# Update BANT tiers only when a non-unknown value is provided
for dimension in ["budget_tier", "authority_tier", "need_tier", "timeline_tier"]:
if dimension in tool_input and tool_input[dimension] != "unknown":
self.lead_data["bant"][dimension] = tool_input[dimension]
if "notes" in tool_input and tool_input["notes"]:
self.lead_data["notes"].append(tool_input["notes"])
return json.dumps({"status": "fit_assessed", "current_bant": self.lead_data["bant"]})
elif tool_name == "score_lead":
# Finalize BANT tiers from this call
for dimension in ["budget_tier", "authority_tier", "need_tier", "timeline_tier"]:
if dimension in tool_input:
self.lead_data["bant"][dimension] = tool_input[dimension]
self.lead_data["summary"] = tool_input.get("summary", "")
# Scoring and decision happen in _calculate_score
result = self._calculate_score()
self.qualification_complete = True
return json.dumps(result)
return json.dumps({"status": "unknown_tool"})
def _process_response(self, response):
"""
Process Claude's response: handle text output and tool calls.
Appends assistant turn and tool results to conversation history.
Returns the text portion of the response (if any) for display.
"""
assistant_message = {"role": "assistant", "content": response.content}
self.conversation_history.append(assistant_message)
tool_results = []
display_text = ""
for block in response.content:
if block.type == "text":
display_text = block.text
elif block.type == "tool_use":
result = self._handle_tool_call(block.name, block.input)
tool_results.append({
"type": "tool_result",
"tool_use_id": block.id,
"content": result
})
# If tools were called, append results and get Claude's follow-up response
if tool_results:
self.conversation_history.append({
"role": "user",
"content": tool_results
})
# Only recurse if qualification isn't complete yet
if not self.qualification_complete:
follow_up = self._call_claude()
return self._process_response(follow_up)
return display_text
def chat(self, user_message):
"""
Send a user message, get agent response.
Returns (agent_reply_text, qualification_complete).
"""
self.conversation_history.append({
"role": "user",
"content": user_message
})
response = self._call_claude()
reply = self._process_response(response)
return reply, self.qualification_completeStep 5: Implement Scoring Logic and Qualification Decision
The _calculate_score method is where raw BANT tiers become a numeric score, a decision label, and a confidence percentage. I calculate confidence based on how many dimensions had actual data vs. unknown — unknown dimensions lower your confidence even if the score is high enough to qualify.
def _calculate_score(self):
"""
Calculate final lead score from BANT tiers using SCORING_RUBRIC.
Returns a dict with score, decision, confidence, and reasoning.
"""
bant = self.lead_data["bant"]
total_score = 0
score_breakdown = {}
known_dimensions = 0
dimension_map = {
"budget_tier": "budget",
"authority_tier": "authority",
"need_tier": "need",
"timeline_tier": "timeline"
}
for bant_key, rubric_key in dimension_map.items():
tier = bant.get(bant_key, "unknown")
rubric = SCORING_RUBRIC[rubric_key]
tiers = rubric["tiers"]
# Look up points for this tier — default to 0 if tier not found
points = tiers.get(tier, tiers["unknown"]).get("points", 0)
total_score += points
score_breakdown[rubric_key] = {"tier": tier, "points": points}
if tier != "unknown":
known_dimensions += 1
# Confidence: percentage of dimensions with actual data, scaled to how strong the score is
data_completeness = known_dimensions / 4.0
score_normalized = total_score / 100.0
confidence = round((data_completeness * 0.6 + score_normalized * 0.4) * 100, 1)
# Qualification decision
if total_score >= QUALIFIED_THRESHOLD:
decision = "QUALIFIED"
next_action = "Schedule a discovery call immediately."
elif total_score >= NURTURE_THRESHOLD:
decision = "NURTURE"
next_action = "Add to email nurture sequence. Follow up in 2 weeks."
else:
decision = "DISQUALIFY"
next_action = "Not a fit at this time. Log and archive."
result = {
"decision": decision,
"score": total_score,
"confidence": confidence,
"breakdown": score_breakdown,
"next_action": next_action,
"summary": self.lead_data["summary"],
"contact": self.lead_data["contact"],
"notes": self.lead_data["notes"]
}
# Store final values back on the lead object
self.lead_data["qualified"] = decision
self.lead_data["score"] = total_score
self.lead_data["confidence"] = confidence
return result
def get_qualification_report(self):
"""Return a formatted string report of the final lead qualification."""
if not self.qualification_complete:
return "Qualification not yet complete."
bant = self.lead_data["bant"]
contact = self.lead_data["contact"]
score = self.lead_data["score"]
decision = self.lead_data["qualified"]
confidence = self.lead_data["confidence"]
report = f"""
╔══════════════════════════════════════════════════════╗
║ LEAD QUALIFICATION REPORT ║
╚══════════════════════════════════════════════════════╝
CONTACT
Name: {contact.get('name', 'Unknown')}
Company: {contact.get('company', 'Unknown')}
Email: {contact.get('email', 'Not provided')}
Title: {contact.get('job_title', 'Unknown')}
BANT SCORES
Budget: {bant['