← Back to Blog

What You'll Build

If you've ever wasted an hour on a sales call that went nowhere, you already know why lead qualification matters. In this developer tutorial, you're going to build a multi-agent system using the Claude API that automatically scores and qualifies B2B leads — two agents working together, one asking the right questions and one doing the scoring math.

By the end, you'll have a working Python script around 150 lines long that takes raw lead data, runs it through a qualification pipeline, and spits out a structured score with a recommended action. This is the exact kind of Claude API multi-agent system we build for clients at Naples AI when they need AI lead qualification automation that actually fits their sales process.

📦 Full Source Code Notice: The complete working code is broken into numbered steps below. Each step builds on the last, so by Step 5 you'll have the entire system assembled and ready to run. Copy each block in order and you're good to go.

Prerequisites

  • Python 3.10 or higher installed
  • An Anthropic API key (get one at console.anthropic.com)
  • Basic familiarity with Python classes and async concepts
  • anthropic and pydantic packages installed (pip install anthropic pydantic)
  • A terminal and a text editor — that's genuinely all you need

Step 1: Set Up Claude API and Project Structure

Start by creating a project folder and setting your API key as an environment variable. You never want to hardcode credentials — that's how keys end up on GitHub.

terminal
mkdir lead-qualifier
cd lead-qualifier
export ANTHROPIC_API_KEY="your-api-key-here"
touch main.py models.py agents.py tools.py

Now install your dependencies. We're using Pydantic to define lead data schemas cleanly, which makes the tool definitions much easier to read later.

terminal
pip install anthropic pydantic

Next, set up the base data models. These represent the lead coming into the system and the qualification result coming out. Keeping these in their own file makes the whole project easier to reason about.

models.py
from pydantic import BaseModel, Field
from typing import Optional


class LeadInput(BaseModel):
    """Represents a raw B2B lead entering the qualification pipeline."""
    company_name: str
    contact_name: str
    email: str
    annual_revenue: Optional[str] = None
    employee_count: Optional[int] = None
    industry: Optional[str] = None
    pain_point: Optional[str] = None
    budget_mentioned: Optional[str] = None
    timeline: Optional[str] = None
    website: Optional[str] = None


class QualificationResult(BaseModel):
    """The structured output from the full qualification pipeline."""
    lead_name: str
    company: str
    qualification_status: str  # "qualified", "unqualified", "needs_nurture"
    score: int = Field(ge=0, le=100)
    score_breakdown: dict
    recommended_action: str
    reasoning: str
    next_steps: list[str]

That's your foundation. LeadInput captures everything a CRM form might collect, and QualificationResult is what your sales team actually sees at the end.

Step 2: Define Lead Qualification Tools and Schemas

This is where the Claude API multi-agent system gets interesting. Tools are how you tell Claude what functions it's allowed to call — think of them as the agent's action menu. We're defining three tools: one to analyze firmographic fit, one to assess budget and timeline, and one to evaluate pain point severity.

tools.py
import anthropic
from models import LeadInput


# Tool schemas tell Claude the exact shape of data it needs to pass back
QUALIFICATION_TOOLS = [
    {
        "name": "analyze_firmographic_fit",
        "description": (
            "Analyzes whether the lead's company size, industry, and revenue "
            "match the ideal customer profile. Returns a fit score from 0-40."
        ),
        "input_schema": {
            "type": "object",
            "properties": {
                "company_name": {
                    "type": "string",
                    "description": "Name of the company being evaluated"
                },
                "industry": {
                    "type": "string",
                    "description": "Industry vertical of the company"
                },
                "employee_count": {
                    "type": "integer",
                    "description": "Number of employees at the company"
                },
                "annual_revenue": {
                    "type": "string",
                    "description": "Approximate annual revenue of the company"
                },
                "fit_notes": {
                    "type": "string",
                    "description": "Brief analysis of why this company does or does not fit"
                }
            },
            "required": ["company_name", "fit_notes"]
        }
    },
    {
        "name": "assess_budget_and_timeline",
        "description": (
            "Evaluates whether the lead has the budget and urgency to move forward. "
            "Returns a readiness score from 0-35."
        ),
        "input_schema": {
            "type": "object",
            "properties": {
                "budget_signal": {
                    "type": "string",
                    "enum": ["strong", "moderate", "weak", "unknown"],
                    "description": "Strength of the budget signal from the lead"
                },
                "timeline_signal": {
                    "type": "string",
                    "enum": ["immediate", "within_quarter", "within_year", "unknown"],
                    "description": "How soon the lead needs a solution"
                },
                "readiness_notes": {
                    "type": "string",
                    "description": "Explanation of the budget and timeline assessment"
                }
            },
            "required": ["budget_signal", "timeline_signal", "readiness_notes"]
        }
    },
    {
        "name": "evaluate_pain_point",
        "description": (
            "Rates the severity and clarity of the lead's stated problem. "
            "Returns a pain score from 0-25."
        ),
        "input_schema": {
            "type": "object",
            "properties": {
                "pain_severity": {
                    "type": "string",
                    "enum": ["critical", "high", "medium", "low", "unclear"],
                    "description": "How severe and urgent the pain point is"
                },
                "pain_clarity": {
                    "type": "string",
                    "enum": ["very_clear", "somewhat_clear", "vague"],
                    "description": "How well the lead articulated their problem"
                },
                "pain_notes": {
                    "type": "string",
                    "description": "Summary of the pain point evaluation"
                }
            },
            "required": ["pain_severity", "pain_clarity", "pain_notes"]
        }
    }
]


def execute_tool(tool_name: str, tool_input: dict) -> dict:
    """
    Routes tool calls to their handler functions.
    In a real system, these handlers might query a CRM or enrichment API.
    """
    if tool_name == "analyze_firmographic_fit":
        return handle_firmographic_fit(tool_input)
    elif tool_name == "assess_budget_and_timeline":
        return handle_budget_timeline(tool_input)
    elif tool_name == "evaluate_pain_point":
        return handle_pain_point(tool_input)
    else:
        return {"error": f"Unknown tool: {tool_name}"}


def handle_firmographic_fit(data: dict) -> dict:
    """
    Scores company fit based on employee count and revenue signals.
    Max score: 40 points.
    """
    score = 0
    employees = data.get("employee_count", 0) or 0
    revenue = data.get("annual_revenue", "") or ""

    # Score based on company size sweet spot (10-500 employees = ideal SMB target)
    if 10 <= employees <= 500:
        score += 20
    elif employees > 500:
        score += 10  # Enterprise — longer sales cycle, partial credit
    elif employees > 0:
        score += 5

    # Revenue signal scoring
    revenue_lower = revenue.lower()
    if any(x in revenue_lower for x in ["million", "m", "$1m", "$5m", "$10m"]):
        score += 20
    elif any(x in revenue_lower for x in ["hundred", "500k", "750k"]):
        score += 10
    elif revenue:
        score += 5

    return {
        "tool": "analyze_firmographic_fit",
        "score": min(score, 40),
        "notes": data.get("fit_notes", "")
    }


def handle_budget_timeline(data: dict) -> dict:
    """
    Converts budget and timeline signals into a readiness score.
    Max score: 35 points.
    """
    score = 0

    budget_scores = {"strong": 20, "moderate": 13, "weak": 5, "unknown": 0}
    timeline_scores = {"immediate": 15, "within_quarter": 10, "within_year": 5, "unknown": 0}

    score += budget_scores.get(data.get("budget_signal", "unknown"), 0)
    score += timeline_scores.get(data.get("timeline_signal", "unknown"), 0)

    return {
        "tool": "assess_budget_and_timeline",
        "score": score,
        "notes": data.get("readiness_notes", "")
    }


def handle_pain_point(data: dict) -> dict:
    """
    Translates pain severity and clarity into a pain score.
    Max score: 25 points.
    """
    severity_scores = {"critical": 15, "high": 12, "medium": 8, "low": 4, "unclear": 0}
    clarity_scores = {"very_clear": 10, "somewhat_clear": 6, "vague": 2}

    score = severity_scores.get(data.get("pain_severity", "unclear"), 0)
    score += clarity_scores.get(data.get("pain_clarity", "vague"), 0)

    return {
        "tool": "evaluate_pain_point",
        "score": min(score, 25),
        "notes": data.get("pain_notes", "")
    }

The scoring logic is intentional here. Firmographic fit tops out at 40 because it's the biggest predictor of deal success. Budget and timeline cap at 35, and pain point clarity caps at 25 — that gives you a 100-point scale total.

💡 Pro tip: In a production AI lead qualification automation setup, handle_firmographic_fit would make a real API call to something like Clearbit or Apollo to enrich the lead data automatically. For now, the scoring logic still works with whatever data you already have.

Step 3: Build the Primary Qualification Agent

The primary agent is the one that reads the lead data and decides which tools to call. It runs a conversation with Claude, handles tool calls in a loop, and collects all the scoring results before passing them along.

agents.py
import anthropic
import json
from typing import Any
from models import LeadInput
from tools import QUALIFICATION_TOOLS, execute_tool


class QualificationAgent:
    """
    Primary agent that analyzes a lead and calls scoring tools.
    Uses Claude's tool_use feature to structure its analysis.
    """

    def __init__(self, client: anthropic.Anthropic):
        self.client = client
        self.model = "claude-sonnet-4-6"
        self.tool_results: list[dict] = []

    def build_system_prompt(self) -> str:
        return """You are a B2B sales qualification specialist. Your job is to analyze 
incoming leads and use the available tools to score them across three dimensions: 
firmographic fit, budget and timeline readiness, and pain point severity.

You MUST call all three tools for every lead. Use the lead information provided 
to make informed judgments when calling each tool. Be honest — an unqualified 
lead is still valuable data."""

    def build_lead_message(self, lead: LeadInput) -> str:
        """Formats lead data into a clear message for the agent."""
        return f"""Please qualify this lead using all three scoring tools:

Company: {lead.company_name}
Contact: {lead.contact_name}
Email: {lead.email}
Industry: {lead.industry or 'Not provided'}
Employees: {lead.employee_count or 'Not provided'}
Annual Revenue: {lead.annual_revenue or 'Not provided'}
Pain Point: {lead.pain_point or 'Not provided'}
Budget Mentioned: {lead.budget_mentioned or 'Not provided'}
Timeline: {lead.timeline or 'Not provided'}
Website: {lead.website or 'Not provided'}

Use all three tools to evaluate this lead thoroughly."""

    def run(self, lead: LeadInput) -> list[dict]:
        """
        Runs the agentic loop until Claude stops calling tools.
        Returns a list of all tool results collected during the loop.
        """
        messages = [
            {"role": "user", "content": self.build_lead_message(lead)}
        ]

        self.tool_results = []

        while True:
            response = self.client.messages.create(
                model=self.model,
                max_tokens=4096,
                system=self.build_system_prompt(),
                tools=QUALIFICATION_TOOLS,
                messages=messages
            )

            # Check if Claude wants to call tools
            if response.stop_reason == "tool_use":
                # Process each tool call in this response
                tool_use_blocks = [
                    block for block in response.content
                    if block.type == "tool_use"
                ]

                # Build the assistant message with the full content
                messages.append({
                    "role": "assistant",
                    "content": response.content
                })

                # Execute each tool and collect results
                tool_result_contents = []
                for tool_block in tool_use_blocks:
                    result = execute_tool(tool_block.name, tool_block.input)
                    self.tool_results.append(result)

                    tool_result_contents.append({
                        "type": "tool_result",
                        "tool_use_id": tool_block.id,
                        "content": json.dumps(result)
                    })

                # Feed results back to Claude so it can continue
                messages.append({
                    "role": "user",
                    "content": tool_result_contents
                })

            elif response.stop_reason == "end_turn":
                # Claude is done calling tools — exit the loop
                break
            else:
                # Unexpected stop reason — exit to avoid infinite loop
                break

        return self.tool_results

The while True loop is the heart of the agentic pattern. Claude runs, calls a tool, gets the result, decides whether to call another tool, and keeps going until it signals it's done with end_turn. That's the Claude API multi-agent loop in its simplest form.

Step 4: Build the Secondary Scoring Agent

The scoring agent receives all the tool results from the primary agent and synthesizes them into a final verdict. This is context passing in action — the second agent never touches the raw lead, it only sees what the first agent decided.

This separation matters. In a real AI lead qualification automation workflow, your scoring logic and your conversational logic should be independently adjustable. Keeping them as separate agents makes that easy.

agents.py (continued — add to the same file)
class ScoringAgent:
    """
    Secondary agent that synthesizes tool results into a final qualification verdict.
    Receives context from the QualificationAgent — never sees raw lead data directly.
    """

    def __init__(self, client: anthropic.Anthropic):
        self.client = client
        self.model = "claude-sonnet-4-6"

    def build_system_prompt(self) -> str:
        return """You are a senior sales strategist reviewing lead qualification scores.
        
Given the scoring results from our qualification tools, you will:
1. Calculate the total score (firmographic + budget/timeline + pain point)
2. Assign a qualification status: "qualified" (70+), "needs_nurture" (40-69), or "unqualified" (below 40)
3. Recommend a specific next action
4. List 2-3 concrete next steps for the sales team

Respond ONLY with a valid JSON object matching this exact structure:
{
  "qualification_status": "qualified|needs_nurture|unqualified",
  "score": ,
  "score_breakdown": {
    "firmographic": ,
    "budget_timeline": ,
    "pain_point": 
  },
  "recommended_action": "",
  "reasoning": "<2-3 sentence explanation>",
  "next_steps": ["", "", ""]
}"""

    def build_context_message(self, tool_results: list[dict], lead: LeadInput) -> str:
        """
        Formats tool results as context for the scoring agent.
        This is the context passing step between agents.
        """
        results_json = json.dumps(tool_results, indent=2)
        return f"""Here are the qualification tool results for {lead.contact_name} at {lead.company_name}:

{results_json}

Please synthesize these results into a final qualification verdict as JSON."""

    def run(self, tool_results: list[dict], lead: LeadInput) -> dict:
        """
        Takes tool results from the primary agent and returns a structured verdict.
        Returns a dict that maps directly to QualificationResult fields.
        """
        response = self.client.messages.create(
            model=self.model,
            max_tokens=1024,
            system=self.build_system_prompt(),
            messages=[
                {
                    "role": "user",
                    "content": self.build_context_message(tool_results, lead)
                }
            ]
        )

        # Extract the text content from the response
        raw_text = response.content[0].text.strip()

        # Strip markdown code fences if Claude wraps the JSON
        if raw_text.startswith("```"):
            raw_text = raw_text.split("```")[1]
            if raw_text.startswith("json"):
                raw_text = raw_text[4:]
            raw_text = raw_text.strip()

        return json.loads(raw_text)

Step 5: Create the Orchestration Loop

This is where everything connects. The orchestrator class holds both agents, runs them in sequence, and assembles the final QualificationResult. The model_dump_json call at the end gives you clean, serializable output you can pipe into a CRM, a webhook, or a database.

main.py
import anthropic
import json
from models import LeadInput, QualificationResult
from agents import QualificationAgent, ScoringAgent


class LeadQualifierOrchestrator:
    """
    Orchestrates the two-agent lead qualification pipeline.
    Routes lead data through the qualification agent, then the scoring agent,
    and assembles a final QualificationResult.
    """

    def __init__(self):
        # Single shared client — both agents use the same connection
        self.client = anthropic.Anthropic()
        self.qualification_agent = QualificationAgent(self.client)
        self.scoring_agent = ScoringAgent(self.client)

    def qualify_lead(self, lead: LeadInput) -> QualificationResult:
        """
        Full pipeline: qualification tools → scoring synthesis → structured result.
        """
        print(f"\n🔍 Qualifying lead: {lead.contact_name} at {lead.company_name}")
        print("─" * 50)

        # Step 1: Run primary agent — calls all three qualification tools
        print("⚙️  Running qualification agent...")
        tool_results = self.qualification_agent.run(lead)
        print(f"✅ Collected {len(tool_results)} tool results")

        for result in tool_results:
            tool_name = result.get("tool", "unknown")
            score = result.get("score", "N/A")
            print(f"   • {tool_name}: score={score}")

        # Step 2: Pass tool results to scoring agent for synthesis
        print("\n⚙️  Running scoring agent...")
        verdict = self.scoring_agent.run(tool_results, lead)
        print(f"✅ Verdict: {verdict.get('qualification_status', 'unknown').upper()} (score: {verdict.get('score', 0)}/100)")

        # Step 3: Assemble the final structured result
        result = QualificationResult(
            lead_name=lead.contact_name,
            company=lead.company_name,
            qualification_status=verdict["qualification_status"],
            score=verdict["score"],
            score_breakdown=verdict["score_breakdown"],
            recommended_action=verdict["recommended_action"],
            reasoning=verdict["reasoning"],
            next_steps=verdict["next_steps"]
        )

        return result

    def qualify_batch(self, leads: list[LeadInput]) -> list[QualificationResult]:
        """Runs the pipeline over a list of leads and returns all results."""
        results = []
        for lead in leads:
            result = self.qualify_lead(lead)
            results.append(result)
        return results


def main():
    # Sample leads representing realistic B2B scenarios
    sample_leads = [
        LeadInput(
            company_name="Gulf Coast Orthopedics",
            contact_name="Dr