---
# === IDENTITY ===
id: consulting/recipes/pilot-execution-playbook/2026
canonical_question: "How do you execute a Signal Stack pilot delivering 10-20 qualified dossiers per week?"
aliases:
  - "Signal Stack pilot execution steps"
  - "How to run a signal-based outreach pilot"
  - "Signal detection pilot delivery process"
entity_type: execution_recipe
domain: consulting > recipes > Pilot Execution Playbook
region: global
jurisdiction: global
temporal_scope: 2026-2027

# === VERIFICATION ===
last_verified: 2026-03-29
confidence: 0.85
version: 1.0
first_published: 2026-03-29

# === TEMPORAL VALIDITY ===
temporal_validity:
  status: evolving
  last_breaking_change: "Initial release — Signal Stack CaaS pilot methodology v1.0"
  next_review: 2026-09-25
  change_sensitivity: high

# === CONSTRAINTS ===
constraints:
  - "No platform work until 3 paying customers validate vertical #1 — pilot proves the vertical, not the platform"
  - "Human-in-the-loop review mandatory for first 100 packages per client — quality > volume"
  - "Pilot customers must have measurable current outreach baselines (open rate, reply rate, meeting rate, close rate)"
  - "GDPR/PECR/CAN-SPAM compliance required for all outbound delivery — verify jurisdiction before first send"
  - "Minimum pilot duration: 4 weeks — shorter pilots produce statistically unreliable conversion data"

# === SKIP CONDITIONS ===
skip_this_unit_if:
  - condition: "User needs Signal Stack theory, not execution"
    use_instead: "consulting/signal-stack/five-layer-pipeline-architecture/2026"
  - condition: "User needs to extract platform from working vertical"
    use_instead: "consulting/recipes/platform-extraction/2026"
  - condition: "User needs to launch a new vertical, not run first pilot"
    use_instead: "consulting/recipes/vertical-launch-checklist/2026"

# === AGENT HINTS ===
inputs_needed:
  - key: vertical
    question: "Which industry vertical is this pilot targeting?"
    type: text
  - key: pilot_customers
    question: "How many pilot customers are lined up?"
    type: choice
    options: ["1", "2-3", "4-5", "not yet identified"]
  - key: current_outreach
    question: "Does the client have measurable outreach baselines (open/reply/meeting rates)?"
    type: choice
    options: ["yes — tracked in CRM", "partially — some metrics available", "no — no baseline exists"]
  - key: signal_sources
    question: "Have signal sources been audited for this vertical?"
    type: choice
    options: ["yes — signal audit complete", "partially — some sources identified", "no — audit not started"]

# === EXECUTION METADATA ===
execution:
  required_inputs:
    - name: "Completed signal audit"
      source: "Signal Architect"
      format: "document"
    - name: "Signal taxonomy with trigger definitions"
      source: "Taxonomy workshop output"
      format: "document"
    - name: "2-3 pilot customer agreements"
      source: "Sales/BD"
      format: "signed agreements"
    - name: "Current outreach baseline metrics"
      source: "Pilot customers"
      format: "spreadsheet"

  outputs:
    - name: "Weekly Dossier Batch"
      format: "PDF dossiers + tracking spreadsheet"
      description: "10-20 qualified dossiers per week with signal evidence, enrichment data, and tailored outreach copy"
    - name: "Pilot Performance Dashboard"
      format: "spreadsheet + weekly report"
      description: "Signal accuracy, dossier quality scores, conversion funnel (open/reply/meeting/close)"
    - name: "Signal Taxonomy Iteration Log"
      format: "document"
      description: "Weekly false positive analysis with taxonomy adjustments"
    - name: "Phase 1 Exit Assessment"
      format: "document"
      description: "Go/no-go decision for platform extraction based on 3 paying customers"

  tools_required:
    - name: "LLM API (Claude or GPT-4)"
      purpose: "Signal classification and dossier generation"
      tier: "paid"
      cost: "$200-500/month per client"
      alternatives: ["Open-source models (Llama 3)", "Rule-based classification only"]
    - name: "Enrichment API (Clearbit/Apollo/LinkedIn)"
      purpose: "Firmographic enrichment and decision-maker identification"
      tier: "paid"
      cost: "$100-300/month"
      alternatives: ["Manual LinkedIn research", "Hunter.io"]
    - name: "Email delivery (Resend/SendGrid)"
      purpose: "Dossier delivery with open/click tracking"
      tier: "free/paid"
      cost: "$0-50/month"
      alternatives: ["Manual email", "Mailgun"]
    - name: "Data source scrapers/APIs"
      purpose: "Pulling signal data from public sources"
      tier: "varies"
      cost: "$0-200/month depending on sources"
      alternatives: ["Manual monitoring", "RSS feeds"]

  credentials_needed:
    - service: "LLM Provider"
      type: "API key"
      where_to_get: "https://console.anthropic.com or https://platform.openai.com"
      free_tier_limits: "Varies — budget $200-500/month"
    - service: "Enrichment API"
      type: "API key"
      where_to_get: "https://clearbit.com or https://apollo.io"
      free_tier_limits: "Apollo: 50 credits/month free; Clearbit: paid only"

  estimated_duration: "4-8 weeks (pilot execution)"
  estimated_cost: "$5K-15K (tools + labor) per pilot cycle"

# === DISTRIBUTION ===
canonical_source: "https://knowledgelib.io/consulting/recipes/pilot-execution-playbook/2026"
suggested_citation: "Source: knowledgelib.io — AI Knowledge Library (verified 2026-03-29)"

# === RELATED UNITS ===
related_kos:
  depends_on:
    - id: "consulting/signal-stack/five-layer-pipeline-architecture/2026"
      label: "The five-layer signal pipeline — Ingest, Detect, Enrich, Generate, Deliver"
  feeds_into:
    - id: "consulting/recipes/platform-extraction/2026"
      label: "Platform extraction triggered by Phase 1 exit criteria"
  related_to:
    - id: "consulting/recipes/vertical-launch-checklist/2026"
      label: "Checklist for launching subsequent verticals after pilot success"

# === SOURCES ===
sources:
  - id: src1
    title: "Signal Stack — A Unified Signal-family Platform"
    author: "Internal methodology"
    url: https://knowledgelib.io/consulting/signal-stack/signal-stack-architecture/2026
    type: primary_research
    published: 2026-03-09
    reliability: high
  - id: src2
    title: "Signal Stack Consulting-as-a-Service — Strategic Intelligence Report"
    author: "Internal methodology"
    url: https://knowledgelib.io/consulting/signal-stack/signal-stack-caas/2026
    type: primary_research
    published: 2026-03-29
    reliability: high
  - id: src3
    title: "Stop Cold Emailing and Start Reading the Exhaust Fumes"
    author: "Internal research"
    url: https://knowledgelib.io/consulting/signal-stack/exhaust-fumes-thesis/2026
    type: primary_research
    published: 2026-03-01
    reliability: high
  - id: src4
    title: "B2B Revenue Waterfall Report"
    author: "Forrester Research"
    url: https://www.forrester.com/report/the-b2b-revenue-waterfall
    type: industry_report
    published: 2024-01-01
    reliability: authoritative
  - id: src5
    title: "How Brands Grow"
    author: "Sharp, B."
    url: https://marketingscience.info/how-brands-grow/
    type: academic_paper
    published: 2010-01-01
    reliability: authoritative
---

# Pilot Execution Playbook

## Purpose

This recipe executes a Signal Stack pilot that delivers 10-20 qualified dossiers per week to 2-3 pilot customers. The pilot validates that signal-driven outreach produces measurably higher conversion than cold outreach — the core thesis that 95% of any market is not buying at any moment, but the 5% in active crisis are detectable through observable corporate distress signals. Phase 1 exit criteria: 3 paying customers in vertical #1. [src1, src3]

## Prerequisites
<!-- Agents: verify ALL prerequisites before executing. Missing prerequisites = failed execution. -->

- [ ] **Signal audit complete** — all available signal sources in the target vertical mapped, scored by accessibility, cost, refresh rate, and signal-to-noise ratio
- [ ] **Signal taxonomy defined** — trigger events, strength scoring, false positive thresholds, and compound signal logic validated by domain advisor
- [ ] **2-3 pilot customers identified** — each with measurable current outreach baselines (open rate, reply rate, meeting-booked rate, close rate)
- [ ] **MVP pipeline operational** — cron jobs pulling data, LLM classification running, enrichment APIs connected, dossier generation working
- [ ] **GDPR/PECR/CAN-SPAM compliance confirmed** — outbound delivery legal in target jurisdictions

## Constraints
<!-- Hard rules. Agents: enforce throughout execution. Violating these = broken output or legal risk. -->

- No platform work until 3 paying customers validate vertical #1. The pilot proves the vertical, not the platform. [src1]
- Human-in-the-loop review mandatory for first 100 packages per client. Quality failures in pilot phase kill the entire engagement. [src2]
- Pilot customers must have measurable current outreach baselines — without baselines, you cannot prove the 2x conversion claim.
- GDPR/PECR/CAN-SPAM compliance required for all outbound delivery. Verify jurisdiction before first send. [src2]
- Minimum pilot duration: 4 weeks. Shorter pilots produce statistically unreliable conversion data. [src5]

## Tool Selection Decision

```
Which path?
├── Technical team available (can write Python)
│   ├── Budget > $500/month
│   │   └── PATH A: Full API stack — LLM + Clearbit/Apollo + Resend
│   └── Budget < $500/month
│       └── PATH B: LLM + free enrichment (Apollo free tier + Hunter.io) + manual delivery
├── No-code team
│   ├── Budget > $500/month
│   │   └── PATH C: Make/n8n + Clay + LLM API — automated but no-code
│   └── Budget < $500/month
│       └── PATH D: Manual pipeline — Google Alerts + ChatGPT + manual enrichment + email
└── Hybrid (one developer + ops person)
    └── PATH E: Python scrapers + LLM API + manual QA — best quality/cost ratio for pilot
```

| Path | Tools | Cost/month | Speed | Output Quality |
|------|-------|------------|-------|---------------|
| A: Full API | Python + Claude/GPT-4 + Clearbit + Resend | $500-800 | 10-20 dossiers/week | Excellent — fully automated |
| B: Budget API | Python + Claude/GPT-4 + Apollo free + manual | $200-400 | 10-15 dossiers/week | Good — some manual steps |
| C: No-code | Make/n8n + Clay + LLM API | $400-700 | 8-15 dossiers/week | Good — limited customization |
| D: Manual | Google Alerts + ChatGPT + manual | $20-50 | 5-10 dossiers/week | Adequate — slow but validates thesis |
| E: Hybrid | Python + LLM API + manual QA | $300-500 | 10-20 dossiers/week | Excellent — best pilot path |

## Execution Flow

### Step 1: Baseline Measurement

**Duration**: 2-3 days
**Tool**: CRM export + spreadsheet analysis

Collect current outreach performance metrics from each pilot customer: total outreach volume per week, open rate, reply rate, meeting-booked rate, close rate, average deal size, and cost per meeting. These baselines are the control group — every pilot metric is measured against them.

Document the baseline in a shared spreadsheet with columns: metric, current value, source, date range, notes.

**Verify**: Baseline spreadsheet complete for all pilot customers with at least 3 months of historical data.
**If failed**: If customer cannot provide baselines, estimate from industry benchmarks (B2B cold outreach: 15-25% open, 1-3% reply, 0.5-1% meeting rate). Document that baselines are estimated, not measured.

### Step 2: Pipeline Calibration

**Duration**: 3-5 days
**Tool**: Python scripts + LLM API + test data

Run the MVP pipeline on a test batch of 50-100 signal events from the target vertical. Classify each as true positive (real buying trigger), false positive (noise), or ambiguous. Calculate initial precision rate.

Calibrate classification thresholds: adjust LLM prompts, add rule-based filters for common false positive patterns, tune compound signal logic (e.g., single signal = low confidence, two correlated signals = high confidence). [src1]

**Verify**: Precision rate > 60% on test batch. If below 60%, the taxonomy needs rework before pilot launch.
**If failed**: Return to taxonomy workshop. Identify the top 3 false positive patterns and add explicit exclusion rules. Re-run test batch.

### Step 3: First Dossier Batch (Week 1)

**Duration**: 5 days (ongoing weekly)
**Tool**: Full pipeline + human review

Generate the first batch of 10-20 dossiers. Each dossier contains: signal evidence (what triggered the alert), company profile (enrichment data), decision-maker identification, tailored outreach copy, and a proof pack (the specific data points supporting the signal). [src2]

Human-in-the-loop review: Signal Architect reviews every dossier in batch 1. Score each on a 1-5 scale for signal accuracy, enrichment completeness, outreach relevance, and proof pack quality. Flag and discard any dossier scoring below 3. [src1]

**Verify**: At least 10 dossiers pass quality review (score >= 3 on all dimensions). Delivery to pilot customers confirmed.
**If failed**: If fewer than 10 pass review, the pipeline is producing too many false positives. Pause delivery, tighten classification, generate a new batch.

### Step 4: A/B Test Package Formats

**Duration**: 2 weeks (runs parallel to weekly delivery)
**Tool**: Email delivery platform with variant tracking

Split dossier delivery into 2-3 format variants to identify what converts best:
- **Variant A**: Full dossier PDF attachment — comprehensive signal evidence + proof pack + recommended action
- **Variant B**: Executive summary email — key signal + one-line proof + CTA to schedule call for full briefing
- **Variant C**: Data-only alert — signal detected, enrichment data, no outreach copy (customer writes their own)

Track open rate, reply rate, and meeting-booked rate per variant per pilot customer. [src3]

**Verify**: Statistical significance on at least one metric after 2 weeks (minimum 30 deliveries per variant).
**If failed**: If volume is too low for significance, extend A/B test by 2 weeks or reduce to 2 variants.

### Step 5: Weekly Taxonomy Iteration

**Duration**: 2-4 hours per week (ongoing)
**Tool**: Spreadsheet analysis + taxonomy update

Every week, review all signals from the past 7 days. Classify outcomes: signal led to meeting (true positive), signal was ignored (unknown), signal was rejected by customer as irrelevant (false positive). Calculate weekly precision and recall.

Update taxonomy based on false positive analysis: add new exclusion rules, adjust signal strength scoring, modify compound signal logic. Document every change in the iteration log with rationale. [src1, src2]

**Verify**: False positive rate decreasing week-over-week. Target: < 30% false positive rate by week 4.
**If failed**: If false positive rate is not improving, the signal sources may be fundamentally noisy. Consider replacing the weakest signal source or adding a second corroborating signal requirement.

### Step 6: Conversion Tracking

**Duration**: Ongoing (weekly measurement)
**Tool**: CRM integration or manual tracking spreadsheet

Track the full conversion funnel for every dossier delivered: dossier sent, opened, replied, meeting booked, proposal sent, deal closed, deal value. Compare against baseline metrics from Step 1.

Key metric: are pilot customers converting signal-based dossiers at > 2x their cold outreach baseline? [src3, src5]

**Verify**: Conversion data available for at least 80% of delivered dossiers by week 4.
**If failed**: If tracking is incomplete, add manual follow-up calls with pilot customers to capture outcomes. Automate tracking in next iteration.

### Step 7: Phase 1 Exit Assessment

**Duration**: 1-2 days
**Tool**: Analysis + presentation

After 4-8 weeks of pilot operation, compile the Phase 1 exit assessment:
- Signal accuracy: precision rate, false positive trend, taxonomy iterations completed
- Conversion performance: dossier-to-meeting rate vs. cold outreach baseline
- Customer satisfaction: NPS or qualitative feedback from pilot customers
- Unit economics: cost per qualified dossier, cost per meeting booked
- Go/no-go recommendation for platform extraction

Phase 1 exit criteria: 3 paying customers in vertical #1 with measurable conversion improvement over baseline. [src1, src2]

**Verify**: Exit assessment document complete with data-backed recommendation.
**If failed**: If fewer than 3 customers are paying, extend pilot by 4 weeks with adjusted taxonomy. If conversion is not beating baseline after 8 weeks, reassess signal sources and vertical selection.

## Output Schema

```json
{
  "output_type": "pilot_performance_report",
  "format": "spreadsheet + PDF summary",
  "sections": [
    {"name": "baseline_metrics", "type": "object", "description": "Pre-pilot outreach performance per customer", "required": true},
    {"name": "weekly_dossier_batches", "type": "array", "description": "Dossier count, quality scores, delivery confirmations per week", "required": true},
    {"name": "signal_accuracy", "type": "object", "description": "Precision rate, false positive rate, taxonomy iteration count", "required": true},
    {"name": "conversion_funnel", "type": "object", "description": "Open/reply/meeting/close rates vs baseline", "required": true},
    {"name": "ab_test_results", "type": "object", "description": "Package format variant performance", "required": true},
    {"name": "exit_assessment", "type": "object", "description": "Go/no-go for platform extraction with supporting data", "required": true}
  ],
  "expected_sections": "6",
  "sort_order": "chronological by pilot week"
}
```

## Quality Benchmarks

| Quality Metric | Minimum Acceptable | Good | Excellent |
|---------------|-------------------|------|-----------|
| Dossier volume (per week) | >= 10 | >= 15 | >= 20 |
| Signal precision rate | > 60% | > 75% | > 85% |
| Dossier quality score (avg) | > 3.0/5 | > 3.5/5 | > 4.0/5 |
| Conversion vs. baseline | > 1.5x | > 2x | > 3x |
| False positive trend | Flat | Decreasing | < 20% by week 4 |
| Pilot customer retention | 2/3 continue | 3/3 continue | 3/3 convert to paid |

**If below minimum**: Pause delivery, return to taxonomy calibration, and extend pilot by 2 weeks.

## Error Handling

| Error | Likely Cause | Recovery Action |
|-------|-------------|----------------|
| Precision < 60% after calibration | Signal taxonomy too broad or sources too noisy | Narrow to top 2-3 highest-quality signal sources, add compound signal requirement |
| Pilot customer stops responding | Dossier quality too low or wrong decision-maker targeted | Direct outreach to customer, request feedback, pivot to different contact |
| Dossier volume < 10/week | Signal sources have low event frequency in target vertical | Add more signal sources or broaden geographic scope |
| A/B test inconclusive | Volume too low for statistical significance | Extend test period or reduce to 2 variants |
| Compliance flag on outbound delivery | PECR/CAN-SPAM violation in delivery method | Immediately pause delivery, review compliance, switch to opt-in only delivery |
| LLM classification quality degrades | Prompt drift or model update | Re-calibrate prompts, pin model version, add regression test batch |

## Cost Breakdown

| Component | Budget Pilot ($5K) | Standard Pilot ($10K) | Premium Pilot ($15K) |
|-----------|--------------------|-----------------------|---------------------|
| Signal Architect labor | $2K (part-time) | $4K (dedicated) | $6K (dedicated + advisor) |
| LLM API costs | $500 | $1K | $1.5K |
| Enrichment APIs | $300 | $600 | $1K |
| Delivery infrastructure | $200 | $400 | $500 |
| Domain advisor | $0 (DIY) | $2K | $3K |
| QA and iteration | $1K | $2K | $3K |
| **Total (4-week pilot)** | **$4K-$5K** | **$10K** | **$15K** |

## Anti-Patterns

### Wrong: Optimizing for volume over quality in week 1
Pushing 20+ dossiers per week before signal accuracy is validated. Result: pilot customers receive irrelevant dossiers, lose trust in the methodology, and the pilot fails despite having a working pipeline. [src2]

### Correct: Cap at 10 dossiers week 1, scale only after quality is validated
Deliver 10 human-reviewed dossiers in week 1. Only increase volume after pilot customer confirms signal relevance on >= 7/10 dossiers.

### Wrong: Skipping baseline measurement
Starting the pilot without documenting current outreach performance. Result: you cannot prove the 2x conversion claim because there is no control group. The pilot produces anecdotal success stories instead of data-backed evidence. [src3]

### Correct: Measure before you move
Spend 2-3 days collecting 3+ months of historical outreach data from each pilot customer before delivering a single dossier.

### Wrong: Not iterating on the taxonomy weekly
Treating the initial signal taxonomy as fixed for the entire pilot. Result: false positive rate stays high, dossier quality plateaus, and pilot customers churn. [src1]

### Correct: Weekly taxonomy iteration is non-negotiable
Review every false positive weekly. Update classification rules, adjust scoring thresholds, and document changes. The taxonomy should visibly improve each week.

## When This Matters

Use when an agent needs to execute or plan a Signal Stack pilot engagement. This is the hands-on delivery recipe — it takes a completed signal audit and taxonomy and turns them into measurable results. The pilot is the validation gate: if it does not produce 3 paying customers with > 2x conversion, do not proceed to platform extraction.

## Related Units

- [Signal Stack Architecture](/consulting/signal-stack/signal-stack-architecture/2026)
- [Platform Extraction](/consulting/recipes/platform-extraction/2026)
- [Vertical Launch Checklist](/consulting/recipes/vertical-launch-checklist/2026)
