MishaBook a demo

Aug 14, 2026

Run a Two-Week Read-Only AI Pilot

A read-only AI pilot is a time-boxed test (typically 10-14 days) where an AI system observes and scores incoming opportunities, customer data, or operational decisions in parallel with human decision-makers, without executing transactions or changing operational state. Performance is measured against a pre-defined rubric before any live authority is granted.

Why Read-Only Pilots Reduce Risk

A read-only pilot isolates the AI's judgment from operational consequences. The human buyer remains the sole decision-maker and executor. This setup answers a single question: does the AI see what the human sees, and does it score opportunities the same way?

Two weeks is the minimum viable window. Shorter tests (3-5 days) don't capture enough variance in incoming deals or customer types. Longer tests (30+ days) delay the go/no-go decision and increase sunk cost bias. Fourteen days captures roughly 2-3 full business cycles for most DTC brands.

Pre-Pilot Setup: Define Your Scoring Rubric

Before the AI sees any data, document how your best buyer makes decisions. This becomes the ground truth. The rubric must be specific enough that two humans could score the same opportunity identically 80% of the time.

Start with your top 5-7 decision factors. For a DTC brand, this might be: customer LTV projection, repeat purchase likelihood, CAC payback period, brand fit, and support burden. Assign weights (must sum to 100). Define thresholds for each factor - not ranges, thresholds.

  • LTV projection: >$500 = 25 points, $200-$500 = 15 points, <$200 = 0 points
  • Repeat likelihood (based on product category + customer segment): >60% = 20 points, 30-60% = 10 points, <30% = 0 points
  • CAC payback: <6 months = 20 points, 6-12 months = 10 points, >12 months = 0 points
  • Brand alignment: Strong fit = 20 points, Neutral = 10 points, Poor fit = 0 points
  • Support complexity: Low = 15 points, Medium = 8 points, High = 0 points

Data Preparation and AI Onboarding

The AI needs access to the same inputs your buyer uses. This typically includes: customer inquiry text, product interest signals, historical customer data (if applicable), and any internal notes or context. Do not give the AI access to your decision outcome yet - that comes after scoring.

Provide 20-30 historical examples (from the past 2-3 months) where you know the outcome. The AI uses these to calibrate its understanding of your scoring rubric. Then it scores new incoming opportunities blind, without seeing your decision until after submission.

Parallel Scoring: The Two-Week Window

During the pilot, every new opportunity gets scored by both the human buyer and the AI, independently and in parallel. The human does not see the AI score before making their decision. The AI does not see the human decision before submitting its score.

Capture the following for each opportunity: AI score, human score, AI reasoning (brief), human reasoning (brief), final decision (human's call), and outcome (if known within the window). Aim for 15-25 scored opportunities over two weeks. This is enough to detect systematic misalignment.

Scoring Agreement Metrics

After two weeks, calculate three metrics: exact agreement rate, directional agreement rate, and scoring variance.

Exact agreement: percentage of opportunities where AI and human scores are within 5 points of each other (on a 0-100 scale). Target: >70%. Below 60% signals the AI has not learned your rubric.

  • Directional agreement: percentage of opportunities where AI and human both recommend accept or both recommend reject (regardless of score magnitude). Target: >80%. Directional misalignment is the bigger risk.
  • Scoring variance: standard deviation of (AI score - human score). Target: <8 points. High variance suggests the AI is weighting factors differently than intended.
  • False positive rate: percentage of opportunities the AI scored high (>70) that the human rejected. Target: <15%. High false positives mean the AI is too optimistic.
  • False negative rate: percentage of opportunities the AI scored low (<40) that the human accepted. Target: <10%. High false negatives mean the AI is too conservative.

Go/No-Go Decision Rules

Use this decision matrix after the two-week window closes. All three conditions must be met to move to limited live authority.

  • Exact agreement >70% AND directional agreement >80%: Proceed to Phase 2
  • Exact agreement 60-70% AND directional agreement 75-80%: Extend pilot by one week, then re-evaluate
  • Exact agreement <60% OR directional agreement <75%: Stop. Revise rubric or data inputs and restart pilot.
  • False positive rate >20% OR false negative rate >15%: Stop. AI is systematically misaligned on risk tolerance.

Phase 2: Limited Live Authority (If Go)

If metrics pass, move to Phase 2: the AI scores and flags opportunities, but the human still approves all decisions. The AI's recommendation appears alongside the human's reasoning. Track whether the human follows the AI recommendation (adoption rate) and whether AI-recommended accepts perform as predicted.

Run Phase 2 for one week. If adoption rate >60% and performance matches prediction, expand to full authority on a subset of opportunity types (e.g., repeat customers only, or LTV >$300 only). Never grant full authority without Phase 2 validation.

Questions

FAQ

What if we don't have a formal scoring rubric yet?

Build it during the pre-pilot week. Interview your top buyer for 2-3 hours. Ask them to score 10 recent opportunities out loud, explaining their reasoning for each. Identify the 5-7 factors that appear in every decision. Assign weights based on how often each factor changed the final call. This is your rubric.

Can we run the pilot on historical data instead of live opportunities?

No. Historical scoring is useful for calibration (step 3), but the pilot must use live, unseen opportunities. Historical data introduces hindsight bias - the human knows the outcome, which changes how they score. Live data forces both the human and AI to make predictions under uncertainty, which is the actual use case.

What if we only get 5-8 opportunities in two weeks?

Extend the pilot to three weeks. Five opportunities is too small a sample to detect systematic misalignment. Aim for at least 15. If your business naturally has low volume, use Phase 1 (calibration) to score 20-30 historical opportunities instead, then run Phase 2 (live scoring) as your actual pilot.

Should the human buyer know they're being scored against the AI?

Yes. Transparency prevents resentment and ensures the human is scoring normally (not defensively). Frame it as: 'We're testing whether this tool can learn how you make decisions. Your score is the ground truth.' This also improves buy-in if the pilot succeeds.

Want this on your account?

Thirty minutes. Bring the number that keeps you up.

More from the blog