All posts
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
13
min read

How to Run a 30-Day AI Support Agent Pilot in B2B SaaS (2026)

Most AI support agent evaluations stall in the same place. The demo goes well, procurement asks for a business case, and nobody can produce one because the numbers came from the vendor's slides rather than from your queue. A 30-day AI support agent pilot fixes that by putting the agent on a slice of your real traffic, with a baseline recorded before it starts and a decision rule agreed before anyone sees results. This guide is the follow-on to our RFP checklist for evaluating an AI support agent: the checklist narrows the field to one or two vendors, and the pilot tells you whether the one you picked earns a wider rollout. The argument is simple: a pilot without a baseline, a scope and a go/no-go rule is a second demo, and you already had one of those.

What is an AI support agent pilot?

An AI support agent pilot is a time-boxed deployment, usually 30 days, in which an AI agent handles a defined slice of live customer support traffic while your team measures its resolution rate, accuracy, escalation behavior and effect on human workload against a baseline recorded before the pilot began. It differs from a proof of concept in one important way: a proof of concept shows the agent can work on sample tickets, while a pilot shows what it does to your queue, your CSAT and your agents' day when real customers are on the other end.

The reason to run one is not curiosity. AI support agents in 2026 make claims about autonomous resolution that are impossible to check from a sales call, because the answer depends on your knowledge base, your product's complexity and how your customers actually write in. The only evidence that transfers to your environment is evidence generated in your environment. A well-run pilot produces that evidence in a month; a poorly run one produces anecdotes and a renewal argument for the vendor.

What should a 30-day pilot include, and what should it leave out?

Scope the pilot to one or two intents that are high volume, low risk and answerable from documentation you already trust, on one channel your team already monitors closely. Leave out billing disputes, security questions, anything regulated, and any intent where a wrong answer costs more than a slow one. The goal is a slice big enough to be statistically meaningful and small enough that a failure is an inconvenience, not an incident.

A practical way to pick the slice: pull the last 90 days of tickets, group them by intent, and rank by volume. Then strike anything that touches money, access, compliance or a bug. What is left is typically how-to questions, configuration and setup help, "where do I find" questions and integration setup. Two or three of those intents will usually account for 20 to 35 percent of volume in a B2B SaaS queue, which is plenty for a pilot.

On channel, pick the one where you can watch every conversation without heroics. For many teams that is the in-product widget or a single shared Slack Connect channel rather than email, because turnaround is visible and a human can step in mid-conversation. Multi-channel rollout is a rollout question, not a pilot question.

Write the scope down as a one-page document with three sections: intents in, intents out, and the rule for what the agent does when a conversation drifts outside scope. Every disagreement in week three traces back to something that was not written down in week zero.

What baseline should you record before day one?

Before the agent handles a single conversation, record four numbers for the in-scope intents over the previous 30 to 90 days: ticket volume, median time to first response, median time to resolution, and CSAT or the equivalent satisfaction signal for those intents. Add a fifth if you can get it: human handle time per ticket for the same intents. Without these, the pilot has nothing to be compared against and the results will be argued about rather than read.

Two details matter here. First, the baseline must be filtered to the in-scope intents, not the whole queue. A pilot that handles setup questions should not be judged against a queue average dominated by bug reports. Second, note the seasonality. If the pilot runs across a release, a holiday or a quarter close, record that, because volume and sentiment will move for reasons unrelated to the agent.

Also record the state of your knowledge before the pilot. Which help center articles cover the in-scope intents, when they were last updated, and where the known gaps are. The agent will only be as good as this material on day one, and a chunk of the pilot's value is finding out which articles are wrong.

How should the 30 days be structured?

Run the pilot in four phases: setup and shadow mode in week one, supervised live traffic in week two, full live traffic on the scoped intents in weeks three and four, and a decision review in the final two days. Each phase has an exit condition, and the pilot does not advance until it is met.

Week 1: connect, configure, shadow

Connect the agent to your knowledge sources and to the systems it needs to read from, typically the help center, the ticketing system and whatever holds account context. Configure the in-scope intents and the escalation rule in plain language. Then run the agent in shadow mode: it drafts answers to live conversations, humans see the draft, and a human still sends every reply.

Shadow mode is where you learn the real accuracy number. Have two people on the team grade every draft on a three-point scale: would send as is, would send with edits, would not send. Track the ratio daily. Exit condition for week one: at least 70 percent of drafts on in-scope intents graded "would send as is" for three consecutive days, and zero drafts that invented a feature, a price or a policy.

Time to live matters at this stage and it is a legitimate pilot metric on its own. If connecting knowledge and configuring intents consumes most of the week, that is information about what a wider rollout will cost. An agent your support team can configure in a day, without an engineering ticket, changes the economics of every later phase.

Week 2: supervised live

Let the agent reply directly on the in-scope intents, but route every conversation into a review queue that a human reads within the hour. The human can correct, follow up or take over. Track the same three-point grade, plus the number of conversations where a human had to intervene and why.

This is the week your knowledge gaps surface. Expect a cluster of failures that trace to an outdated article or an undocumented edge case rather than to the agent. Fix the article, not the agent, and log each fix. Exit condition: human intervention on fewer than one in five in-scope conversations, and no incident where the agent said something you had to apologize for.

Weeks 3 and 4: full live on the scoped intents

Remove the review queue for in-scope conversations and let the agent run. Humans still see everything that escalates. This is the only two-week window that produces numbers comparable to your baseline, so protect it: do not add intents, do not change the escalation rule, do not switch channels. If you change the scope in week three, you have started a new pilot.

During these two weeks, the one thing to watch actively is escalation quality. When the agent hands off, does the human get the full context, or does the customer repeat themselves? A handoff that loses context turns every escalation into a worse experience than having no agent at all.

Final two days: decision review

Pull the four baseline metrics for the same intents over weeks three and four and put them next to the pre-pilot numbers. Add the pilot-only metrics: resolution rate, accuracy grade, escalation rate and time to live. Then apply the decision rule you wrote before the pilot started.

Which metrics decide whether the pilot worked?

Four metrics decide it: resolution rate on in-scope intents, accuracy as graded by your own team, escalation rate and escalation quality, and change in human handle time. Every other number is context. Volume deflected is not on the list, because deflection counts conversations the agent ended, not problems it solved, and the two are not the same thing.

Resolution rate

The share of in-scope conversations the agent closed without a human touching them and without the customer reopening or re-contacting within 72 hours. The re-contact window is the part most teams skip and the part that separates a real resolution from a polite brush-off. Set the target before the pilot; for narrowly scoped how-to intents with decent documentation, a majority of conversations fully resolved is a reasonable bar, and anything below a third suggests the scope or the knowledge is wrong.

Accuracy

Your team's grade, not the vendor's dashboard. Sample at least 50 conversations per week from weeks three and four and grade them on the same three-point scale used in shadow mode. The number that matters most is the count of confidently wrong answers, because a single invented policy does more damage than twenty "would send with edits" drafts.

Escalation rate and quality

Escalation rate is the share of in-scope conversations handed to a human. On its own it is ambiguous: a low rate could mean the agent is capable or that it is answering things it should not. Read it alongside accuracy. Escalation quality is whether the human received the conversation history, the customer's account context and the agent's own assessment of what it could not do. Score it on the same sample of 50.

Human handle time

For the in-scope intents, did median handle time per ticket that still reached a human go down, stay flat or go up? It should go down, because the agent has already gathered context and answered the easy part. If it went up, the agent is generating clean-up work, and that cost has to be counted against any resolution gains.

How do you make the go/no-go decision?

Write the decision rule before day one, in the form "we expand if resolution on in-scope intents is at or above X, confidently wrong answers are at or below Y per hundred, and handle time on escalated tickets did not rise." Then apply it literally on day 29. The point of writing it in advance is that by day 29 someone on the team will like the tool and someone will not, and neither opinion should be the deciding vote.

Three outcomes are possible. Expand: the rule was met, and the next step is adding intents or a second channel under the same measurement discipline. Extend: the rule was narrowly missed and the misses trace to knowledge gaps you fixed in week two, so run weeks three and four again with the corrected knowledge before deciding. Stop: the rule was missed on accuracy or escalation quality, which are the two failures that do not improve by adding more documentation.

Be honest about one more thing in the review: how much of the team's time the pilot consumed. An agent that resolved half of in-scope conversations but needed a full-time person to babysit it has not yet earned a rollout. Time to live and ongoing maintenance effort belong in the business case alongside resolution rate.

Where does Worknet fit in a pilot like this?

Worknet is built for exactly the pilot shape described here. It runs inside your product, in Slack and in Microsoft Teams, connects to Salesforce, Zendesk, HubSpot and other systems through API and MCP, and takes its configuration in plain English from the support team rather than through an engineering project, so week one is measured in days rather than consumed by integration work. It answers from your company knowledge, reads the user's screen and account state to answer in context, takes permitted actions on the user's behalf, and hands off to a human with the full conversation when it reaches the edge of its scope.

Two things we will say plainly. Worknet's pricing is quote-based, so this post makes no cost claim and no claim about how it compares on price to any other vendor. And the metrics above are the ones we would ask you to hold us to in a pilot, including the 72-hour re-contact window on resolution rate, because a number we cannot defend in your queue is not a number worth quoting.

Conclusion

A 30-day AI support agent pilot works when it has a written scope, a filtered baseline, four phases with exit conditions and a decision rule agreed before the first conversation. Run that way, it produces a business case built from your own tickets rather than a vendor's deck, and it surfaces the knowledge gaps that would have hurt you in any rollout. If you have finished the RFP stage and want to see what a pilot looks like on your own product, book a demo and we will scope the first two intents with you.

FAQs

Frequently Asked Questions

How long should an AI support agent pilot run?

Thirty days is the practical minimum for a B2B SaaS support pilot. It allows a week for setup and shadow mode, a week of supervised live traffic, and two weeks of full live traffic on the scoped intents, which is the only window that produces numbers comparable to a pre-pilot baseline. Shorter pilots tend to end before the knowledge fixes from week two have had a chance to show up in the results.

What is the difference between an AI support agent pilot and a proof of concept?

A proof of concept shows the agent can answer a set of sample tickets, usually chosen by the vendor, in a controlled setting. A pilot puts the agent on live customer traffic for a defined slice of intents and measures resolution rate, accuracy, escalation behavior and human handle time against a baseline. A proof of concept tells you the agent works; a pilot tells you what it does to your queue.

Which metrics should an AI support agent pilot track?

Four metrics decide the outcome: resolution rate on in-scope intents with a 72-hour re-contact window, accuracy as graded by your own team on a sample of at least 50 conversations per week, escalation rate together with escalation quality, and the change in human handle time on tickets that still reach an agent. Volume deflected is context, not a decision metric, because it counts conversations ended rather than problems solved.

Which support intents should be in scope for a first pilot?

Start with one to three intents that are high volume, low risk and answerable from documentation you trust, such as how-to questions, setup and configuration help, and integration setup. Exclude billing disputes, access and security questions, anything regulated, and any intent where a wrong answer costs more than a slow one. In most B2B SaaS queues the chosen intents will cover 20 to 35 percent of volume, which is enough for a meaningful pilot.

When should you stop an AI support agent pilot early?

Stop early if the agent produces a confidently wrong answer about a policy, price or feature in live traffic and the vendor cannot show why it will not recur, or if escalations are reaching humans without conversation context so customers have to repeat themselves. Both failures are about trust and handoff design rather than knowledge coverage, and neither improves by adding more documentation. Low resolution rates on their own are usually a scope or knowledge problem and a reason to adjust, not to stop.

Question text goes here

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique. Duis cursus, mi quis viverra ornare, eros dolor interdum nulla, ut commodo diam libero vitae erat. Aenean faucibus nibh et justo cursus id rutrum lorem imperdiet. Nunc ut sem vitae risus tristique posuere.

Question text goes here

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique. Duis cursus, mi quis viverra ornare, eros dolor interdum nulla, ut commodo diam libero vitae erat. Aenean faucibus nibh et justo cursus id rutrum lorem imperdiet. Nunc ut sem vitae risus tristique posuere.

Question text goes here

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique. Duis cursus, mi quis viverra ornare, eros dolor interdum nulla, ut commodo diam libero vitae erat. Aenean faucibus nibh et justo cursus id rutrum lorem imperdiet. Nunc ut sem vitae risus tristique posuere.

Question text goes here

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique. Duis cursus, mi quis viverra ornare, eros dolor interdum nulla, ut commodo diam libero vitae erat. Aenean faucibus nibh et justo cursus id rutrum lorem imperdiet. Nunc ut sem vitae risus tristique posuere.

Question text goes here

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique. Duis cursus, mi quis viverra ornare, eros dolor interdum nulla, ut commodo diam libero vitae erat. Aenean faucibus nibh et justo cursus id rutrum lorem imperdiet. Nunc ut sem vitae risus tristique posuere.

No items found.
Question text goes here

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique. Duis cursus, mi quis viverra ornare, eros dolor interdum nulla, ut commodo diam libero vitae erat. Aenean faucibus nibh et justo cursus id rutrum lorem imperdiet. Nunc ut sem vitae risus tristique posuere.

Question text goes here

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique. Duis cursus, mi quis viverra ornare, eros dolor interdum nulla, ut commodo diam libero vitae erat. Aenean faucibus nibh et justo cursus id rutrum lorem imperdiet. Nunc ut sem vitae risus tristique posuere.

Question text goes here

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique. Duis cursus, mi quis viverra ornare, eros dolor interdum nulla, ut commodo diam libero vitae erat. Aenean faucibus nibh et justo cursus id rutrum lorem imperdiet. Nunc ut sem vitae risus tristique posuere.

Question text goes here

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique. Duis cursus, mi quis viverra ornare, eros dolor interdum nulla, ut commodo diam libero vitae erat. Aenean faucibus nibh et justo cursus id rutrum lorem imperdiet. Nunc ut sem vitae risus tristique posuere.

Question text goes here

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique. Duis cursus, mi quis viverra ornare, eros dolor interdum nulla, ut commodo diam libero vitae erat. Aenean faucibus nibh et justo cursus id rutrum lorem imperdiet. Nunc ut sem vitae risus tristique posuere.

How to Run a 30-Day AI Support Agent Pilot in B2B SaaS (2026)

written by Ami Heitner
September 24, 2026
How to Run a 30-Day AI Support Agent Pilot in B2B SaaS (2026)

Ready to see how it works?

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
🎉 Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.