🧩 Complete Guide

SaaS Pilot Programs: How to Run a Proof of Concept That Converts

Most pilots are not lost at the readout. They are lost in week two, when the six people who were going to try it properly did not, and nobody noticed until the meeting.

📅 Updated September 2026 ⏱ 13 min read ✍️ By Kompassify
A comparison of three evaluations that are often confused: a free trial where an individual tests fit and it ends by lapsing, a pilot or proof of concept where a buying committee tests a business case and it ends in a decision, and a beta where existing users test an unfinished feature and it ends in a release

The pilot was agreed with real enthusiasm. Eight named users, four weeks, a shared channel, and a readout in the diary. Week one goes well: everyone logs in, there are questions, somebody imports a real dataset and says it looks promising.

Week two is quiet. Week three is silent. On the Thursday before the readout somebody from the customer side asks whether the pilot could be extended, because people have been busy and they want to give it a fair run. Sometimes the extension produces a decision. Usually it produces a second quiet month and a polite note in the autumn saying priorities have shifted.

Nothing went wrong with the product. What went wrong is that a pilot is an onboarding problem wearing a commercial costume, and it was run as a commercial process alone. This guide covers what separates a pilot from a trial and from a beta, the three things to agree before day one, how to scope it, the week-two usage cliff that decides most pilots, the mid-pilot checkpoint, how to run a readout that produces an answer, and what to measure while it is running.

Key Takeaways

  • A pilot is defined by its ending. Agreed criteria, a date, and a named decision maker. Without all three it is an extended trial.
  • Pilots die of inactivity, not of disappointment. The usual cause of failure is that too few named participants used it enough to judge.
  • Scope narrow enough to finish one real cycle. A broader pilot does not produce more proof, only more calendar.
  • The mid-pilot checkpoint is the whole intervention. By the readout it is already decided.
  • Measure participation, not sentiment. How many named users reached a real outcome is the number that predicts the decision.

What a Pilot Actually Is

Pilot program, in one paragraph

A pilot program, also called a proof of concept, is a time-boxed evaluation in which a prospective customer uses your product for real work, with real data and a named group of people, in order to decide whether to buy. What makes it a pilot rather than a long trial is the decision attached to the end: agreed success criteria, a fixed date, and a person who will answer yes or no. Remove any one of those and the evaluation will expire rather than conclude.

The word is used loosely, which causes real damage because the three things it gets confused with need opposite handling.

Free trial Pilot / POC Beta
Who takes part One person, self-selected A named group inside a buying committee Existing customers who opted in
What is being proven Whether the product fits this individual Whether it delivers an agreed business outcome Whether an unfinished feature is ready
Who decides The user, informally A named sponsor, formally Your product team
How it ends It lapses A readout and a decision A release
Typical failure One user never reaches value Too few participants used it to judge Feedback too thin to act on

Two consequences follow. First, the playbook that works for free trial conversion — optimise for one person reaching value quickly — is necessary but not sufficient in a pilot, because a pilot is judged on the group. Second, a beta program is aimed at learning and can tolerate an unfinished experience; a pilot is aimed at proving, and cannot.


Three Things to Agree Before Day One

Almost every rescued pilot turns out to have been rescued at the start rather than in the middle. These three agreements take an hour and they determine whether the evaluation can end at all.

1. Success criteria in the customer’s numbers

Two or three outcomes, written down, stated in the units the customer already uses, and observable within the pilot window. Approvals completed in under a day instead of three. Two hundred records processed without a manual correction. The Monday report produced without exporting anything. Criteria like the team finds it easy to use cannot be disproved, which means they cannot be proved either, and a pilot that cannot be failed is a pilot that cannot be passed.

Where the criteria are not visible in your product, decide in advance how they will be observed. The alternative is a readout in which both sides recall the last four weeks differently.

2. Scope: one real cycle, one real dataset, one real team

Pick the narrowest slice of work that, if it works, transfers to everything else. The instinct on both sides is to prove as much as possible, which produces a pilot spread thinly across five teams and conclusive about nothing. A pilot that proves one workflow end to end with genuine data is more persuasive than one that touched six workflows with test records, and it is dramatically easier to finish.

3. A date and a named decision maker

The readout goes in the calendar before the pilot starts, with the person who will decide present at that meeting. If nobody can name the decision maker in advance, that is the most useful thing you will learn all quarter, and it is better learned in week zero. Agree what happens on each outcome too, including what a “no” means, because a pilot with no defined negative outcome tends to end in an extension instead of an answer.

A pilot without a defined ending does not fail. It evaporates. Extensions feel generous and are usually the moment a pilot stops being a live evaluation and becomes an unattended account.


Choosing the Cohort

The list of participants is often assembled by whoever is available, which quietly decides the outcome. Three rules make the group representative enough to prove something and small enough to support.

Small and named. Six to ten people, identified individually rather than as a team. “The operations team” produces a pilot where everyone assumes someone else is doing it. A list of names produces accountability and, just as importantly, gives you a denominator to measure participation against.

Include a sceptic. A pilot staffed entirely by enthusiasts proves that enthusiasts like your product. The objection that will be raised in the buying decision is better raised in week two, where you can still answer it.

Include the person who does the work daily. Pilots evaluated only by managers measure impressions of the product rather than the product. The daily user is also the one whose behaviour tells you whether the value is real, because they are the only participant for whom using it is not a favour.


The Week-Two Cliff

Pilot usage follows a shape so consistent it can be planned for. Week one is busy, because the pilot is new and there is a kick-off. Week two collapses, because ordinary work resumes and nothing in the participant’s day depends on the pilot. Weeks three and four are quiet. Then there is a small spike just before the readout, when people remember the meeting.

What pilot usage really looks like Active participants per week, in a pilot that later ended without a decision 8 4 0 Week 1 Week 2 Week 3 Week 4 Readout Kick-off energy Meeting panic Guidance in the product lets them restart alone here Mid-pilot checkpoint while there is still time The readout does not decide the pilot. Week two does, and nothing in the sales process can see it.

Usage concentrates at both ends. The middle is where evidence was supposed to be gathered.

The cliff is not a sign of disinterest. It is what happens to any voluntary activity once novelty fades: the participant intended to use the product, then a real deadline arrived. By the time they return, they have forgotten the setup they did in week one and face a small relearning cost, which is precisely the point at which people postpone.

The remedy is unglamorous and effective: make returning cheap. A participant coming back in week three should find their place immediately, with the next step visible, without needing to ask anybody. That means in-product guidance that survives the gap — a checklist showing what they have done and what remains, contextual help on the screens they only saw once, and a short prompt on the path to the pilot’s success criteria. Everything that depends on a human being available will fail here, because the relearning happens at 7:40 on a Tuesday.

It is also the reason pilot participants should be onboarded as carefully as paying customers rather than more casually. The pilot cohort has less patience than a new customer, not more: they have not yet committed, and the cost of quietly abandoning is zero. The activation principles in our guide to increasing user activation apply here with the volume turned up.


The Mid-Pilot Checkpoint

Put a checkpoint at the end of week two, in the calendar from day one, and treat it as the most important meeting of the pilot. The readout is a reporting event; the checkpoint is the only moment where the outcome can still be changed.

Bring participation data rather than impressions: how many named users have logged in, how many have reached each success criterion, where people stopped. That last item is the one that produces useful conversation, because it converts a vague concern into something fixable. Three participants stopped at the same configuration screen is a solvable problem. “It has not been used much” is not.

Kompassify product tour analytics showing started, finished and skipped tours plus per-step completion rates, the participation data a mid-pilot checkpoint needs

Per-step completion turns the mid-pilot checkpoint into a factual conversation: who reached which criterion, and where people stopped.

The checkpoint also protects you from the extension request. When a mid-pilot review has already established that four of eight people are active, an extension can be negotiated as a re-engagement plan with a new date, rather than granted as a vague continuation. Extensions offered without a change to the plan generally produce a second period of identical usage.

A useful discipline. If participation is below half the cohort at the checkpoint, address participation before touching anything else. There is no product objection worth answering when the objection is that nobody has looked yet.


Running the Readout

The readout works best as a short, structured meeting rather than a demonstration. Go through the agreed criteria one at a time and state plainly whether each was met, including any that were not. Conceding a miss early is what makes the rest credible, and a pilot that met two of three criteria with a clear explanation of the third converts more often than one presented as a flawless success.

Two additions make the meeting land. Show the participation picture — who used it, how much, what they achieved — because a sponsor who has heard nothing for a month needs to know the evidence is real. And bring the rollout sketch: what the first ninety days after a yes would look like, who would be onboarded in what order, and what it would require from them. A decision maker saying yes is not buying the pilot, they are buying the deployment, and the sketch answers the question they are actually weighing. The handover principles in our guide to the sales to onboarding handoff apply from the moment that yes arrives, and pilots convert faster when the transition looks already designed.

Where the pilot is heading toward a multi-team deployment, the account-and-seats structure in the enterprise onboarding guide is the right shape for that ninety-day sketch.


What to Measure During a Pilot

1. Named-user participation

The share of the agreed cohort that used the product at all this week. It is the earliest signal available and the one most strongly related to the eventual decision. Track it as a fraction of the named list, never as total sessions, because two enthusiastic users can make an inactive pilot look healthy.

2. Criteria reached

For each agreed success criterion, how many participants have actually achieved it in the product. This is the readout, assembled continuously rather than the night before, and it turns the final meeting into a confirmation rather than a reveal.

3. Time to first real outcome

Days from access to the first participant completing genuine work, not a sample. It is the pilot version of time to value, and when it runs past the first week the cliff will do the rest. Where it is slow, the cause is almost always setup: data, permissions, or an integration nobody owned.

4. Drop-off point

The step where participants most often stop. In a pilot this is worth more than any aggregate, because the group is small enough that a single blocked screen explains the whole curve — the same reasoning behind an onboarding funnel, applied to a cohort of eight.


Seven Ways Pilots Fail

1. No decision maker in the room

The pilot is run by an enthusiastic champion who turns out not to hold the budget. Everything goes well and nothing happens. Name the decider before day one or accept that this is exploration, not an evaluation.

2. Criteria written after the fact

When success is defined at the readout, both sides define it as whatever happened, and the meeting becomes a negotiation about interpretation rather than a decision.

3. Scope that cannot finish

Five teams, three workflows, two integrations and four weeks. Breadth feels safer and guarantees that nothing is proved conclusively.

4. Sample data instead of real data

A pilot run on demo records proves the product works on demo records. Real data is where the edge cases live, and edge cases are what the sceptic will ask about.

5. Onboarding left to the kick-off call

One live session, no guidance afterwards, and a participant returning in week three who cannot remember where anything is. This is the most common and the most fixable of the seven.

6. No visibility into usage

A pilot tracked purely as a deal stage means the first evidence of trouble arrives at the readout, when nothing can be done. Somebody on your side should be able to see participation weekly.

7. The automatic extension

Granting more time without changing the plan repeats the same month. If an extension is warranted, it should come with a new cohort, a smaller scope or a new date — something different.


Running a Pilot With Kompassify

The part of a pilot that most often goes unmanaged is the participant experience between the kick-off and the readout, because it happens inside the product where the commercial process cannot see it.

With Kompassify you can give the pilot cohort its own guidance layer without engineering work: a short walkthrough on first entry that goes straight to the agreed success criteria rather than a tour of everything, a pilot onboarding checklist whose steps are the criteria themselves so progress is visible to the participant and to you, contextual help on the screens people only saw once during the kick-off, and a prompt that brings a returning participant back to the next step in week three instead of the home screen. Because flows are segment-targeted, all of it applies to the pilot group only and disappears when the account converts to a normal deployment.

Per-step completion data also turns the mid-pilot checkpoint into a factual conversation: you can say which criteria have been reached and by whom, and where in the flow people stopped, four weeks before anyone has to form an opinion about it.

Stop losing pilots in week two

Give every pilot cohort a guided path to its own success criteria, see which participants got there, and fix the blocked step while the evaluation is still running. No code required. Free up to 100 monthly active users, plans from $129/month, GDPR-compliant and EU-hosted.

Start for free →

The One-Paragraph Version

A pilot is an evaluation with a decision attached, which means agreed criteria in the customer’s own numbers, a fixed date and a named decision maker, all settled before day one. Scope it to one real cycle of real work with a small named cohort that includes a sceptic and someone who does the job daily. Then plan for the shape pilot usage actually takes: a busy first week, a collapse in the second, and a spike of panic before the readout. Make returning cheap with in-product guidance that carries a participant back to the next step without a human, hold a mid-pilot checkpoint at the end of week two with participation data rather than impressions, and treat low participation as the problem to solve before any product objection. Run the readout against the criteria one by one, concede what was missed, and bring a ninety-day rollout sketch, because the sponsor is deciding about the deployment rather than the pilot. Measure named-user participation, criteria reached, time to first real outcome and the drop-off point — and treat any extension offered without a change of plan as the beginning of a second quiet month.


Frequently Asked Questions

What is a pilot program in SaaS?

A pilot program, often called a proof of concept, is a time-boxed evaluation in which a prospective customer uses your product with real work, real data and a defined group of people, in order to decide whether to buy it. The defining feature is not the length or the price but the decision attached to the end of it: a pilot has agreed success criteria, a date and a named person who will say yes or no. An evaluation without those three things is not a pilot, it is an extended trial that will expire quietly.

What is the difference between a pilot and a free trial?

A free trial is self-serve, individual and open-ended in intent: one person explores the product to see whether it fits, and the trial ends by lapsing. A pilot is committee-based, scoped and decision-bound: a group uses the product against agreed criteria, and it ends in a meeting where somebody decides. They also fail differently. A trial fails when a single user does not reach value, while a pilot usually fails because too few of the named participants used it enough for anyone to judge.

How long should a SaaS pilot last?

Long enough to complete one real cycle of the work your product supports, and no longer. For most products that is four to six weeks; for products tied to a monthly or quarterly process it is one full cycle plus a week. Pilots that run longer rarely produce more evidence, because usage concentrates in the first two weeks and the last one. If a pilot needs three months to prove anything, the proof being sought is usually too broad, and narrowing the scope is a better answer than extending the calendar.

Should a pilot be paid or free?

A paid pilot buys commitment: when a budget line exists, participants are named, time is allocated and someone internally is accountable for a result. A free pilot lowers the barrier to starting and raises the risk of drifting, because nobody in the customer's organisation has to defend the time spent. Where the implementation requires real effort from the customer, charging something, even a nominal implementation fee, generally improves completion rates more than any amount of follow-up.

Why do pilots fail even when the product works?

Because insufficient use was made of them. The most common pilot failure is not a product that disappointed but a cohort that never engaged: accounts were created, two of the eight named users logged in, and at the readout nobody could speak to the outcome. Enthusiasm is highest in the first days and drops sharply once ordinary work resumes, so the practical job during a pilot is keeping a small group active for a few weeks, which is an onboarding problem rather than a sales one.

What should the success criteria of a pilot be?

Two or three outcomes stated in the customer's own numbers, agreed in writing before the start, and measurable inside the pilot window. Good criteria look like a specific process completing faster, a specific volume being handled, or a specific error disappearing. Bad criteria are satisfaction, ease of use, or whether the team likes it, because they cannot be disproved and therefore cannot be proved either. If the criteria cannot be observed in the product, decide before the pilot how you will observe them.

Who should own the pilot internally?

One named person on each side, with the vendor side owning momentum and the customer side owning the decision. The frequent failure is a pilot owned by a sales representative who tracks it as a deal stage and sees nothing that happens inside the product, so the first signal of trouble is the readout itself. Pairing the commercial owner with whoever can see usage, and agreeing a mid-pilot checkpoint in advance, fixes most of it.