What a 6-Week AI PoC Should Deliver — and When to Kill It

What a 6-Week AI PoC Should Deliver тАФ and the Exit Criteria That Kill It Early

TL;DR

  • About one-third of organizations have begun scaling AI across the enterprise, and roughly 95% of generative AI deployments deliver no measurable impact to the P&L.
  • The cause is a proof of concept built to provide a demonstration when the buyer needs to make a decision.
  • A six-week AI PoC hands over a decision package: a working build on a controlled slice of your own data, a cost model you can budget against, a security blueprint your team can review, and a plan for the next build.
  • It also carries five kill gates with written thresholds, covering data, baseline, cost, workflow, and the decision itself.
  • A stopped PoC is not a failed engagement, and it costs a fraction of the build it prevents.

Most AI PoCs are designed to impress rather than to decide

A proof of concept is not a demonstration. It is a decision, and the decision is what you are paying for. The best-run PoCs are the ones that can fail, because they carry written thresholds, weekly gates, and a delivery team willing to say stop in week two. A strong proof of concept doesn’t exist to force a “yes.” It exists to tell the truth early.

The numbers behind that claim hold up across the research, even where the framing varies. MIT’s Project NANDA published The GenAI Divide: State of AI in Business 2025 in July of that year. It found that 95% of organizations saw no measurable return from generative AI, while only 5% of integrated pilots extracted significant value. McKinsey’s State of AI survey puts about one-third of organizations at the point of beginning to scale, leaving close to two-thirds that have not yet. Gartner reviewed hundreds of implementations and reports that at least half of generative AI projects were abandoned after the proof-of-concept stage by the end of 2025. It names four causes: poor data quality, inadequate risk controls, escalating costs, and unclear business value.

None of those is a technology failure. The pattern behind them is procedural and repeats often enough that AWS includes it in its own guidance. PoCs get treated as technical demos designed to impress, when the organization needs a rigorous experiment designed to produce learning. A vendor demonstrates a capability, and the team gets interested. Someone then launches a PoC with no written hypothesis, no agreed baseline, no named owner, and no exit criteria. The result is interesting but inconclusive, which is the worst outcome on the menu. An inconclusive PoC spends the budget and leaves the decision unmade.

What follows is pilot purgatory, where the project is neither killed nor deployed. It continues to draw low-level attention while the organization moves on to the next thing. We treat this as the first of the seven enterprise AI adoption risks, and the instrument described below is how you prevent it.

The reframe is straightforward. You are not buying a prototype; you are buying a decision. That decision carries the most value when it comes back negative, because a negative answer saves money. This article covers what a serious AI proof-of-concept development engagement hands over after six weeks. It also covers the five gates along the way that give the engagement permission to stop.

PoC, prototype, MVP, and pilot: the distinction that sets your scope

The vocabulary needs to be settled before the gates, because scope disputes usually turn out to be definitional disputes with a budget attached. SumatoSoft’s published formulation is the cleanest one available. A PoC tests feasibility in a controlled scope. An MVP is a minimal but working product you put in front of users. A pilot runs a validated solution with a limited group before full rollout. They arrive in that order, and starting with an MVP before feasibility is settled is how budgets get spent on ideas that were never going to work.

StageThe question it answersWhen you run itWhat it produces
PoCCan this work here, on our data and inside our constraints?Before committing build budgetEvidence and a go/no-go recommendation
PrototypeDoes the workflow and interface hold up with users?Once feasibility is settledA testable interaction model
MVPWill users adopt it and pay for it?Once the case has been madeA minimal working product in market
PilotDoes it survive live conditions at limited scale?Once the product is readyOperational evidence before full rollout

One distinction carries more weight than the rest. Traditional software PoCs test whether something can be built, while AI PoCs test whether a probabilistic system can be trusted. In conventional software, the code either works or it does not. Behavior that holds in the test environment tends to hold in production. Machine learning and generative systems do not behave that way because the same prompt can return different outputs on different days when the inputs are slightly different. That variance is why an AI PoC needs an evaluation set. An evaluation set is a fixed collection of test cases with known correct answers, and it replaces the demonstration that once impressed a room. For the fuller treatment, see our guidance on the difference between a PoC and an MVP.

The decision package: what a six-week PoC hands over

Most AI PoC services stop at “you get a prototype,” which is not enough for a serious buying decision. A prototype shows that something can work once, under conditions the builder selected. A decision package tells you what happens next and what it costs. SumatoSoft’s published program delivers four artifacts. The framework in this article adds two more that turn the package into evidence.

A working build on a controlled slice of your own data. Sample data proves nothing about your business. The difficulty in enterprise AI usually lives in the shape of your records rather than in the model. The privacy default runs opposite to what teams expect. Synthetic or fully anonymized data comes first, and production records containing personally identifiable information enter the environment only after your information security and legal teams approve the ingestion in writing. AWS names skipping that approval as one of the most common failures in generative AI PoC work.

A cost model you can budget against. The pilot ran on one document and production runs on a million. That gap is where business cases go to die. The model covers token economics, per-call ceilings, retrieval volume, and the unit cost of a single transaction under agreed usage assumptions. SumatoSoft’s published position is unambiguous: “We never let the system run as an open meter. Cost visibility and usage limits are part of how we govern delivery.”

A security blueprint your team can review. Your security function reads this document before anything scales. It covers data ingress and retrieval flow, access boundaries and permission layers, isolation options, and audit logging with retention rules. Any system that can call an external service or trigger an internal action needs those limits designed in from day one rather than added afterward.

A plan for the next build. Build phases with deliverables, integration points, evaluation checkpoints, acceptance criteria, and release governance. An approved PoC then leads to an MVP without having to rediscover the environment from scratch.

An evaluation set and results against a named baseline. This artifact separates evidence from theater. The evaluation set exists before the build starts. The baseline is a current number rather than a feeling about current performance. The readout reports quality, coverage, failure modes, and refusal behavior rather than a screenshot of a good day.

Ownership of everything produced. The code, prompts, architecture, and delivered assets belong to your company under the project agreement. A PoC you do not own has limited evidentiary value, because you cannot inspect or extend it.

The decision package

The five kill gates

A gate that cannot fail is a status meeting with a better name. The thresholds get written and signed before the build starts. A criterion agreed after the results are visible has become a rationalization. Gartner’s April 2026 research points in the same direction from the investment side. It found that organizations reporting successful AI initiatives invest up to four times more of their revenue in foundational areas such as data quality, governance, AI-ready people, and change management. The gates below are ordered by how cheap the kill is, and week one is the cheapest place in the program to stop.

The five kill gates

Gate 0, before week one: Is there a hypothesis, a baseline, and an owner?

This is the pre-gate, and it offers the cheapest exit available: choosing not to start. Three things have to exist in writing. The hypothesis must be specific enough to be wrong, so “automatically classify and route the 40% of tickets that ask routine billing questions” qualifies where “reduce support costs” does not. The baseline must name a current number, because a result you cannot compare against is an anecdote. The owner must be a business leader accountable for the outcome, rather than an innovation team without a path into operations.

The threshold: all three, in writing, signed by the accountable owner. If it fails: do not run the PoC. You would be buying an experiment with no way to read the result. Our five-step framework for planning custom AI automation covers the scoping work that has to happen first.

Gate 1, week one: Does the data exist, and can you reach it?

Data readiness kills more PoCs than any other single factor. Organizations tend to discover the problem three weeks in, when the timeline slips and the scope expands. Confidence then drains away before any result exists. Gartner found that 63% of organizations either lack the right data management practices for AI or are unsure whether they have them. Our own survey of 72 executives puts data quality and consistency at the top of the readiness-gap list, cited by 58% of respondents.

The audit is therefore the first deliverable rather than the first obstacle. It answers four questions. Can you access the data? Is it clean enough to test against? Does it contain PII that requires handling before ingestion? Is there enough of it?

The threshold: accessible, sufficient, and clean enough to test within the PoC window, rather than fixable eventually. If it fails: you have a data project rather than an AI project. The honest move is to stop or convert it, while saying which one. An AI readiness assessment answers this question for a fraction of PoC cost when you already suspect the answer.

Gate 2, weeks two and three: Does it beat the baseline on your data?

This is the technical gate. The thin slice runs against the evaluation set built at Gate 0, measuring quality, coverage, failure modes, and refusal behavior across the full test set. A curated demonstration proves nothing here, because a probabilistic system that impressed once may not repeat.

The margin carries more weight than the metric. Beating the baseline by a rounding error reads as a no. The production version will perform worse than the PoC version, not better, once volume, edge cases, drift, and adversarial inputs arrive. Model performance degrades the moment a system goes live, which is why the gate requests headroom.

The threshold: situational and written at Gate 0. The method is fixed even where the number is not. Choose a margin wide enough to withstand the degradation from a controlled test to a live system. If it fails: stop. This is the gate that teams most want to extend to tune the results, and extending here is the mechanism by which six weeks becomes six months. Where the approach looks sound, but the implementation needs work, custom machine learning development is a separate conversation from this PoC.

Gate 3, week four: Can it reach production, and what does it cost to run?

Two questions share a gate because they tend to fail together. The integration probe builds a thin but live connection to the system that worries you most, whether that is the core platform or the legacy back office. It proves authentication, data contracts, retry semantics, and failure behavior against the running system rather than against a mock. Failed machine learning pilots frequently trace back to this point, where the data science work was sound, and the integration with existing databases was never attempted.

The cost model then projects run costs at production volume. AI programs without cost discipline routinely overrun their projections by multiples. Our breakdown of AI development costs covers the drivers in more depth.

The threshold: an integration path that has been demonstrated, plus a run cost the business case survives at projected volume. If it fails: stop, or re-scope toward a use case whose economics work. A system that performs well technically and negatively on unit economics is still a no.

Gate 4, week five: Will the workflow change?

This is the gate that almost no competing framework includes, and our first-party research directly supports it. We surveyed 72 executives and functional leaders across more than 30 industries. Workflow redesign was the single biggest factor in moving AI from pilot to production, named by 61% of respondents. Data readiness followed at 22%, executive sponsorship at 14%, and MLOps at 3%.

The question at week five is therefore not whether the model works, since Gate 2 settled that. The question is whether anyone will change how they work because of it. Week five is where you find out, rather than three months after launch.

The evidence is behavioral. The business owner commits to the process change in writing. The people who use the system engage with its outputs rather than routing around them. An operating budget with a named owner is in place for the period after the PoC ends.

The threshold: a named owner who has committed to the workflow change, plus evidence of engagement from the people who would use it. If it fails: stop. You have proved the technology and disproved the adoption, and the second finding kills projects more reliably than the first.

What moves AI from pilot to production

Gate 5, week six: The decision

The sprint ends with your decision to proceed, refine, or stop. Three outcomes rather than two is a meaningful distinction, and “refine” is not a polite way to stop. It is a bounded second pass with a new threshold and a new date. Refine is available only where the failures sat in adoption or in the economics. A no-go on data readiness or on beating the baseline terminates the engagement, because those two failures describe conditions that a second iteration cannot change inside a PoC budget.

The output format carries weight here as well. The engagement ends with a go/no-go recommendation, backed by data and delivered as a written readout. It covers what the PoC set out to prove, what worked, what the system will cost to run, and whether the case is strong enough to proceed. If the recommendation cannot fit on a page with the attached evidence, the PoC is not finished.

What a “no” is worth

Every competing framework treats the kill as the disappointing outcome, then reassures the reader that learning has value. The arithmetic is more persuasive than the reassurance. A PoC costs a defined, fixed amount agreed before the work starts. The build-it gates cost an order of magnitude more. The research supplies the probability. Roughly two-thirds of organizations have not begun scaling AI, and about 95% report no measurable P&L return.

S&P Global Market Intelligence found that 42% of companies abandoned most of their AI initiatives in 2025, up from 17% the year before. The average organization scrapped 46% of its proofs of concept before production. Under those odds, the expected value of finding out early becomes the largest single line in the business case.

Consider the arithmetic with published ranges. A six-week PoC in the low-to-mid five figures gates a production build that SumatoSoft prices at roughly $100,000 to $400,000 and up. Ongoing monitoring and retraining add a further 15% to 20% of build cost per year. Stopping at Gate 1 costs a fraction of the PoC. At Gate 5, it costs the PoC. Stopping in month nine, with half a build delivered and a team committed, costs the PoC plus six figures. These figures are illustrative arithmetic rather than a study finding. The conclusion holds when you substitute your own build estimate, because the ratio between the two numbers drives the result.

The cost of finding out late

The reframe worth taking to your CFO is a comparison. The question is not what happens if the PoC says no. The question is what it would have cost to find that out in month nine rather than in week two.

One condition attaches to all of this. The arithmetic works only where the PoC can fail. A PoC run by a vendor who needs a yes, with criteria written after the results arrive, produces a yes of no evidentiary value. The exercise was never a test. A stopped PoC is not a failed engagement. If the sprint shows that the data, the economics, or the delivery conditions aren’t strong enough yet, it has done its job. It has saved you from a bigger mistake.

Which instrument do you need?

Two shapes of proof of concept

Not every question needs six weeks. A spike answers one question and finishes inside four weeks, because there is a single unknown to resolve. That unknown might be an integration your team has never attempted, or a model that has never run against your data. It might equally be an unmeasured workload or an untested compliance constraint. A full PoC is what this article describes. It runs discovery, a build on a bounded slice of your data, an integration probe, a cost model, and a decision package, with five gates along the way. It takes six weeks because five separate conditions can kill it, and each one takes a week to answer honestly.

The routing rule is short. If you can name the single thing you do not know, buy a spike. If you cannot, you need the six weeks, because “we are not sure it will work” describes five questions rather than one.

ShapeThe question it answersDurationWhat you getBuy it when
SpikeCan this one thing work?Up to four weeksA technical answer and a recommendationYou can name the single unknown
Full PoCShould we build this at all?Six weeksThe decision package plus five gates with written thresholdsThe unknown is all of it
Two shapes of PoC

When it should not be a PoC at all

Four situations call for something other than either shape.

Where the data is not accessible, you have a data project. Running it as a PoC means paying PoC rates to discover that in week three. Where you already know the approach works technically and the open question is adoption, you want a pilot or an MVP rather than a feasibility test. When you have not shortlisted vendors, remember that a PoC tests the finalist rather than the field, and spending scarce trial capacity across a long list wastes both. Where the decision is cheap to reverse, and you could back out inside a week, run the thing and see, because insurance on a small loss is a poor trade.

The honest flip side is that six weeks is sometimes too short. Where the answer requires a production-condition test with live users under live load, you are describing a pilot. A pilot takes longer. The predictive maintenance program we ran for a wind energy operator ran for 8 weeks alongside live operations. It was a pilot rather than a PoC, answering a different question with different evidence.

Frequently asked questions

What should an AI proof of concept deliver?

A decision package rather than a prototype. That means a working build on a controlled slice of your data, a cost model at production volume, a security blueprint, and a plan for the next build. It also means an evaluation set with results against a named baseline, plus ownership of all produced outputs.

How long should an AI PoC take?

Four to six weeks, depending on shape. A spike resolving a single unknown is completed within four weeks. A full PoC with discovery, an integration probe, and five gates runs six weeks. Anything longer has usually stopped being a PoC and become an unbudgeted build.

How much does an AI PoC cost?

Fixed scope and fixed price, agreed in the contract after a discovery call. A SumatoSoft PoC typically falls in the low-to-mid five figures. The exact number depends on the complexity of the use case and the readiness of your data. The comparison worth making is against a full build, which runs from roughly $100,000 upward.

What is the difference between an AI PoC and an MVP?

A PoC tests feasibility in a controlled scope, answering whether the technology can work inside your constraints. An MVP is a minimal but working product put in front of users to determine whether they will adopt and pay for it. Feasibility comes first.

When should you kill an AI PoC?

At the first gate whose threshold it misses. Data that cannot be accessed or cleaned inside the window kills it in week one. A result that fails to clear the baseline by a defensible margin kills it in week three. An integration that cannot be built kills it in week four, as does a run cost the business case cannot absorb.

What exit criteria should an AI PoC have?

Written thresholds attached to each gate, agreed before the build starts. Avoid industry-standard accuracy percentages, since useful thresholds are situational. Set the number from your baseline, the margin needed to survive production degradation, and the point at which the economics stop working.

Who owns the code and IP from a PoC?

Under a SumatoSoft project agreement, your company does. The code, the prompts, the architecture, and the delivered assets belong to you. That ownership is what allows the PoC to serve as evidence rather than as a demonstration you rented.

Our PoC succeeded, so why will it not go to production?

Usually Gate 4. The technology cleared its threshold, and the workflow never changed. Our research identifies workflow redesign as the single largest factor separating pilots that reach production from those that stall. Cost is the second common answer, where run economics at production volume were never modeled.

Conclusion: buy the decision, not the demo

The proof of concept is the cheapest decision in an AI program. It is also the only one where being wrong costs almost nothing. Teams that run it as a sales demonstration receive a yes and a stalled project. Teams that run it as an experiment, permitted to fail, receive an answer they can act on either way, and the answer arrives while acting on it remains inexpensive.

SumatoSoft has spent 14 years and built more than 350 custom products this way across 25+ countries, under ISO 27001 and ISO 9001 certifications. The AI Pilot & Prove program runs inside our Agentic Development Lifecycle. That is a governed delivery model where AI operates within defined boundaries from day one, and every engagement ends with a stated willingness to recommend stopping. The vendor worth hiring is the one who told you how to kill the project before you signed for it.

A strong proof of concept doesn’t exist to force a “yes.” It exists to tell the truth early.

Run a PoC that is allowed to fail

Bring your riskiest assumption. Six weeks, five gates, and a decision package: a working build on your own data, a budgetable cost model, a security blueprint, and a plan for the next build. Or a clear and evidenced no. Fixed scope and fixed price, agreed in the contract after a discovery call. You own everything we produce. ISO 27001 certified.

Scope your AI PoC

Let’s start

You are here
1. Submit your project brief
2. Connect with our strategy team
3. Finalize scope & investment
4. Start achieving your goals

If you have any questions, email us info@sumatosoft.com

    Please be informed that when you click the Send button Sumatosoft will process your personal data in accordance with our Privacy notice for the purpose of providing you with appropriate information.

    Vlad Fedortsov (Account Manager)
    Vlad Fedortsov
    Account Manager
    Book an intro call
    Thank you!
    Your form was successfully submitted!
    SumatoSoft logo
    If you have any questions, email us info@sumatosoft.com

      Please be informed that when you click the Send button Sumatosoft will process your personal data in accordance with our Privacy notice for the purpose of providing you with appropriate information.

      Vlad Fedortsov (Account Manager)
      Vlad Fedortsov
      Account Manager
      Book an intro call
      Thank you!
      We've received your message and will get back to you within 24 hours.
      Do you want to book a call? Book now
      SumatoSoft clients logo