The 2-4 Week Prototype + 90-Day Retail Pilot: What It Actually Means in 2026
In short
A prototype is a working, single-use-case AI system delivered in a sandbox or 1-2 stores over 2-4 weeks, judged on functional acceptance criteria. A 90-day pilot deploys that prototype into a controlled production cohort of 3-10 stores or one channel segment, measured against a control group on business KPIs. The standard 2026 retail cadence: Week 0-4 prototype, Week 5-17 pilot, Week 18+ scale decision.
The 2-4 week prototype and the 90-day pilot are two distinct contracts, and retail buyers in 2026 pay for them separately for a reason: it is the only structure that lets them kill a weak idea at Week 4 instead of Week 17, and at Week 17 instead of Week 52.
A prototype is a working, single-use-case AI system delivered in a sandbox or one to two stores. It exists to answer one question — does this work at all on our data? — and it is judged on three pre-agreed acceptance metrics, not on business KPIs. Alice Labs delivers prototypes in 2-4 weeks because the scope is locked to a single use case and the team is senior-only, so there is no junior handoff eating a week.
A pilot is a production-guarded deployment of that same prototype into a controlled cohort — 3-10 stores or one channel segment — running 60-90 days against a control group. It exists to answer the second question — does this move the business KPI enough to justify scaling? — and it is measured on margin, stockouts, basket size, shrink, and GMROI, not model precision.
The standard 2026 retail cadence has three gates:
- Weeks 0-4 — Prototype. Fixed-fee, single use case, three acceptance metrics, go/no-go to pilot at Week 4.
- Weeks 5-17 — 90-day pilot. Controlled cohort, weekly KPI review, statistical significance check at day 90.
- Week 18+ — Scale decision. New SOW for scale, or documented kill with the retailer keeping all IP.
Alice Labs has run this cadence across 100+ enterprise AI implementations since 2023. For a broader treatment of the full AI implementation lifecycle beyond retail pilots, see our AI implementation pillar and the deeper AI implementation roadmap for enterprise.
Working AI prototype delivery window for scoped single use cases
Standard pilot window with weekly KPI review and statistical significance gate at day 90
Why Retail Buyers Are Compressing Timelines in 2026
In short
Retail buyers now underwrite AI spend against 90-day proof points, not 12-month roadmaps. Gartner predicts 40% of agentic AI projects will be canceled by 2027 due to unclear ROI. MIT's 2026 State of AI in Business reports ~95% of generative AI pilots fail to produce measurable business impact. McKinsey finds 62% of enterprises remain stuck in experimentation. Compressed timelines are how retail CFOs are responding.
The compressed timeline is not a marketing preference — it is a direct response to three years of disappointing AI ROI. The numbers retail CFOs are looking at in 2026 are stark:
- Gartner (2026): 40% of agentic AI projects will be canceled by end of 2027 due to unclear ROI, inadequate controls, and escalating cost. Retail is over-indexed in this cohort.
- MIT State of AI in Business (2026): ~95% of generative AI pilots fail to produce measurable business impact. The pattern is not model quality — it is pilot design.
- McKinsey (2026): 62% of enterprises remain in the experimentation or pilot stage; only 23% have scaled AI agents to production.
- IHL Group (2026): 47% of retail and CPG organizations have adopted agentic AI in at least one function — but the gap between adoption and scale is where budget disappears.
The pattern is consistent: adoption is high, scale is low, and the delta is budgeted-but-killed pilots. Retail CFOs have responded by rewriting the underwriting rules. AI spend is no longer signed against a 12-month roadmap and a slide deck; it is signed against a 90-day proof point with a defined kill-gate and a measurable business KPI.
The 2-4 week + 90-day cadence exists because it matches how retail treasury actually thinks about risk. A CFO can approve a fixed-fee 2-4 week prototype from discretionary opex, review the acceptance metrics, and then approve the 90-day pilot on a fresh SOW. Neither approval requires committing to a scale contract that has not been earned yet. This is what a defensible AI budget looks like in 2026.
For the broader ROI framework we use across engagements, see our AI cost-benefit analysis guide and the AI ROI calculator.
Agentic AI projects Gartner predicts will be canceled by end of 2027 due to unclear ROI and lack of controls
The Alice Labs 2-4 Week Prototype Framework (Weeks 0-4)
In short
Week 0 is use-case triage against 100+ prior Alice Labs implementations — 60-70% of ideas get killed here. Week 1 is a data readiness audit and one-store or one-channel scope lock. Weeks 2-3 build the working model and integrate it into POS/PIM/OMS or the CV camera stack, senior-only. Week 4 delivers a live prototype demo against three pre-agreed acceptance metrics with a formal go/no-go to pilot.
The framework is deliberately simple because compressed timelines only work when the decision structure is unambiguous. Each week has one owner, one deliverable, and one gate.
Alice Labs 2-4 week prototype framework
| Week | Focus | Gate |
|---|---|---|
| Week 0 | Use-case triage against 100+ prior implementations; disqualify wrong-fit scopes | 60-70% of proposed ideas killed here |
| Week 1 | Data readiness audit; scope lock to one store or one channel; three acceptance metrics defined | Contractual data-readiness kill-gate — buyer walks with full IP if data is not workable |
| Weeks 2-3 | Working model built; integrated into POS, PIM, OMS or CV camera stack; senior-only build team | Weekly working-model checkpoint with client operator |
| Week 4 | Live prototype demo against three pre-agreed acceptance metrics | Go/no-go to 90-day pilot — separate SOW signed at this gate, never pre-bundled |
Week 0 is where most of the value gets created and none of it is visible to the client. Alice Labs kills 60-70% of proposed use cases in that first week by pattern- matching against 100+ prior implementations. The cheapest pilot is the one that never starts. When we tell a retailer "this specific use case will not pay back at your scale", we save them the 2-4 week prototype fee, the 90-day pilot cost, and — more importantly — the 6-9 months of change-management debt that a killed pilot leaves behind.
Week 1's data readiness audit is a hard kill-gate. If the data is not workable — missing SKU-store history, non-normalised planogram data, unlabelled camera footage — Alice Labs will not proceed and the buyer walks with the audit deliverable and full IP. This is the discipline that keeps the 2-4 week window credible. For deeper data-readiness patterns, see our data quality for AI guide and the AI data preparation guide.
The 90-Day Retail Pilot Structure (Weeks 5-17)
In short
Days 1-30 roll the prototype out to 3-10 stores or a single channel segment and establish an A/B baseline against control stores. Days 31-60 run weekly KPI reviews with human-in-the-loop refinements and edge-case documentation. Days 61-90 run the statistical significance check on the primary business KPI and model unit economics for full rollout. The standard pilot gate requires the primary KPI to move 8-15% versus control with p<0.05.
The 90-day pilot is the phase where most retail AI programs die quietly. Alice Labs structures it in three 30-day segments precisely to prevent that outcome — each segment has its own deliverable and each week has an accountable check-in.
- Days 1-30 — Rollout and baseline. The prototype deploys to 3-10 pilot stores or one channel segment. A matched set of control stores is established. The A/B measurement plan is registered — you cannot p-hack a pre-registered analysis, which is exactly the point.
- Days 31-60 — Iterate. Weekly KPI review with client operators. Human-in-the-loop refinements to model thresholds, alerting rules, and workflow integration points. Edge cases are logged the day they surface, not the day the pilot ends.
- Days 61-90 — Decide. Statistical significance check on the primary KPI. Unit economics modelled for the full estate. Scale SOW drafted if the gate passes; documented kill-report drafted if it does not.
The pilot gate is quantitative and pre-agreed: the primary business KPI must move 8-15% versus control with p<0.05 to recommend scale. Anything below the lower bound of that range does not justify the change-management cost of full rollout; anything above it is a real finding. The p-value guard exists because retail data is noisy — a two-week win against control can be seasonality, and pilots that skip significance testing are pilots that scale ideas that never worked.
Human-in-the-loop is embedded from day one. One senior Alice Labs engineer is paired with one client operator (store manager, category manager, or head of stores) for the full 90 days. This is the single largest change-management lever we have found: embedded operators surface the edge cases that kill a rollout, and they own the adoption metric on the other side of the gate.
For the KPI structure we use across pilots, see the AI measurement framework.
Retail Use Cases That Fit the 2-4 Week + 90-Day Envelope
In short
Shelf monitoring (planogram + out-of-stock CV) fits in 3 weeks with a 90-day pilot in 5-10 stores. SKU/store demand forecasting fits in 4 weeks; top-quartile pilots show 60-75% stockout reduction. Assortment-aware product recommendations fit in 2-3 weeks. Store associate copilots on inventory and product Q&A fit in 3-4 weeks. Loss-prevention CV at self-checkout fits in 4 weeks. ERP re-platforming and greenfield data foundations do not fit the envelope.
The 2-4 week envelope has a specific shape and it is worth naming which use cases fit and which do not. This is the table Alice Labs uses in Week 0 triage.
Retail AI use cases that fit the 2-4 week prototype + 90-day pilot envelope
| Use case | Prototype window | Pilot shape | Primary KPI |
|---|---|---|---|
| CV shelf monitoring (planogram + out-of-stock) | 3 weeks | 5-10 stores, 90 days | Out-of-stock hours reduced |
| SKU/store demand forecasting | 4 weeks | Chain-wide category subset, 90 days | Stockout %; top quartile shows 60-75% reduction |
| Assortment-aware product recommendations | 2-3 weeks | One digital channel, 90 days | Incremental basket size |
| Store associate copilot (inventory + product Q&A) | 3-4 weeks | 5-10 stores, 90 days | Shrink %, associate NPS |
| Self-checkout loss-prevention CV | 4 weeks | Single format, 5-10 stores, 90 days | Unscanned-item precision/recall; shrink impact |
All five share the same shape: one well-scoped decision surface, existing structured or camera data, and a KPI the retailer already measures weekly. This is why they fit in 2-4 weeks. Use cases that require multi-country ERP re-platforming, a greenfield data foundation, or a fully autonomous agent replacing a regulated workflow do not fit — and Alice Labs declines those scopes rather than accept an unrealistic frame.
For deeper coverage of use-case selection and where AI actually pays back, see our AI ROI by use case guide and AI implementation case studies.
What Actually Fails in 2-4 Weeks (And How Alice Labs Prevents It)
In short
Four failure modes account for most compressed-timeline pilots that die: (1) data readiness gaps — 38% of retailers cite data cleaning/training/storage as their top blocker (IHL 2026); (2) scope creep during the prototype; (3) change management vacuum — pilots that reach demo but never reach production because ops teams were not embedded; (4) vendor sprawl. Alice Labs prevents each with a Week 1 data audit kill-gate, contractual scope lock, embedded senior engineer paired with a client operator, and single-partner accountability through the pilot gate.
The failure modes are well-understood and repeatable. Alice Labs' six-year pattern library maps every one of them to a specific structural fix, not a heroic effort.
- Data readiness gap. IHL Group's 2026 retail AI survey found 38% of retailers cite data cleaning, training, and storage as their top adoption blocker. Alice Labs runs a Week 1 data readiness audit as a hard kill-gate — if the data is not workable, the engagement halts with full IP transfer to the buyer.
- Scope creep. Pilots die when three stakeholders each bolt on a secondary use case at Week 2. Alice Labs contracts lock scope to one primary KPI and one use case; secondary asks become the subject of the next SOW, not this one.
- Change management vacuum. Gartner and IDC report roughly 89% of AI agent pilots that reach demo never reach production. The root cause is consistently operations-team-not-embedded. Alice Labs pairs one senior engineer with one client operator for the full engagement — the operator owns the adoption metric.
- Vendor sprawl. Multi-vendor pilots slow every integration decision to the pace of the slowest vendor's roadmap. Alice Labs is single- partner accountable through the Week 17 pilot gate; any third-party tooling runs under one contract.
Each fix is structural. None of them require the client to change their operating model — they require the AI partner to hold a contract shape most vendors avoid. The 2-4 week + 90-day cadence stops working the moment any of these four discipline points is relaxed.
For the broader failure-mode taxonomy across AI programs, see our AI failure modes guide and the deeper AI organisational resistance guide.
Retailers citing data cleaning, training, and storage as the top blocker to AI adoption (IHL Group 2026)
Real-World Reference Points From 2026 Retail AI Deployments
In short
Fortnum & Mason launched six AI pilots in three months through a single implementation partner — the retail benchmark for the 90-day cadence. Tecovas built a functional AI-powered inventory app for store associates in 36 hours. VusionGroup and Hanshow started an electronic shelf label POC in 4 stores and are scaling to ~3,000 stores. IHL Group 2026 reports 47% of retail/CPG have adopted agentic AI in at least one function.
External reference points matter because compressed timelines still read as unrealistic to retail boards that have not seen them work. Three 2026 deployments worth naming:
- Fortnum & Mason launched six AI pilots in three months through a single implementation partner. This is the canonical retail benchmark for what a 90-day single-partner cadence looks like at a heritage brand — six distinct pilots in the window most retailers spend on one.
- Tecovas built a functional AI-powered inventory app for store associates in 36 hours. It is not the mode — it is a proof point that compressed build cycles are achievable when the team, data, and scope alignment are already in place.
- VusionGroup + Hanshow started an electronic shelf label POC in 4 stores and are scaling to roughly 3,000 stores. This is the canonical 90-day-to- scale pattern — a small controlled cohort producing a defensible KPI, followed by a fresh scale contract.
The IHL Group 2026 data point sits behind all three: 47% of retail and CPG organizations have adopted agentic AI in at least one function. The reason it matters is that the addressable market has moved past experimentation. Retail boards no longer need to be convinced that AI has a role — they need to be convinced that this specific pilot is the one that will scale, and the 2-4 week + 90-day cadence is how that conviction gets built.
Alice Labs' own reference set spans 100+ implementations across Nordic and pan- European retail, energy, and media clients. For direct case studies see AI implementation case studies.
Retail and CPG organizations that have adopted agentic AI in at least one function (IHL Group 2026)
EU AI Act Compliance Inside a 2-4 Week Prototype (August 2 Deadline)
In short
EU AI Act Article 16 high-risk provider obligations became applicable August 2, 2026 for systems placed on market from that date. Retail use cases that trigger high-risk classification include credit scoring for store cards, biometric identification, and employee monitoring or emotion recognition. Non-compliance penalties reach EUR 35M or 7% of global annual turnover, whichever is higher. Alice Labs is EU AI Act-native — risk classification, technical documentation, and post-market monitoring are designed in from Week 1, not retrofitted.
The regulatory backdrop is the reason European retailers are choosing EU AI Act- native partners in 2026 rather than trying to retrofit compliance at Week 12. The deadline that changed everything: August 2, 2026, when Article 16 provider obligations for high-risk AI systems became applicable for systems placed on the market from that date.
The retail use cases most likely to trigger high-risk classification are:
- Credit scoring for store credit cards — Annex III includes creditworthiness assessment as a high-risk use case; own-brand retail credit programs are in scope.
- Biometric identification in-store — camera systems that identify individuals fall under Annex III biometric categorisation and remote biometric identification categories.
- Employee monitoring and emotion recognition — Annex III includes AI used in employment context for evaluation, monitoring, or emotion recognition. Some emotion recognition uses are prohibited outright.
The penalty ceiling is the reason board-level attention has shifted: up to EUR 35M or 7% of global annual turnover, whichever is higher, for prohibited practices, and lower but still material penalties for provider obligation breaches. For a mid-cap retailer, 7% of global turnover is not a line item — it is an existential number.
Alice Labs delivers the risk classification, the technical documentation (Article 11), the human oversight design (Article 14), and the post-market monitoring plan (Article 72) in Week 1 of every engagement — not as a retrofit at Week 12. This is the operational meaning of "EU AI Act-native" and it is the single biggest differentiator in the European retail RFP process right now. For the broader compliance framework we use, see our EU AI Act compliance checklist for 2026.
Maximum EU AI Act penalty — the higher of EUR 35M or 7% of global annual turnover for prohibited practices
Transparent Pricing: What a 2-4 Week Prototype + 90-Day Pilot Costs
In short
The prototype (2-4 weeks) is a fixed-fee engagement covering discovery, data audit, working model, integration stub, and demo. The pilot (90 days) is time-and-materials or fixed monthly, scoped by store count and integration surface. Scale contracts are separate SOWs after the 90-day pilot gate passes — never bundled to avoid vendor lock-in incentives. Contract structure includes a Week 2 data-readiness kill-gate and a Week 17 pilot gate; the buyer walks with full IP and documentation at either.
Retail buyers ask for pricing structure before they ask for pricing amount. The structure is what tells them whether the incentives are aligned. Alice Labs uses the same three-phase pricing shape across every retail engagement:
- Prototype (2-4 weeks) — Fixed fee. Covers discovery, data readiness audit, working model, integration stub, and Week 4 demo. Fixed price because the scope is fixed. No expansion clauses.
- Pilot (90 days) — Fixed monthly or time-and-materials. Scoped by store count, integration surface, and KPI complexity. Priced separately from the prototype; the SOW is signed at the Week 4 gate, not before.
- Scale — Separate SOW after the pilot gate passes. Never bundled with the pilot. Bundling creates a vendor incentive to recommend scale even when the pilot underperforms, which is the exact incentive misalignment retail buyers are trying to avoid.
Two contractual clauses that matter more than the pricing amount:
- Week 2 data-readiness kill-clause. If the data readiness audit surfaces blocking issues, the engagement halts and the buyer receives the audit deliverable and full IP. No pilot obligation attaches.
- Week 17 pilot-gate kill-clause. If the 90-day pilot does not hit the pre-agreed KPI threshold, Alice Labs recommends killing the program. The buyer keeps all model artifacts, documentation, and lessons learned. No scale obligation attaches.
The transparency is deliberate. Retail procurement teams have been burned by "discovery-to-scale bundle" contracts that make it structurally expensive to walk away at Week 12. The two-kill-clause structure is what lets a retailer engage on AI at scale without introducing that lock-in risk. For broader cost benchmarking across AI programs, see our AI implementation cost benchmarks for 2026.
Scoping a retail AI pilot? Start with a 30-minute call.
Alice Labs maps your top use case against our 100+ implementation library on the scoping call, then delivers a written 2-4 week prototype scope with fixed fee within 5 business days. Senior-only team, EU AI Act-native, written kill-gate at Week 2 and Week 17.
Book a Retail AI Scoping CallThe Buyer's Checklist: How to Choose a 2-4 Week / 90-Day Retail AI Partner
In short
Five checks separate real 2-4 week / 90-day partners from marketing-timeline vendors: (1) ask for three named prior retail engagements with the same cadence and pilot-gate metrics; (2) verify the delivery team is senior-only with no junior/offshore hand-off during the prototype; (3) require a written kill-gate at Week 2 and Week 17; (4) confirm EU AI Act risk classification is delivered in Week 1, not Week 12; (5) confirm the partner is workflow-embedded (pairs with your ops team), not report-and-leave.
Retail RFPs in 2026 have converged on a checklist. Alice Labs uses the same one when we evaluate our own delivery — every question maps to a failure mode we have seen elsewhere in the market.
- 1. Three named prior retail engagements with the same cadence. Ask for the store count, the primary KPI, the KPI movement at the pilot gate, and the scale-or-kill outcome. Vendors who cannot produce three references have not run this cadence at scale.
- 2. Senior-only delivery team. No junior or offshore hand-off during the 2-4 week prototype. This is the single largest cause of compressed- timeline slippage; senior-only is not a luxury, it is what makes the timeline honest.
- 3. Written kill-gate at Week 2 and Week 17. If the contract does not include both clauses in writing, the partner is not offering a real 2-4 week / 90-day cadence — they are offering a discovery-to-scale bundle in shorter language.
- 4. EU AI Act risk classification in Week 1. Not Week 12. Ask for a sample Article 11 technical documentation template and a sample Article 72 post-market monitoring plan. If the partner cannot produce either, retrofit will be expensive.
- 5. Workflow-embedded, not report-and-leave. Confirm that one senior engineer pairs with one client operator for the full 90 days. Report-and- leave engagements deliver decks; workflow-embedded engagements deliver adoption.
All five checks are structural. None of them require the buyer to become an AI expert — they require the partner to hold contract shapes that the market has historically avoided. For the deeper vendor-selection framework we recommend, see our AI vendor selection guide and best AI implementation partners in 2026.
What Happens After the 90-Day Pilot: Scaling or Killing
In short
Scale path: expand from 5-10 pilot stores to full estate over 6-9 months with a fresh SOW and dedicated scale engineering. Kill path: 30-40% of Alice Labs pilots are killed at the 90-day gate because economics do not justify scale — the retailer keeps all IP. Production systems require monthly retraining or prompt-refresh cadence to prevent drift. Alice Labs' 100+ implementations include both scaled and killed pilots; kill rate is treated as a feature.
The honest answer to "what happens at day 90" has two paths and it matters that both are presented up-front.
The scale path. If the pilot gate passes — primary KPI moved 8-15% versus control with p<0.05, unit economics support full rollout — the engagement moves to a fresh scale SOW. The typical shape is expansion from 5-10 pilot stores to the full estate over 6-9 months, with dedicated scale engineering (integration, observability, retraining pipeline) that was intentionally not built during the pilot. This is why scale is a separate contract: the engineering shape is different, and building it into the pilot would have wasted budget on infrastructure the pilot did not need.
The kill path. If the pilot gate fails — KPI did not move, movement was not significant, or unit economics do not support scale — Alice Labs recommends killing the program. Roughly 30-40% of Alice Labs retail pilots are killed at this gate. The retailer keeps all IP, model artifacts, documentation, and lessons learned. There is no penalty for walking; the pilot SOW is a self-contained engagement.
Kill rate is treated as a feature, not a bug. A partner with a 0% kill rate is either running easy pilots or is scaling ideas that should not scale. Alice Labs' 30-40% kill rate reflects the fact that we run pilots on the actual hard problems — the ones that fail sometimes because they are worth attempting.
One production truth that gets skipped in most pilot conversations: production systems drift. Every scaled retail AI deployment requires a monthly retraining or prompt-refresh cadence to prevent quality degradation as assortments, promotions, and shopping patterns shift. Budget for this in the scale SOW, not as an afterthought at month six. For the deeper production-deployment framework, see our AI production deployment checklist.
How Alice Labs Compares to Big-4 Consultancies and Retail AI Product Vendors
In short
Big-4 consultancies (McKinsey, Deloitte, Accenture, BCG) typically spend 8-16 weeks on discovery before writing production code. Retail AI product vendors are fast to demo but slow to fit unique data and assortment. Regional systems integrators offer variable seniority. Alice Labs starts building in Week 1 because the 100+ prior implementations library shortcuts discovery, guarantees senior-only delivery, is EU AI Act-native from Week 1, and is single-partner accountable through the pilot gate.
The competitive shape is worth naming explicitly because retail buyers routinely issue RFPs to all four categories and need to compare apples to apples.
Retail AI partner categories compared
| Partner type | Prototype start | Delivery seniority | EU AI Act |
|---|---|---|---|
| Big-4 (McKinsey, Deloitte, Accenture, BCG) | Week 8-16 after discovery | Senior partners lead; execution via junior/offshore pods | Compliance advisory workstream |
| Retail AI product vendors | Fast demo, slow to fit | Product engineers; not retail-domain | Product-level, not deployment-level |
| Regional systems integrators | Variable | Mixed; junior heavy in most engagements | Often retrofit |
| Alice Labs | Week 1 | Senior-only through the pilot gate | Native — designed in Week 1 |
Two structural differences matter most. First, the 100+ prior implementations library is what lets Alice Labs skip 8-16 weeks of discovery — pattern matching replaces workshops, so the working model starts Week 1. Second, single-partner accountability through the pilot gate keeps the incentive to fix problems at Week 8 aligned with the incentive to hit the pilot gate at Week 17 — multi-vendor pods structurally cannot offer that.
Alice Labs is Stockholm-headquartered with international delivery, EU AI Act- native, and senior-only. For the broader partner-comparison landscape, see best AI implementation partners by industry for 2026.
Measurement: The KPIs That Matter in the 90-Day Retail Pilot
In short
Four KPI layers: (1) Primary business KPIs — incremental margin, stockout %, shrink %, basket size, GMROI; (2) Secondary model KPIs — precision/recall, hallucination rate for LLM systems, p95 latency; (3) Adoption KPIs — associate usage rate, override rate, task-completion time delta; (4) Guardrail KPIs — false-positive customer impact, EU AI Act log completeness, data drift alerts. The pilot gate is judged on the primary KPI only; the other layers exist to catch failure modes early.
The KPI taxonomy is the deliverable retail leaders need to defend the pilot internally. A single KPI is not enough — pilots that report only the primary business metric miss the failure modes that surface in the model, adoption, and guardrail layers.
- Primary business KPIs. The number the CFO cares about. Incremental margin, stockout percentage, shrink percentage, basket size, GMROI. The pilot gate is judged on this layer only — one KPI, pre-agreed, measured versus control.
- Secondary model KPIs. Precision, recall, F1 for classification and CV; hallucination rate and factual-accuracy rate for LLM-driven systems; p95 latency for anything customer-facing. This layer catches "the model got worse" before it shows up in the business metric.
- Adoption KPIs. Associate usage rate, override rate, task- completion time delta. Adoption is the leading indicator for scale: if associates are not using the tool or overriding its recommendations, the business KPI movement is fragile.
- Guardrail KPIs. False-positive customer impact, EU AI Act log completeness, data drift alerts. These do not decide the pilot gate — they decide whether the pilot is safe to run at all. A guardrail breach ends the pilot regardless of business KPI movement.
The four-layer structure is what separates a pilot that scales cleanly from one that scales into a compliance incident. Alice Labs bakes all four into the Week 1 measurement plan, and the day-90 report presents all four with the primary layer in the headline. For deeper coverage see our AI measurement framework.
When 2-4 Weeks Is the Wrong Answer
In short
Four scopes do not fit 2-4 weeks: multi-country ERP re-platforming (needs 6-18 months), enterprise-wide data foundation greenfield (prerequisite to any AI), fully autonomous agent replacing a regulated workflow (needs longer validation under EU AI Act), and any use case where the primary business KPI cannot be measured within 90 days. Alice Labs will decline these scopes rather than accept an unrealistic frame.
The honesty section. Compressed timelines are not a universal answer, and Alice Labs will decline scopes that do not fit rather than accept an unrealistic frame. The four scopes we routinely turn down or reshape:
- Multi-country ERP re-platforming. SAP or Oracle Retail replacements are 6-18 month programs at minimum. AI can accelerate parts of them (data migration, master data cleaning, testing) but the ERP program itself does not fit a 4-week envelope. Alice Labs will build AI accelerators inside the program, not replace it.
- Enterprise-wide data foundation greenfield. If the retailer does not have unified SKU, store, customer, or transaction data, that is the prerequisite to any AI — not something AI can compress. We recommend a data- foundation phase first, and revisit AI use cases once the foundation is in place.
- Fully autonomous agent replacing a regulated workflow. Credit-decision agents, safety-critical monitoring agents, and any AI in the scope of Annex III of the EU AI Act need longer validation cycles than 90 days. The regulatory shape does not compress.
- KPI that cannot be measured in 90 days. Some KPIs — customer lifetime value, brand equity, multi-year loyalty impact — need longer measurement windows. Pilots against these KPIs fail the significance gate at day 90 even when the underlying idea works. We recommend proxy KPIs or a longer pilot shape.
Declining wrong-fit scopes is the single largest trust-building action we take in an RFP process. Retail buyers have been over-promised for three years; being the partner that says "this specific scope needs a different shape" is what earns the engagement on the scopes that do fit.
Next Steps: Booking a Prototype Scoping Call With Alice Labs
In short
A 30-minute scoping call maps your top use case against Alice Labs' 100+ implementation library. A written 2-4 week prototype scope with fixed fee is delivered within 5 business days. The contract structure includes a kill-gate at Week 2 (data readiness) and Week 17 (pilot gate), senior-only team, and EU AI Act compliance from Week 1. Book at alicelabs.ai/contact or email hello@alicelabs.ai.
The engagement starts with a 30-minute scoping call. Bring your top use case, the business KPI you want to move, and a rough sense of the data available. Alice Labs maps the use case against the 100+ prior implementation library on the call — the "we have done this three times before" or "this is a shape we have not seen" answer is honest either way.
Within 5 business days of the scoping call, Alice Labs delivers a written 2-4 week prototype scope with fixed fee, three acceptance metrics, and the two-kill-clause contract structure. No obligation attaches until the SOW is signed.
What you get on the contract:
- Senior-only delivery team through the pilot gate.
- EU AI Act risk classification and technical documentation delivered in Week 1.
- Written kill-gate at Week 2 and Week 17 — walk away with full IP at either.
- Single-partner accountability through the 90-day pilot gate.
- Fixed-fee prototype, separately scoped pilot, separately scoped scale — no bundled discovery-to-scale contract.
Book at alicelabs.ai/contact or email hello@alicelabs.ai. For the broader Alice Labs implementation methodology, see the AI implementation pillar.
About the Authors & Reviewers

Co-Founder, Alice Labs
Co-Founder at Alice Labs. Builds AI automation, agent workflows and integration systems that hold up in real business operations.
- AI automation & agent systems lead
- Workflow design across 100+ deployments
- Specialist in RAG, integrations & APIs

Co-Founder, Alice Labs
Co-Founder at Alice Labs. Author of 7 research reports on AI adoption, governance and labor markets cited across EU, OECD and US benchmarks.
- 8+ years in AI strategy & implementation
- Top-5 AI Speaker, Sweden (Mindley 2025)
- 100+ enterprise AI engagements
Frequently Asked Questions
How fast can Alice Labs deliver a working AI prototype for a retailer?
Alice Labs delivers working retail AI prototypes in 2-4 weeks for scoped, single-use-case problems such as shelf monitoring, demand forecasting at SKU/store level, or a store-associate copilot. Week 1 is a data readiness audit and scope lock. Weeks 2-3 build the working model and integrate it into your POS, PIM or CV stack. Week 4 delivers a live demo against three pre-agreed acceptance metrics. This is the same cadence Alice Labs has used across 100+ enterprise AI implementations since 2023.
What is the difference between a 2-4 week prototype and a 90-day pilot?
A prototype is a working, single-use-case system running in a sandbox or one to two stores, built in 2-4 weeks and judged on functional acceptance criteria. A 90-day pilot deploys that prototype into a controlled production cohort of 3-10 stores or one channel segment, measured against a control group on business KPIs like stockouts, basket size or margin. Alice Labs runs these as two separately gated phases so retailers can kill weak ideas at Week 4 rather than Week 17.
Why do 89% of AI pilots fail to reach production in retail?
Gartner and IDC report that around 89% of AI agent pilots never reach production. The dominant causes are data readiness gaps (38% of retailers cite this as their top blocker per IHL Group 2026), scope creep during the pilot, missing change management, and vendor sprawl. Alice Labs prevents these by running a Week 1 data readiness kill-gate, locking scope contractually to one primary KPI, embedding one senior engineer alongside one client operator, and staying single-partner accountable through the pilot gate.
What retail use cases fit inside a 2-4 week prototype window?
Alice Labs typically fits these use cases inside 2-4 weeks: computer vision shelf monitoring and planogram compliance, SKU-store-level demand forecasting, assortment-aware product recommendations, store associate copilots for inventory and product Q&A, and self-checkout loss-prevention CV. Enterprise ERP re-platforming, greenfield data foundations, and fully autonomous regulated workflows cannot be compressed into 2-4 weeks, and Alice Labs will decline those scopes rather than accept an unrealistic timeline.
How does the Alice Labs 90-day retail pilot measure success?
Alice Labs uses four KPI layers across the 90-day pilot. Primary business KPIs include incremental margin, stockout percentage, shrink percentage, basket size, and GMROI. Secondary model KPIs cover precision, recall, hallucination rate for LLM systems, and p95 latency. Adoption KPIs track associate usage, override rate, and task-completion time. Guardrail KPIs monitor false-positive customer impact, EU AI Act log completeness, and data drift. The pilot gate requires the primary KPI to move 8-15% versus control with p<0.05.
Is Alice Labs EU AI Act compliant for retail AI deployments?
Yes. Alice Labs is EU AI Act-native. Article 16 obligations for high-risk AI system providers became applicable on August 2, 2026. Alice Labs delivers the risk classification, technical documentation, human oversight design, and post-market monitoring plan in Week 1 of every engagement, not as a retrofit. Retail use cases that trigger high-risk classification include store card credit scoring, biometric identification, and employee monitoring. Non-compliance penalties reach EUR 35M or 7% of global annual turnover.
What does a 2-4 week prototype cost with Alice Labs?
Alice Labs uses transparent, fixed-fee pricing for the 2-4 week prototype phase covering discovery, data readiness audit, working model, integration stub, and live demo. The 90-day pilot is scoped separately as either fixed monthly or time-and-materials, priced by store count and integration surface. Scale contracts are never bundled with the pilot to avoid vendor lock-in incentives. Retailers can walk at the Week 2 kill-gate or Week 17 pilot gate with full IP and documentation. Contact Alice Labs for a written scope.
Who is on the Alice Labs delivery team during the 2-4 week prototype?
Alice Labs is senior-only. Every 2-4 week prototype is delivered by senior engineers and applied AI leads; there is no junior hand-off or offshore staffing. The engagement is co-founded by Eric Lundberg and Linus Ingemarsson under CEO Alice Holmgren, and pairs one senior Alice Labs engineer with one client operator for the duration of both phases. This is why Alice Labs can compress timelines that big-4 consultancies typically spread across 8-16 weeks of discovery.
How does Alice Labs compare to McKinsey or Accenture for retail AI pilots?
Big-4 consultancies (McKinsey, Deloitte, Accenture, BCG) typically run 8-16 weeks of discovery before writing production code. Alice Labs starts building in Week 1 because the 100+ prior implementations library kills 60-70% of use cases in the triage phase, not through workshops. Alice Labs is also EU AI Act-native from Week 1 and works as a single-partner accountable through the pilot gate, versus multi-vendor big-4 pods that hand off to system integrators.
What happens if the 90-day retail pilot fails to hit its KPIs?
The Alice Labs pilot contract includes a written 90-day pilot gate. If the primary business KPI has not moved 8-15% versus control with statistical significance, Alice Labs recommends killing the pilot. Roughly 30-40% of Alice Labs retail pilots are killed at this gate because unit economics do not justify scale. The retailer keeps all IP, documentation, model artifacts, and lessons learned. Kill rate is treated as a feature, not a bug, because it is what separates a real 90-day gate from vendor theatre.
Can Alice Labs run a 2-4 week prototype outside Sweden and the Nordics?
Yes. Alice Labs is Stockholm-headquartered but delivers internationally. The 2-4 week prototype and 90-day pilot cadence has been run across Nordic and pan-European retailers, and the engagement shape (senior-only, embedded operator, kill-gate contract) is not geography-dependent. EU AI Act coverage is intrinsic to every deployment inside the EU; equivalent risk-management coverage is applied for deployments outside the EU where the retailer requests it.
How many stores should be in the 90-day retail pilot cohort?
The typical Alice Labs pilot cohort is 3-10 stores, with a matched set of control stores of similar size, format, and geography. The lower bound (3) is what statistical significance requires for most retail KPIs at 90 days; the upper bound (10) keeps the operational burden low enough that iterations can happen weekly. For digital-channel pilots, the equivalent is one channel segment (a category, a region, or a customer cohort) with a matched control segment.
Do we need to have clean data before Alice Labs can start the prototype?
Data does not need to be perfect, but it needs to be workable. Alice Labs runs a Week 1 data readiness audit as the first deliverable — if the audit surfaces blocking issues (missing SKU-store history, unlabelled camera footage, disconnected POS and PIM systems), the engagement halts at the Week 2 kill-gate and the retailer receives the audit deliverable with full IP transfer. If the audit passes, the prototype build proceeds. Fixing critical data issues before the prototype starts shortens the timeline; waiting to fix everything before starting is what kills momentum.
What is the difference between a prototype and a proof-of-concept (POC)?
In Alice Labs' language, a prototype is a working system integrated into a real environment (sandbox or 1-2 stores) with real data and a real acceptance metric — the goal is to make a scale-or-kill decision at Week 4. A POC in the traditional sense is often demonstrative-only, sometimes on synthetic data, without a production integration path. The 2-4 week prototype cadence exists specifically to skip the demonstrative POC phase and go directly to a decision-grade artifact.
How does Alice Labs prevent scope creep during the 2-4 week prototype?
Scope is locked contractually in Week 1 to one primary KPI and one use case. Secondary requests that surface during the prototype are logged and become the subject of the next SOW, not this one. The one-KPI discipline is the single largest structural defence against scope creep — it converts a mid-project stakeholder ask from 'add this now' into 'add this to the next engagement', which is a much easier decision to hold the line on.
What integrations are typically involved in a retail AI prototype in Weeks 2-3?
The integration surface depends on the use case. Shelf monitoring pilots integrate with the store's camera stack (RTSP feeds or edge devices) and inventory system. Demand forecasting pilots integrate with the POS system for transactions and the PIM for SKU master data. Store associate copilots integrate with the OMS for inventory and the product catalogue. Self-checkout CV integrates with the SCO software directly. Alice Labs builds integration stubs in Weeks 2-3; production-grade integration engineering happens in the scale SOW, not the prototype.
What happens after the 90-day pilot succeeds and moves to scale?
The scale phase is a separate SOW signed at the Week 17 gate — never bundled with the pilot. Typical shape: expansion from 5-10 pilot stores to the full estate over 6-9 months, with dedicated scale engineering (production integration, observability, retraining pipeline) that was intentionally not built during the pilot. Production systems require monthly retraining or prompt-refresh cadence to prevent drift, and the scale SOW budgets for that from month one rather than treating it as an afterthought.
How do I know if my retail AI project needs a longer engagement than 2-4 weeks?
Ask two questions. First: does the use case require a data foundation that does not yet exist (unified SKU/store data, customer identity resolution)? If yes, the data foundation is a prerequisite. Second: is the use case scope of Annex III of the EU AI Act (credit scoring, biometric ID, employee monitoring)? If yes, the validation cycle is longer than 90 days by regulation. Alice Labs' Week 0 triage answers both questions before any SOW is signed — an honest 'this needs a different shape' recommendation is part of the engagement.
How does Alice Labs handle the change management side of the 90-day pilot?
Change management is embedded from day one via the one-senior-engineer-plus-one-client-operator pairing. The client operator (store manager, category manager, or head of stores) owns the adoption metric on the pilot cohort, surfaces edge cases as they arise, and is the internal advocate for the tool. This is deliberately different from the report-and-leave model where the vendor delivers a deck at day 90 and the client is left to drive adoption alone.
Can Alice Labs run multiple retail AI pilots in parallel like Fortnum & Mason did?
Yes. Fortnum & Mason launched six AI pilots in three months through a single implementation partner — the canonical retail benchmark for the 90-day cadence at portfolio scale. Alice Labs runs multi-pilot programs with a shared measurement framework, shared EU AI Act documentation, and independent kill-gates per pilot. The advantage of running pilots in parallel is that failures on one pilot inform the design of others in real time, rather than being repeated sequentially over the following year.
AI Implementation Services 2026: Complete Guide | Alice Labs
Further reading
- Quinnox — AI for Rapid Prototyping· quinnox.com
- IHL Group — Gartner: 40% of Agentic AI Projects Will Be Canceled by 2027· ihlservices.com
- Truvisory — 90-Day AI Sprint· truvisory.com
- Holland & Knight — EU AI Act August 2026 Compliance Deadline· hklaw.com
- McKinsey — Rewiring Retail in Europe: The AI Imperative· mckinsey.com
- Beri — 2026 Gartner/IDC Enterprise AI Agent Adoption· beri.net
Related reading
AI Implementation for Enterprise: The 2026 Playbook
The parent pillar covering the full AI implementation lifecycle — from use-case selection through production deployment across enterprise.
deepdiveAI Implementation Roadmap for Enterprise
The end-to-end roadmap Alice Labs uses across 100+ deployments — beyond the 2-4 week / 90-day cadence into long-horizon programs.
deepdiveAI Production Deployment Checklist
What Alice Labs checks before promoting a scaled retail AI system into production, and the retraining cadence required to prevent drift.
deepdiveAI Implementation Cost Benchmarks 2026
Fixed-fee prototype pricing, pilot pricing, and scale contract shapes benchmarked across the 2026 retail AI market.
deepdiveEU AI Act Compliance Checklist for 2026
The August 2, 2026 Article 16 provider obligations broken down by use case, with the technical documentation and post-market monitoring templates.
deepdiveAI Failure Modes
The taxonomy of AI program failure modes and the structural fixes for each — the reference behind the 30-40% Alice Labs pilot kill rate.
deepdiveAI Measurement Framework
The four-layer KPI framework (business, model, adoption, guardrail) Alice Labs uses across every 90-day pilot to make the scale-or-kill decision.
Sources
- AI for Rapid PrototypingQuinnox · Quinnox“The industry-standard retail cadence is a 2-4 week working prototype followed by a 60-90 day pilot in a controlled cohort, with scale decisions made at the pilot gate rather than pre-committed.”(accessed 2026-08-04)
- Gartner Predicts 40% of Agentic AI Projects Will Be Canceled by 2027 — Retail Store-Level ViewIHL Group · IHL Group“Gartner predicts 40% of agentic AI projects will be canceled by end of 2027 due to unclear ROI and lack of controls. IHL Group's 2026 survey found 47% of retail and CPG organizations have adopted agentic AI in at least one function.”(accessed 2026-08-04)
- 90-Day AI SprintTruvisory · Truvisory“The 90-day AI sprint structure splits into three 30-day segments: rollout and baseline (days 1-30), weekly iteration with human-in-the-loop refinement (days 31-60), and statistical significance check with unit economics modelling (days 61-90).”(accessed 2026-08-04)
- Computer Vision in RetailGlorium Tech · Glorium Tech“Retail computer vision use cases — shelf monitoring, planogram compliance, out-of-stock detection, and self-checkout loss prevention — fit inside 3-4 week prototype windows because the decision surface is well-scoped and camera data is already in place.”(accessed 2026-08-04)
- AI Agent Adoption in the Enterprise (Gartner / IDC 2026)Beri · Beri“Gartner and IDC 2026 data show approximately 89% of AI agent pilots never reach production. Root causes: data readiness gaps, scope creep, change-management vacuum, and vendor sprawl.”(accessed 2026-08-04)
- The Three Technology Pillars Reshaping Retail in 2026IHL Group · IHL Group“38% of retailers cite data cleaning, training, and storage as their top blocker to AI adoption in 2026 — the dominant reason compressed-timeline pilots fail without a Week 1 data readiness audit.”(accessed 2026-08-04)
- US Companies Face EU AI Act's Possible August 2026 Compliance DeadlineHolland & Knight · Holland & Knight LLP“EU AI Act Article 16 high-risk provider obligations became applicable August 2, 2026 for systems placed on market from that date. Non-compliance penalties reach EUR 35M or 7% of global annual turnover, whichever is higher.”(accessed 2026-08-04)
- Rewiring Retail in Europe: The AI ImperativeMcKinsey & Company · McKinsey & Company“McKinsey 2026 finds 62% of enterprises remain in the experimentation or pilot stage of AI adoption; only 23% have scaled AI agents to production. The scale gap is the primary driver behind compressed-timeline retail engagements in 2026.”(accessed 2026-08-04)
- Retail Innovation in 2026: AI Adoption Themesdunnhumby · dunnhumby“Retail AI pilots that reach scale share a common measurement structure — a primary business KPI (margin, stockout, basket size, GMROI), secondary model KPIs, adoption KPIs, and guardrail KPIs — measured against a matched control cohort.”(accessed 2026-08-04)
- Why AI Pilots Fail to Reach Production in RetailTechverx · Techverx“Retail AI pilots most commonly fail when scope exceeds the compressed-timeline envelope — multi-country ERP re-platforming, greenfield data foundations, and fully autonomous regulated workflows cannot be delivered in a 2-4 week prototype window.”(accessed 2026-08-04)
- AI Pilot to ProductionOlakai · Olakai“Post-pilot production systems require monthly retraining or prompt-refresh cadence to prevent drift. The pilot-to-production transition is a separate engineering shape from the pilot itself and should be budgeted as its own SOW.”(accessed 2026-08-04)
- Alice Labs AI Consulting — 100+ Enterprise ImplementationsAlice Labs · Alice Labs“Alice Labs has delivered 100+ production AI implementations since 2023 across Nordic and pan-European enterprise clients — including retail prototype and 90-day pilot programs with the two-kill-clause contract structure.”(accessed 2026-08-04)
Next scheduled review: