BazBiff

GUIDES

How to Choose Your First High-Value AI Use Case

71% of CPG leaders now use AI yet most pilots stall. Here is the 3-filter framework FMCG brands use to choose where to start and avoid expensive dead ends.

26 Jan 202620 min readBy BazBiff Team

Updated June 2026

Here's the stat that should embarrass the industry: 71% of CPG companies now use AI in at least one function, according to McKinsey's 2025 State of AI report. Yet only 23% are scaling AI agents in even one area. That means three quarters of FMCG brands are running experiments that never graduate. The pilots stall, the case studies stay internal, and the spend sits on the balance sheet without a return.

The technology isn't the problem. The tools have matured. The cost has dropped sharply. In practice, brands pick the wrong starting point — an application that's too complex, too data-hungry, or too far removed from a decision someone actually needs to make. Your first AI initiative sets the template for everything that follows. Choosing it well is a strategic call, not a technical one.

This article gives you a practical 3-filter framework for selecting your first high-value AI use case, identifies the four FMCG applications that consistently pass all three filters, and shows you what a 90-day pilot looks like in practice.

The Bottom Line - 71% of CPG companies use AI, but only 23% are scaling — that gap is a selection problem, not a technology problem (McKinsey, 2025). - Apply three filters before you commit: well-defined process, available data, measurable outcome within 90 days. - Demand forecasting, quality inspection, workflow automation, and promo reporting automation consistently pass all three. - Brands that skip workflow redesign see 4x higher AI failure rates (BCG, 2026).
Decision gate showing stalled versus scaled AI pilots
Decision gate showing stalled versus scaled AI pilots

Why Do Most FMCG AI Pilots Stall?

Most FMCG AI pilots stall because the use case was chosen for the wrong reasons — not because the technology failed. The most common culprits are excitement ("everyone's doing demand forecasting"), seniority ("the MD wants to see AI in the warehouse"), and vendor momentum ("the software company already has a demo ready"). None of those are bad conversation starters, but none of them tell you whether the application will actually work in your organisation, with your data, on your timeline.

McKinsey's 2025 research makes the scale of the problem concrete. The same report that found 71% of CPG companies using AI also found that scaling from one function to multiple functions remains the dominant challenge. Brands that have moved beyond the pilot stage have one thing in common: their first initiative was chosen deliberately, with explicit criteria, not opportunistically.

BCG's 2026 research on AI agents in supply chains adds a structural dimension worth paying attention to. Consumer goods brands that don't redesign their workflows around AI — treating it as a bolt-on rather than a reason to rethink the process — see failure rates four times higher than brands that treat AI implementation as process change first and technology second.

The failure isn't random. It's predictable, and it's avoidable if you apply the right filter before you start.

What Is the 3-Filter Framework for Choosing an AI Use Case?

This framework doesn't rank use cases by potential value, because potential is almost impossible to assess before you have any traction. It ranks them by the probability of actually delivering value within 90 days. An application that delivers 30% of its potential but is live, measurable, and building internal confidence is worth more than one promising 10x returns that's still in a steering committee six months later.

Apply all three filters in sequence. A use case needs to pass all three to qualify as a viable first proof of value.

Filter 1: Is the Process Well-Defined?

AI doesn't invent process. It accelerates and improves processes that already exist. If you can't document the current process in a one-page flowchart — inputs, steps, decision points, outputs — you're not ready to apply AI to it yet. You need to sort the process first.

A well-defined process has:

  • A clear trigger (what starts it)
  • Defined inputs (what data or materials enter the process)
  • A documented set of steps that are followed consistently, even if manually
  • A known output with an accepted quality standard
  • An owner who can describe the current failure modes

Demand forecasting passes this test easily. Inputs are sales history and promotional plans, the step is generating a volume projection, and the output is a weekly or monthly forecast used for procurement. Promo reporting passes it too: the inputs are sales-out data and promotional spend, the output is a reconciled report, and someone on your team produces it manually every week. A vague ambition to "use AI for customer insights" doesn't pass. No defined process, no consistent input, no agreed output.

Worth noting: The fastest way to assess Filter 1 is to ask the person who currently runs the process to walk you through it in real time. If it takes more than 20 minutes and requires several "it depends" qualifications, it's not well-defined enough to automate. That isn't a reason to abandon AI. It's a reason to standardise the process first — which itself creates value before a line of AI code is written.

Filter 2: Is the Data Available and Clean?

AI is a pattern-recognition engine. Without data that reflects the pattern you're trying to learn, it has nothing to work with. This filter is the one most brands underestimate, because data problems are invisible until you actually try to use the data.

A data source passes this filter if:

  • It's accessible without significant manual effort
  • It covers at least 12–24 months of history (for time-series applications like forecasting)
  • It's structured well enough to be queried or exported without a rebuild
  • The quality is good enough to establish a reliable baseline — not perfect, but not so noisy that a human couldn't use it either

You don't need perfect data to start. Every AI project improves data quality as a side effect, because you discover and fix problems during implementation. What you do need is data that's good enough to train a model and measure its performance against a baseline.

Put plainly: if your sales data lives in three different formats across two systems and a spreadsheet, you've got a data problem that needs solving before an AI pilot. Solving it is itself a valuable project. See our guide on why spreadsheet reporting becomes a growth bottleneck for a framework on tackling the data infrastructure problem first.

Comparison table for data-ready and data-not-ready AI use cases
Comparison table for data-ready and data-not-ready AI use cases

Filter 3: Is There a Measurable Outcome Within 90 Days?

The 90-day constraint isn't arbitrary. It's the maximum window in which a pilot can maintain organisational momentum, budget approval, and stakeholder attention without a visible result. Beyond 90 days, proofs of value accumulate debt: political capital spent, team bandwidth consumed, and growing suspicion that it's a research project rather than a business initiative.

A measurable outcome within 90 days requires:

  • A baseline metric you can measure today, before the AI is in place
  • A target that's numerically specific — not "improve accuracy" but "reduce forecast error from 28% to below 18%"
  • A data capture mechanism that shows progress week by week, not just at the end

Demand forecasting has this. You measure mean absolute percentage error (MAPE) today, run the AI model in parallel for four weeks, and compare. Promo reporting automation has this: you measure hours spent per week building the report now, and measure again after automation. Quality inspection has this: you measure defect escape rate and inspection throughput today, then track both after deployment.

A use case that can't be measured in 90 days is either too complex for a first initiative, or not important enough to measure. Both are reasons to choose something else first.

Which FMCG Applications Pass All Three Filters?

These aren't the only FMCG AI applications that work. They're the ones that consistently satisfy all three filters across different categories, sizes, and data maturity levels — which is why they're the right place to start.

Complexity versus speed grid for selecting the first AI use case
Complexity versus speed grid for selecting the first AI use case

1. Demand Forecasting

Demand forecasting is the highest-frequency AI application in FMCG for a straightforward reason: the problem's well-defined, the data exists, and the outcome is directly measurable. McKinsey's research on AI-driven demand forecasting found it reduces forecast errors by 20–50% — a range that reflects how much room for improvement exists in most organisations still running spreadsheet-based forecasts.

P&G and Unilever have been doing this at scale for years, with dedicated data science teams and proprietary models. The equivalent at your scale looks different. You're not building a bespoke model from scratch. You're connecting 18–24 months of SKU-retailer sales history to a forecasting tool that accounts for promotional calendars and seasonality — and you're comparing its output against your current manual forecast over a 4-week parallel run.

The AI ingests historical sales data, promotional uplift factors, and seasonal patterns, then produces a volume forecast that outperforms a spreadsheet model by accounting for non-linear interactions between variables. A promotion during a school holiday performs differently from the same promotion in February. The AI learns this. A formula doesn't.

90-day pilot structure: Establish MAPE baseline in week one, run the AI forecast in parallel through weeks two to eight, compare outputs against actuals, and make the go/no-go decision in week twelve with two months of evidence.

2. Quality Inspection

Computer vision for quality inspection passes all three filters cleanly. The process is well-defined (compare product or packaging against specification), the data is images or sensor readings already being generated on your production line, and the outcome is measurable in defect escape rate, inspection throughput, and labour hours per unit inspected.

The business case compounds in two directions: fewer defects reaching retail reduces returns, complaints, and brand damage; faster inspection throughput reduces the production bottleneck and labour cost. For food and beverage brands with compliance requirements, the audit trail created by AI inspection is an additional benefit.

The data requirement is lower than most teams expect. A few thousand labelled images covering your main defect categories is typically enough to train an initial model with useful accuracy. The model improves as more examples are labelled during deployment.

3. Workflow Automation

Not all workflow automation is AI — and the distinction matters for choosing the right tool. Understanding the difference between AI and automation in an FMCG context helps you avoid deploying the wrong solution. Rule-based automation handles predictable, structured tasks. AI-powered automation handles tasks where the inputs vary, exceptions are common, and context changes the right answer.

For FMCG operations, the highest-value AI workflow applications are:

  • Purchase order processing: extracting and validating order data from varied supplier formats, flagging anomalies, and routing exceptions without manual intervention
  • Customer query triage: classifying incoming retailer or distributor queries by type and urgency, routing to the right team, and surfacing relevant account history
  • Promotional compliance checking: comparing promotional execution data against agreed terms and flagging discrepancies before the invoice dispute stage

All three have a well-defined current process (even if it's manual), data that exists somewhere, and outcomes measurable in processing time, error rate, or cost per transaction.

4. Promo Reporting Automation

Promo reporting automation is the most accessible first AI initiative for brands that haven't yet addressed their data infrastructure. If your team's currently spending hours each week assembling the same report from multiple sources — sales-out, promotional spend, stock levels, margin — automating that assembly and generation passes all three filters with minimal implementation risk.

The process is well-defined because it's being done manually right now, so it can be documented. The data is available because someone's pulling it from somewhere every week. The outcome is measurable in hours saved and report frequency. Most brands find that automated promo reporting enables daily or real-time visibility that was previously only possible weekly.

This is the starting point we come back to most often, because it builds the data infrastructure that makes more ambitious AI applications feasible later. Automated reporting isn't the destination. It's the foundation. For more on why manual reporting creates structural constraints, see our piece on why spreadsheet reporting becomes a growth bottleneck.

What Does a 90-Day Pilot Structure Look Like?

A 90-day pilot isn't a proof of concept. It's a structured business experiment with a clear hypothesis, a measurement framework, and a pre-agreed decision point. That distinction matters because it changes how you resource it, how you communicate it internally, and what happens at day 90.

Ninety-day AI pilot timeline from baseline to decision
Ninety-day AI pilot timeline from baseline to decision

Month 1: Baseline and Data Validation

The goal of month one isn't to build anything. It's to establish the ground truth against which the AI's performance will be measured.

Weeks 1–2: Document the current process in detail. Measure the baseline metric. Identify the data sources and assess their quality. Appoint a process owner who'll own the outcome at day 90.

Weeks 3–4: Clean and validate the data. Fix the most significant quality issues. Define the success threshold — the specific metric value at which you'll consider the pilot successful and proceed to scale.

A common mistake is skipping the baseline measurement and trying to reconstruct it retrospectively. That makes the evaluation at day 90 almost impossible to defend to a sceptical stakeholder. Measure now, even if the current performance is embarrassing.

Month 2: Parallel Run

Run the AI model alongside the existing process, not instead of it. The parallel run serves two purposes: it generates a comparison dataset (AI output vs. human output vs. actual outcome), and it gives the team time to build confidence in the model's behaviour before depending on it.

Weeks 5–8: Deploy the AI model in shadow mode. Continue the existing process in parallel. Log every AI prediction or output alongside the human-generated equivalent.

Weeks 9–10: Review the comparison data. Identify the categories of error the AI is making. Adjust configuration or training data to address the most significant gaps. Start transitioning lower-risk decisions to AI-generated outputs.

The parallel run is also when workflow redesign happens. BCG's 2026 research found that brands failing to redesign workflows see 4x higher failure rates — and that finding applies directly here. The question isn't just whether the AI is accurate. It's whether your team's daily workflow has been updated to act on the AI output, rather than generating the manual version first and checking it against the AI second.

Month 3: Evaluation and Decision

Weeks 11–12: Compare AI performance against the baseline from month one. Measure the primary outcome metric and at least two secondary metrics (cost, time, error rate). Produce a one-page evaluation brief that states: what the baseline was, what the AI achieved, what the gap is relative to the target, and what would need to change to close it.

The decision at day 90 has three possible outcomes:

  1. Scale: Performance meets or exceeds the target. Proceed to full deployment and expand scope.
  2. Extend: Performance is below target but the trajectory's positive and the root cause is identified. Extend by 30 days with a specific fix in place.
  3. Stop: Performance is below target with no clear path to improvement. Document what you learned, apply it to the next use case selection, and run the 3-filter process again.

All three outcomes are legitimate. A well-structured pilot that stops at day 90 has generated real value — in data understanding, process documentation, and a sharper view of what the next initiative should be.

Why Does Workflow Redesign Determine AI Success?

The BCG finding is the most consistently underestimated factor in FMCG AI implementation. It's worth dwelling on.

AI doesn't deliver value by existing. It delivers value when the people and processes around it are redesigned to act on its output. A demand forecast that's more accurate than the spreadsheet model, but still being overridden by the same gut-feel adjustments as before, hasn't changed the business. Similarly, a quality inspection system that flags defects but routes them into the same manual review queue as before hasn't changed throughput.

The workflow redesign question to ask at the start of every pilot — not the end — is: "If the AI is right 90% of the time, what decisions will we make differently, and who currently needs to approve those decisions?" If the answer requires a policy change, an approval workflow change, or a shift in who owns a particular decision, those changes need to be scoped and planned before the pilot starts. Discovering them as blockers after the AI is already performing is a costly way to learn.

That's also why the process owner appointed in month one needs decision-making authority, not just technical knowledge. Their job isn't to evaluate the AI. It's to change the process to use it.

For more on how to approach the transition from manual to AI-assisted workflows without creating parallel process overhead, see our guide on moving from manual processes to automated workflows.

How Do You Choose Between the Four Use Cases?

If you've read this far and you're still unsure which application is right for your organisation first, use this shortcut.

Start with promo reporting automation if: Your team spends more than three hours per week assembling the same reports. Your data exists but sits in multiple places. You haven't yet built a unified data layer.

Start with demand forecasting if: You've got at least 18 months of sales history at SKU-retailer level. A commercial or supply chain team makes procurement decisions based on volume projections. Your forecast error is above 20% MAPE.

Start with quality inspection if: You've got a production line generating images or sensor data that isn't currently being used. You have a documented defect specification. Your quality team is spending significant time on manual inspection tasks.

Start with workflow automation if: You've got a high-volume, repetitive process that currently requires manual handling of variable inputs — orders, queries, documents. You can measure the current process in transactions per hour or cost per transaction.

When you're unsure, start with promo reporting automation. It's the lowest-risk first pilot, it builds the data infrastructure that makes the other three more tractable, and it produces visible value fast enough to maintain stakeholder confidence for the next stage.

Decision tree for choosing a first high-value AI use case
Decision tree for choosing a first high-value AI use case

Frequently Asked Questions

Why do most FMCG AI pilots fail?

Most FMCG AI pilots fail because they target the wrong starting point. Teams either pick a problem that's too complex — undefined process, messy data, no clear metric — or something so small it delivers no visible value. McKinsey's 2025 State of AI report found that while 71% of CPG companies use AI in at least one function, only 23% are scaling AI agents in even one area. The gap between experimenting and delivering is almost entirely a selection problem, not a technology problem.

What is the best first AI use case for an FMCG brand?

The best first AI use case is whichever one passes all three filters: a well-defined process, available and reasonably clean data, and a measurable outcome within 90 days. For most consumer goods brands, demand forecasting, promo reporting automation, quality inspection, and workflow automation consistently pass all three filters and generate returns visible enough to secure the next phase of investment.

How long should an AI pilot last?

Ninety days is the right ceiling for a first proof of value. Month one focuses on baseline measurement and data validation. Month two runs the AI model in parallel with the existing process. Month three evaluates performance against the baseline and makes the go/no-go decision. Pilots that run longer than 90 days without a measurable checkpoint lose stakeholder confidence and rarely recover momentum.

How much data do you need before starting an AI pilot?

You don't need perfect data to start, but you do need enough structured, accessible data to establish a baseline and train or configure a model. For demand forecasting, most AI vendors recommend a minimum of 12–24 months of sales history at SKU-retailer level. For quality inspection, a few thousand labelled images is typically sufficient. If you can't describe your data source in one sentence, it's probably not ready.

What is the difference between AI and automation for FMCG?

Automation follows fixed rules to execute a predefined sequence of steps — it does the same thing every time. AI learns from data and can handle variability, exceptions, and pattern recognition that rules-based automation can't. For FMCG, this distinction matters most in demand forecasting and quality inspection. See our guide to AI vs automation for FMCG for a fuller breakdown.


Ready to Choose Your First AI Use Case?

The 3-filter framework takes less than a day to apply across your current process landscape. The output is a ranked shortlist of use cases that are genuinely ready to pilot — not theoretically interesting, but practically deliverable in 90 days.

If you'd like a structured walkthrough with your team, book a use case selection call. You can also learn more about the BazBiff team and how we approach this work.

Explore the full library of AI and operations content on our blog.