The monthly review went fine. The consensus number for the brand came in close to actual, the accuracy slide showed something the room could live with, and the meeting moved on to promotions. Three weeks later the club account was out of the hero SKU for most of a week, and pallets of the seasonal flavour were sitting at the 3PL against a code date.
Both of those were in the forecast, cancelling each other out. The demand forecasting mistakes named in a post-mortem (aggregation, promotional lift, shipments used as demand, stockouts, calendars) are five readings of one thing: the history each was scored against had been rebuilt by hand from feeds that disagree, so a bad assumption and a moved input look identical. CPG demand planning starts with the sales history argues that case in full. Each of the five still has its own signature, and which one you are looking at decides what gets fixed first.
Five demand forecasting mistakes, and what each one is counting
None of these is a modelling error. Each one is a decision about what a number means, taken by whoever assembled the file, then inherited by the model as fact.
Accuracy measured at a level nobody orders at
Accuracy reported for the brand, for the month, is an average of errors pointing in opposite directions. A surplus on the seasonal flavour nets against a shortfall on the hero SKU and the total lands where it should. The production run at the co-packer and the truck to the club account are committed at SKU, account and week, which is the level where both of those errors were real and where both of them cost money.
Reporting at the total is a habit rather than a position anyone defends out loud. It persists because SKU-level accuracy is harder to produce: it needs the history keyed to one product across every channel that product sells in, and most brands cannot assemble that in the two days before the meeting.
Promotional volume left in the baseline
A secondary display drove four weeks of volume last spring. Where the history does not record that those weeks were promoted, the model reads them as demand and expects them again. The display does not repeat, the forecast runs high, and the miss gets written up as the category softening.
The version that costs more is the promotion that does repeat, because the lift was never separated from base and so nobody can say what the spend bought. Free cases raise the same question at smaller scale: a case given away is either a unit of demand or a cost line, and whichever answer last cycle used, it was not written down anywhere the next person can find it.
Shipments forecast as if they were demand
A shipment is your demand plus the retailer's inventory decision on top of it. A chain building stock ahead of a reset reads as growth, and the quiet weeks afterwards read as decline, so a model trained on shipment history learns a seasonality that belongs to somebody else's warehouse policy. The mistake is rarely a deliberate choice: shipments are the series that arrives daily, in your own item numbers, from a system you control, and sell-through is the one that arrives late, by their item number, on their week.
Stockout weeks recorded as demand of zero
A SKU unavailable at an account for three weeks leaves three weeks of low sales in the history, and the model reads them as evidence that demand fell. The next forecast comes down, the buy comes down with it, and the item runs out again against a smaller plan. Censored demand only ever pushes one way, which is why an item that keeps disappointing is sometimes an item nobody could buy.
Correcting for it needs two things the history rarely carries: on-shelf availability by account and week, and a flag that survives the next rebuild rather than living in a colour on a spreadsheet cell.
One calendar per source, and a week that lands in two months
Retailer files run to a week ending Saturday, grouped four and five weeks to a period. Your ERP closes on the calendar month. A marketplace settles days after the order was placed. Convert between them at the wrong point and a week of sales moves into a month it does not belong to, which surfaces as a seasonal pattern the business does not have and as a year-over-year comparison between periods of different lengths. A fifty-third week, which arrives every few years, does the same thing to an annual number and does it to everyone at once.
What the five have in common
Each one is a definition nobody wrote down: what counts as a sale, which week it belongs to, whether a promoted week is comparable to a base week, whether a zero means no demand or no stock. The definitions exist, because someone applied them when they built the file. They live in a formula, a lookup tab, or one person's head, and they change when that person is on holiday and somebody else assembles the history.
That is also why accuracy reviews stall. Comparing this cycle's error against last cycle's assumes both were scored on the same history, and a history rebuilt each month from four feeds gives you two different pasts to compare. A planner defending a call in that meeting cannot separate their own judgement from an input that moved underneath them, so the hour goes on the data and the assumption never gets discussed.
We built Permute to hold the definitions rather than the file. We connect the ERP, whether that is NetSuite or another, alongside the retailer portals, Shopify, Amazon Seller Central and the distributor spreadsheets that only ever arrive by email, and resolve every partner identifier to one product with its pack conversions dated. The rules for a promoted week, a return, and an unavailable week are held as explicit statements rather than as columns in whichever workbook was open, so a planner can ask for weekly units by SKU and account for the last two years, net of returns, on the retail calendar, and get it back with the source rows attached: this feed, this file, received at this time. The Governance layer is the part that changes an accuracy review, because it records which definition produced a number, who changed an assumption and when, and which people are entitled to see the rows underneath, so a broker working one account sees that account and a co-packer sees the products they make. When a retailer restates a month, the history moves and the workspace says that it moved. That is what a context layer actually does, and it sits underneath Claude or ChatGPT rather than in place of either.
We are not a forecasting engine. The method, the model and the judgement about next spring stay in your planning tool or with the planner who has run this category for six years. What we supply is the input, and once the input holds still an error can be attributed to a method, an assumption or an event, which is what an accuracy review is for.
Rebuild the history behind one bad month
Connect your ERP, retailer feeds and marketplaces, then pull weekly units by SKU and account for the month that missed.
How to tell which mistake you are making
Take the SKU and account pair that missed by the most last quarter (the one the sales lead still brings up) and audit the history behind it rather than the forecast on top of it.
- The error for that pair is available at SKU, account and week, without taking the brand total and allocating it down.
- Every promoted week in its history is flagged as promoted, including the display that was added late and never reached the trade calendar.
- Weeks when the item was unavailable at that account are marked unavailable rather than sitting in the history as demand of zero.
- Shipments and sell-through are two named series, and you can say which of them the forecast was scored against.
- A retailer restatement from two months ago appears as a restatement with the date it arrived, rather than having been absorbed into the current number.
- Rebuilding last quarter's history today produces the file you used at the time, or a list of what changed and why.
- The person accountable for the accuracy number can open the rows behind it, and a broker on one account cannot open the others.
The last item is a permissions question rather than a data one, and how the data is handled covers that scoping, including what an AI agent is allowed to read. Fail three or more of the seven and the next method change will be evaluated against a history that moves underneath it, which produces another review like the last one.
What to fix before the next planning cycle
Fix them in the order that removes argument: one product key across every channel, then one calendar with the conversion stated, then the flags for promotion and availability, then a measurement level you commit to for a year rather than one chosen per slide. Only the last of those is a forecasting decision, and it is the shortest of the four once the first three exist.
For the mapping and calendar work in order, How to automate POS data analysis walks through it feed by feed. If most of your volume is direct, Demand forecasting automation for ecommerce brands is the closer starting point, because the channel data arrives daily and the failure looks different. On the stock side, Inventory planning automation: what to automate first covers the assembly the same numbers feed, and AI inventory planning: what works and what breaks covers what a general model does with these exports before it runs out of window.
Across the consumer brands we work with, the forecast improves later than people expect and by less. What improves in the first cycle is the conversation: the error is decomposed into a method, an assumption and an event, and the planner spends the meeting defending a call instead of a file. The rest of the shelf is in our writing for consumer brands.
Bring the month that missed
We will look at how the history behind it was assembled, and what it would take to reproduce it next quarter.
Questions planners ask about forecast accuracy
Which accuracy metric should we report?
Report bias separately from error, because they call for different fixes. A forecast that is high in the same direction month after month is a decision problem, usually a commercial number nobody wanted to argue with, while error scattered in both directions is a signal problem in the history.
Whichever measure you choose, score it at the level the order is placed at and compare it against a naive forecast such as last year, same week. Without that comparison a number on a slide says nothing about whether the model earns its place.
Should we correct the history for stockouts?
Flag rather than overwrite. Keep the recorded sales as they happened, add an availability flag by account and week, and hold any unconstrained estimate as a separate series so a reader can see which one a forecast used.
Estimating what would have sold is judgement work and it stays with the planner. What the layer underneath contributes is the flag surviving the next rebuild, so the estimate does not have to be made again from memory.
Who should own forecast accuracy when sales and supply disagree?
One number, owned by whoever carries the consequence of it being wrong, which in most brands is supply. The commercial view still belongs in the process, entered as a named override with its author and its reason rather than blended into the statistical forecast.
Recorded that way, overrides can be scored on their own. A team that finds its adjustments consistently make the forecast worse tends to make fewer of them.
Do we need new planning software to fix these?
Four of the five are definition and assembly problems that sit upstream of any planning tool, and they arrive intact in a new one. A planning tool owns the projection and the replenishment logic, and it reads whatever history you hand it.
Where a brand has no planning tool at all, the honest first purchase is usually the history rather than the software, because a tool bought first spends its first year being blamed for inputs.