How Kelynto weighs evidence.

Kelynto does not guess at causes, and it does not ask a language model for them. A rules-based engine tests each candidate cause against the week's data and credits it only when the evidence clears a stated bar. This page sets out each step, the thresholds behind it, and what the method does not do.

Describes engine version 2.1.0. Each statement was checked against the engine's source code and its tests on October 3, 2026. The thresholds are defaults, and a workspace administrator can change the ones marked as settings.

The method, in order.

Every weekly run goes through the same steps. Nothing in them is learned from another retailer's data.

  1. Variance decomposition

    Variance is actual sales minus plan. The week's total is rolled up by region, by category and by store, and every rollup reconciles to the total. Each store and category, which this page calls a cell, is then split into two parts: a store-wide component and what is left for the category itself.

    The store-wide part exists because of a simple observation. When most of a store's categories move together, something hit traffic. By default "most" means 60 percent of the store's categories, and the size of the effect is the median category variance, taken after setting aside categories that have their own explanation, so one runaway category cannot distort it.

    Setting: the share of categories that must move together.

  2. Store-wide causes share the traffic effect

    Signals that act on a whole store, such as severe weather, road work, a competitor opening or a store operations issue, share the store-wide effect in proportion to how well each one fits. Each is capped at the largest effect that kind of cause can plausibly have. A lane closure cannot be credited with a 40 percent collapse.

  3. Stockouts are measured, not inferred

    A stockout is read directly from your in-stock file: this week's in-stock rate for a store and category against the median of its own previous weeks. A drop of at least four points counts, and the expected loss is the plan scaled by the size of the drop. This is what separates a stockout from a demand problem.

    It needs your in-stock file. Without one, no stockout is ever credited.

  4. A cause that covers many stores is judged on all of them

    A promotion, a local event, a storm stock-up and a category-wide shift each cover a set of stores and categories. The cause is judged on the net of everything it covers, never on the cells that happen to sit above plan. Looking only at the winners would credit the cause with ordinary week-to-week noise.

    The cause is credited only if it passes a gate. The lift must point the way the cause would push. It must stand at least 3 standard errors clear of zero, and the bar rises when many causes are tested in the same week, so that testing a hundred of them does not let one through by chance. It must be at least half your percentage threshold and worth at least your minimum amount.

    Groups are judged largest first, each on what the earlier groups left behind. So a promotion at ten stores is measured against the same category at the stores without it. If the whole category ran above plan everywhere, that is a category-wide shift, and the promotion is credited with none of it.

    When the gate fails, nothing is credited, and the review says the cause was on file and what it covers did not move beyond normal variation.

  5. Evidence fit

    Each credited cause carries a fit score between 0 and 1, the product of three things.

    fit = signal strength × plausibility of the size × statistical evidence

    Signal strength combines how severe the signal is, how close it is to the store and how much of the week it covered. For an event, expected attendance counts too. Plausibility compares the size of the effect with the range that kind of cause can produce, and marks it down when it falls outside. Statistical evidence rises from the gate at 3 standard errors to full strength at 6.

    The amount credited to a cause is shared among its cells in proportion to each cell's own move in that direction, and never more than the cell actually moved.

  6. Confidence

    Confidence blends this week's fit with how often planners have confirmed that type of cause.

    confidence = confirmation rate × (0.45 + 0.55 × fit)

    The confirmation rate starts from a prior for each type of cause. Inventory availability starts at 0.85, severe weather at 0.70 and a promotion at 0.60. Each planner confirmation or rejection in your workspace updates it, with the prior weighted as eight earlier verdicts. A confidence of 0.62 or more is shown as high, 0.42 or more as medium, and anything lower as low. The band decides how firmly the commentary words the claim.

    It is a ranking aid for the planner. It is not a probability validated on real retailer data.

  7. The noise guard

    Store and category numbers wobble every week for no reason at all. After every cause has been tested, what is left in a cell is reported only when it clears three bars: your percentage threshold (4 percent of plan by default), your minimum amount (750 in your currency by default), and three times that category's ordinary spread between stores that week. Anything smaller is counted as noise and not shown as a finding.

    The spread is measured from your own data each week, so the guard needs no tuning to start. Without it, a fixed percentage bar flags cells in a week in which nothing happened.

    Settings: both thresholds and the multiple. A multiple of 0 switches the guard off.

  8. The unexplained bucket is explicit

    Variance that clears the noise guard and has no supported cause is listed as not yet explained. It is never assigned to the nearest plausible story. A category that moved across the chain with no cause on file appears as one finding, for example "Produce above plan across 119 stores, not yet explained", not scattered over a catch-all list.

    One rule holds for every run and is enforced by automated tests: explained plus unexplained plus noise always equals the variance. The review states the share that is explained, so the remainder is visible to whoever reads it.

    An unexplained line is useful. It is a missing signal, a threshold to tune, or a cause your planner already knows and can label.

  9. The planner's verdict: Confirm, Reject or Correct

    Allocations are grouped into findings, such as "Produce stockouts at 9 stores". A finding is a candidate, not a verdict. The planner confirms it, rejects it, or corrects it to a different cause. A correction can also name three causes the engine does not detect by itself: a pricing change, a plan or forecast error, or a data issue.

    Each verdict is recorded with who gave it and when. Verdicts feed the cause hit rate on the scorecard and update the confirmation rate for that type of cause in your workspace. That is the whole learning loop: there is no trained model behind attribution today.

  10. Forward signals carry a direction, never an amount

    The radar lists upcoming signals near each store for the next three weeks, each with the direction it should push sales and an impact score built from the same signal strength used above. It does not forecast a sales figure, and it does not feed your plan.

    The reason is checkability. A dollar forecast from this evidence would be unvalidated. A direction can be checked every single week: once the week has actuals, a flagged store either moved the way the radar said or it did not. Only stores that moved by at least half your percentage threshold are counted, and a store, signal and week flagged by several runs counts once.

    Settings: the horizon (1 to 8 weeks) and the minimum impact for an item to be listed.

  11. Every number in the draft is checked

    The commentary is written from the week's facts only: totals, findings and radar lines, never raw rows. Every number in a draft, whether money, a percentage, a count, a date or a number written as a word, must match a number in those facts within the rounding its own format implies. One number that does not match fails the draft, and the deterministic template version is used instead.

    The check covers numbers and currency symbols. It does not cover direction words such as above or below, names, or fractions written as words. That is one reason a planner reads every review.

What the method does not do.

A method is easier to trust when its edges are stated. These are the edges.

  • It does not prove causation. A credited cause is the best-supported candidate among the causes on file. No experiment is run, which is why a person confirms each one.
  • It only knows the causes it is given. Inventory availability, severe weather, storm stock-up, road work, competitor openings and closings, store operations, local events, promotions and local signals you supply. Anything else shows up as not yet explained.
  • It does not model cannibalisation, halo, price elasticity or assortment changes.
  • It does not use a trained model. There is no ground truth until planners label real weeks. The labels it collects are what a learned model would later need.
  • Its defaults have not been tuned on a real retailer. No retailer's data has been through the engine yet, which is why the thresholds above are settings.
  • It has no accuracy record. No figure on real retailer data exists. What will be published, and how it will be judged.
  • Weather in the radar is only as good as the public forecast inside the horizon.
  • One time zone per workspace. A chain that spans several zones chooses the one most of its stores are in.

How the method is tested.

On generated data with a known answer, because that is the only data so far where the right answer is known.

  • Weeks in which nothing happened. Ordinary noise only, with promotions and signals on file that did nothing. The test requires that no causal finding is produced.
  • A promotion that was already in the plan. The engine must not credit it with the cells that happen to sit above plan.
  • Planted effects of a known size. Promotions, weather, stockouts, road work and a competitor opening must still be found, at the size planted.
  • The arithmetic. Explained plus unexplained plus noise must equal the variance in every run.

What that does and does not show.

It shows the logic behaves as designed. It does not show accuracy on your business.

  • The honest measure is your planners' verdicts on your own weeks: the cause hit rate and the edit rate.
  • A practical first test is a backtest: run a few of your past weeks and set each draft beside what your team actually wrote.
  • The method is open to challenge. Every finding carries the evidence it rests on, and this page states the rule that produced it.

Benchmarks: the plan for published results and the four pilot measures

Four ways to take the next step.

  1. Explore

    See the product

    The sample workspace is not open to visitors yet. The product tour shows nine real screens from it.

    Take the product tour

  2. Evaluate with your data

    Run it on your own weeks

    A backtest on your past weeks, then about six live Monday reviews, judged on four measures agreed before it starts.

    See what a pilot involves

  3. Deploy in your environment

    Check the architecture

    What runs where, what crosses the boundary, and how far each path has been proven. Azure first.

    Read the deployment notes

  4. Talk to us

    Twenty minutes with the founder

    Your process and your questions. Or write to marc@kelynto.com.

    Book a 20-minute conversation

Kelynto is an early-stage company opening its first pilots with retail planning teams. The product is built and tested on synthetic data. There are no customers and no results yet, and this site says so wherever it matters.