Benchmarks.

No benchmark results are published yet. This page fixes, in advance, what will be published and how it will be judged, so that when results arrive they can be read against a method that was written down first.

Status on October 3, 2026: no benchmark results exist. No retailer's data has been through Kelynto, so there is nothing to report on real data.

The only figures available today are local test measurements on synthetic data. None of them is shown on this page, because none of them is a result.

Results

None published. This is where results will appear, one table per dataset, each row with its sample size, its baseline and the engine version that produced it.

A result will be published here only when it comes from data described under Datasets, was produced under the holdout rule below, and, where a retailer's data is involved, the retailer has agreed in writing to its publication.

Datasets

Three kinds of data can say something about Kelynto. They answer different questions, and results will always name which kind they come from.

DatasetWhat it can showWhat it cannot showExists today
Generated weeks with planted effectsWhether the engine finds a cause of a known size, and stays silent in a week in which nothing happened.Accuracy on a real business. The effects are ones we wrote.Yes. It is what the automated tests run on.
Retailer backtestHow a draft for a past week compares with the commentary the retailer's own team wrote for that week.Whether a planner would have accepted the draft at the time.No. No retailer's data has been loaded.
Live pilot weeksWhat planners did with real drafts: how much they rewrote, which causes they confirmed, whether radar calls came true.How the result transfers to a different retailer.No. No pilot has run.

For any retailer dataset the publication will state the number of stores, categories and weeks, which optional files were supplied (in-stock, promotions, local signals), and which signal sources were switched on. The retailer will be named only with its permission.

Methodology

What is being measured is the method described on the Methodology page: variance decomposition, the gate a cause must pass, the noise guard, the confidence bands and the direction-only radar. Results will be produced by running the product as a customer would run it, through the same loaders and the same weekly run.

  • Every setting that differs from its default will be listed beside the result.
  • The commentary writer will be named: the deterministic template, or a language model and which one.
  • Planner verdicts will be the planners' own. A verdict entered by Kelynto staff on a planner's behalf will be counted separately and labelled.

Holdout strategy

The rule is that no evaluated week may influence the settings it is evaluated under.

  • Holdout in time. Thresholds are fixed before the evaluated weeks are run. Confirmation rates for a week are learned only from verdicts on earlier weeks, which is how the product works in normal use.
  • No tuning on the test weeks. If a setting is changed after seeing results, the weeks already seen stop counting as holdout and will be reported as such.
  • Generated data. Published runs will use data generated from seeds that were not used while the engine was being developed.
  • All weeks reported. A week will not be dropped because its result is poor. Weeks excluded for a data fault will be listed with the reason.

Metrics

The four measures are the ones on the product's own scorecard, so a published figure and a customer's figure mean the same thing. Each is reported with its sample size and is withheld below a minimum.

MetricDefinitionMinimum sample
Draft edit rateThe word-by-word difference between the draft and the text the planner published. 0 means sent unchanged, 1 means rewritten. Reported as the median of up to the last four reviews.2 reviews with planner edits
Cause hit rateFindings planners confirmed, out of those they reviewed. Also reported by type of cause.10 reviewed causes
Radar direction accuracyFlagged stores that moved the way the radar said, among those that moved materially.10 stores that moved materially
Would use every weekThe planner's weekly answer on a scale of 1 to 5.3 answers

Reported beside them: the share of material variance that was explained and the share left unexplained, how often radar items were new to the planner, and the minutes a review took against the team's own stated starting point. On generated data, two more: how many planted causes were found at the planted size, and how many findings appeared in weeks in which nothing happened.

Baselines

A number means little without something to hold it against. Each result will be reported next to the baseline that fits it.

  • The team's own commentary. In a backtest, the causes named in the draft are compared with the causes the retailer's team named for the same week.
  • A fixed threshold with no gate. The same data run through a plain rule: flag every store and category beyond a fixed percentage of plan and credit the nearest signal. This shows what the gate and the noise guard add, or do not.
  • Chance, for direction. Radar direction accuracy is read against a coin toss, and against always predicting the usual direction for that type of signal.
  • The pilot targets. 20 percent or less of the draft rewritten, 60 percent or more of causes confirmed, 75 percent or more of radar calls right, and 4 of 5 or higher on would-use. These are targets set before a pilot. They are not results.

Limitations

These will apply to anything published here, and each publication will repeat the ones that bear on it.

  • Generated data proves that the logic behaves as designed. It says nothing about accuracy on a real business.
  • A planner's verdict is a judgement, not ground truth. A confirmed cause can still be wrong, and a rejected one right.
  • The first retailers will be few and will have chosen to take part. Their results will not be a sample of retail.
  • A result from one retailer does not transfer to another with different categories, signals or data quality.
  • Minutes saved is self-reported by the planner against a starting point the team states itself.
  • The radar's weather signals are only as good as the public forecast inside its horizon.

Version tested

Every result will carry the engine version that produced it, and results from different versions will not be pooled. The engine version is stored on every weekly run, so a figure can always be traced to the code that produced it.

The current engine version is 2.1.0. No results have been published for it or for any earlier version.

What exists today

Three sets of figures exist. All three come from synthetic data on a development machine, and none is a benchmark.

  • Automated tests on generated weeks. They pass or fail: planted causes must be found, and a week of pure noise must produce no causal finding. They are a check on the logic, described on the Methodology page.
  • Timing measurements. How long loading, the weekly run and page reads took at 500 and 2,000 synthetic stores on one small development machine. They guide engineering and are shared during a technical review. A development machine is not a benchmark environment, so they are not quoted here.
  • The sample workspace scorecard. Its figures come from a synthetic retailer with simulated planner reviews. They show what the screen looks like. They are not results.

The first real figures will come from a pilot, measured on the scorecard above.

See what a pilot involves How Kelynto weighs evidence Ask about a backtest