Skore and tabular AI: how to check a business forecast
Evaluate agent-built business forecasts with prediction-time inputs, appropriate data splits and a purchasing baseline. What Skore's checks can establish.
By Clairevue · · 5 min read

Suppose a wholesaler wants an AI-built model to predict next week’s sales before placing Monday’s purchase order. The historical test looks good, but one input column contains the stock left at the end of the week being predicted.
Nobody knows that number on Monday. The model has information it won’t have when someone needs the forecast.
Skore offers agents and evaluation tools for spreadsheet-shaped business data; before buying stock on their advice, test the forecast using only what the business knew at the decision date.
What Skore offers
Probabl’s October 6 launch announcement describes Skore agents that build, validate and track predictive models. In Assist mode, a person drives the workflow while the agent helps; in Delegate mode, the agent builds and validates with a chosen level of check-ins.
The output is a predictive model trained on tabular data, such as rows of historical orders. An AI coding agent can write its training code, while the trained model produces the numerical predictions.
Probabl reports reaching a working, deployable model 1.78× faster than general-purpose AI coding agents in an internal benchmark. That measures development speed. The announcement doesn’t give enough experimental detail to reproduce the comparison, and we haven’t tested Skore’s speed or forecast accuracy.
Separately, the open-source Skore library documents evaluation reports and ways to save or retrieve them. It can compare models and expose scores across different data splits. Those capabilities make experiments easier to inspect; they don’t establish that every input column belongs in a real purchasing forecast.
A good score can come from leaked answers
Scikit-learn’s data leakage guide documents a feature-selection example using 200 synthetic rows with 10,000 random features and randomly assigned binary labels. There’s no genuine predictive relationship to learn.
Selecting promising features using the whole dataset before splitting it gives a reported test accuracy of 76%. Selecting features using only the training subset gives 50%, around chance. These are the documentation’s synthetic results, not a Skore benchmark or a business forecast we ran.
The first procedure lets the test answers influence which features survive. The model then faces a test it has already helped design.
For the wholesaler, the corresponding problem could be a column updated after Monday: the week’s closing stock, or a delivery recorded as completed later. A planned delivery that staff already knew about on Monday is different. Check when each value became available, rather than judging its safety from the column name.
Fit missing-value rules and feature selection only on each training subset, then apply them to its test subset. A scikit-learn pipeline can keep those steps together during cross-validation. It can’t turn next week’s closing stock into an input available today.
Make the split match the forecast
The Skore evaluation API lets the developer choose how to divide training and evaluation data. Without an explicit choice, it normally uses a shuffled 80/20 split; a skrub learner with a preconfigured splitter uses that configuration instead.
Random splitting can suit some prediction tasks. For this purchasing forecast, it mixes later observations into training while testing on earlier ones, a situation Monday’s buyer can’t reproduce.
Test past-to-future: train on earlier periods and forecast later weeks. Only include training outcomes that had finished before each test forecast’s issue date. Scikit-learn’s time-series forecasting example shows how to use earlier observations and evaluate on later periods; its bike-rental results aren’t an estimate of this wholesaler’s performance.
Skore has a documented temporal-overlap check. It checks datetime-typed columns in pandas training and test tables, flagging cases where the latest training timestamp reaches or passes the earliest test timestamp. A passed timestamp check doesn’t establish when a delivery-status field became available in the business. Record which checks ran and explain omissions, including explicit ignores and fast-mode skips of slower checks. An empty warning list doesn’t show whether the developer ran every applicable check.
Compare the forecast with the purchasing rule you use now
For the hypothetical wholesaler, try a simple baseline that predicts next week using the last completed week’s sales. Evaluate it and the proposed model on the same later weeks, with the same information available at each forecast date.
Report error in units the buyer understands. Examine under-prediction and over-prediction separately, because too little stock and unsold inventory have different consequences. If products were out of stock, recorded sales also need that context before treating them as the demand the business could have served.
An agent can generate many variations quickly. Keep the final test period away from that search. Use separate validation periods to choose features and settings, then evaluate the selected model on the held-back period. Google’s machine-learning guidance explains how repeated tuning against a test set gradually makes it less useful as a test of new data.
Keep the dated input snapshot and exact split with the saved report, alongside the pipeline version and any skipped-check explanations. Record the time and compute used to produce it. Request these records from the pilot so another reviewer can inspect how the developer obtained its results.
Before letting a forecast change purchase orders, run it alongside the current buying process. Save its prediction before the week starts, then compare it with the outcome and the baseline. Keep a buyer reviewing the order until those forward-looking results justify changing the decision process.