Appearance
The old economics of forecasting
Forecasting has always lived two lives. One life is research: foundation models that take any signal and answer what comes next. The other life is operations: a demand planner defending a spreadsheet, a process engineer watching thresholds that never change. The gap between them has always been economic, not technical.
The old economics go like this: one bespoke model per use case, months of expert work per model. So teams model the few hundred series where the money is, and cover the rest with safety margins. Extra inventory. Extra headroom. Extra tolerance, acted on after the window has closed. That margin is the cost of a decision nobody could forecast, and you pay it every cycle.
A time series foundation model (TSFM) changes the arithmetic. Trained once across vast, varied signals, it generalizes to a series it has never seen. Give it a window of measurements and it tells you what comes next, how far current behavior sits from normal, which past period resembles this one, which settings best serve a target. The same weights handle a shampoo line, a card network and a retail catalogue.
Google's timesfm-3.0-pytorch model card sits on Hugging Face as open weights. IBM Granite time series models report 44M+ downloads before this launch, a level of usage that only a generalizing design reaches. Different vendors, same bet: a model that has seen enough different signals will handle one it has never seen.
IBM says it ran these models before selling them, inside its own operations first, then with design partners in cement, steel, pulp and paper, food and telecommunications. The writeup reports productivity gains of 5 to 10x, and forecasting work that waited for specialists moved to the domain experts who own the decision. When I read their account of a grocery demand planner pointing one shared model at the whole catalogue, a two-year SKU and a three-month one in the same job, a launch with no history starting from the SKUs it resembles, that is the part that lands. One model becomes a model factory.
The reanalysis trick
Foundation models need a lot of training data, and most forecasting domains do not have a forever dataset. The interesting work is happening where that gap gets closed synthetically. CloudCast v2, a weather paper from arXiv, shows the pattern in the domain where forecasts are hardest to fake.
The weather problem is an initialization problem. Short-range nowcasting preserves observed cloud placement for the first hours, then loses skill as clouds form, dissipate and deform. Longer lead times need atmospheric evolution, but operational numerical weather prediction does not represent the satellite-observed cloud state when it starts. One model sees the truth but cannot evolve it. The other can evolve the atmosphere but starts from the wrong sky.
CloudCast v2 splits the problem in two. First it trains on the Copernicus European Regional Reanalysis, a synthesized but physically consistent record, and learns how cloud fields evolve. Then it adapts to satellite-derived observations using conditional flow matching, a generative method that turns noise into a cloud-cover forecast, conditioned on the observed initial field and NWP inputs.
Result: 10% lower mean absolute error than its predecessor, CloudCast v1, across the 1-12 hour range. In fractions skill score, a neighborhood-based measure of spatial agreement, it overtakes v1 after about 3-6 hours, depending on cloudiness category. What that means in practice: the observation-initialized model extends useful skill from the usual 1-3 hour nowcasting window to a full 12 hours, and keeps the spatial detail that satellites actually see.
Google's TimesFM line plays the same trick. Its 3.0 generation trains with SimGen, a generator that synthesizes training windows, then zero-shots to unseen series. Reanalysis for weather, a generator for general series. Same structure: learn dynamics from a source that has everything, adapt to the stream that has the truth.
Key numbers
- 5-10x: productivity gain IBM reports from design partners once forecasting work moved from specialists to domain experts.
- 44M+: downloads behind the IBM Granite time series models before this launch.
- 10%: lower mean absolute error for CloudCast v2 vs v1 over 1-12 hours, enough to matter when one point of accuracy is worth millions.
- 100,000+: series per night on CPU for the million-parameter TTM model, meaning the long tail finally gets a forecast at all.
Four models, not one
No single model serves a shampoo line, a card network and a retail catalogue alike. IBM ships a portfolio, and the design decisions reveal something: the best model for a job is often the smallest one.
The four Granite models are in Early Access, all callable through the same Flink SQL functions. The differences matter, and they are architectural, not marketing:
| Model | How it reads a series | Returns | Best when |
|---|---|---|---|
| PatchTST-FM | Patch by patch, like a language model reads text; each variable in its own channel | A full forecast distribution | A planner needs percentiles, e.g. reorder points at the 90th |
| FlowState | A running summary updated with every point; continuous in time | Forecasts on mixed frequencies | Seconds-level SCADA and hourly market data in one model |
| TTM | Tiny mixing networks along time and variables, no attention | Point forecasts at scale | 100,000+ series nightly on CPU |
| TSPulse | Paired time and frequency views, one small multi-task model | Anomaly scores, classifications, gap-fills, similar past cases | "Have we seen this before?" questions |
PatchTST-FM keeps each variable in its own channel so one noisy signal cannot drag the rest down. FlowState's dynamics are continuous in time, so it reads seconds-level SCADA and hourly market data without resampling. TTM drops attention entirely, and a million-parameter model covers a hundred thousand series nightly on CPU. TSPulse pairs time and frequency views to answer the question every operator asks.
Small was a decision, not a compromise. Inference runs natively inside Confluent Cloud, or from open weights on your own CPUs, with zero cloud ingress or egress. The model is matched to the decision: TTM for bulk, PatchTST-FM for distributions, FlowState for mixed frequencies, TSPulse for pattern matching.
Quick Take: Forecasting was never bottlenecked on accuracy; it was bottlenecked on per-series specialist labor, and foundation models turn that whole practice into a function call.
The streaming path
The interesting part is where the model runs. Confluent Cloud now hosts the Granite models natively, and the whole pipeline, stream to forecast to alert, is one SQL call:
sql
SELECT
AI_FORECAST(
load_kw,
event_time,
JSON_OBJECT('model' VALUE 'ttm', 'horizon' VALUE 12)
) OVER (
ORDER BY event_time
RANGE BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW
) AS forecast
FROM meter_readings;Change the model value and the same call runs any of the four. No separate ML stack to build, no GPU to provision, no warehouse to copy the series into. Confluent manages serving, scaling and runtime operations. Inference results land on Kafka topics, replayable, governed by the same schemas and access controls as everything else on the platform. Access opens on Confluent Cloud on AWS, with Confluent Platform bringing the same setup to on-premises and hybrid environments later.
Statefulness is the move that makes it work. The next value only means something against recent history, and an anomaly only exists against a running sense of normal. Flink keeps that state keyed per series and fault tolerant, so each model gets the history it needs without a database hit per call. Do it right and a signal decays in seconds instead of days. A pump caught drifting today is a work order. The same pump next week is an outage.
Where the value lands
The use cases read like a tour of everything forecasting was supposed to fix.
Forecasting turns service level into policy. The demand planner gets a distribution, not a line, so replenishment fires from it, a markdown lands before stock ages, a reorder lands before the shelf empties. Fewer stockouts, fewer markdowns, working capital freed from the tail of the catalogue.
Anomaly detection gets transfer. A card that bought groceries in the same three postcodes for two years funds a wallet abroad at 3am. The model keeps a sense of normal per card and scores the payment in flight. Protection starts day one because what the model learned elsewhere transfers to new products and corridors with no labeled case. The same customer on an honest holiday sails through. Tightening rules declines honest customers, and that revenue is gone too.
Optimization turns a forecast into a simulator. Andrés runs process at a shampoo plant whose mixing line streams temperature, agitator speed, dosing rate and viscosity. Conditioned on the settings he controls, the forecast answers what-if: energy at this mixing speed, throughput at this temperature, viscosity in spec or not. An optimizer searches that space against a KPI he names, respects his constraints, and explains its recommendation. A recommendation he cannot interrogate, he will not act on.
Common pitfalls
The pattern is young enough that the failure modes are predictable. Five show up constantly.
First, zero-shot does not mean zero-config. A foundation model still needs the right context length and the right frequency. Feed it hourly data, ask for a daily forecast, and the mode mismatch produces a confident wrong answer. Match the frequency, give it enough history, check what normalization the weights expect.
Second, the cold start. New SKUs, new products, new regions have no history, and the model leans on analogy to similar series. TSPulse exists to answer exactly this question. But if nothing maps the new series to the family it resembles, the forecast is mean-reverting mush. Similarity retrieval is part of the pipeline, not an extra.
Third, wrong-size model for the job. Teams standardize on one model for everything and either burn GPU budget on the long tail or starve the few series that move millions. The million-parameter TTM exists because most of your series are not worth a GPU call, and the accuracy math cuts both ways: every point of accuracy is worth millions, until you multiply it across ten thousand low-stakes series.
Fourth, copying data out of the stream to score it. Extraction to a batch store re-creates the exact decay the tooling was designed to remove. By the time the batch lands, the pump that was drifting is an outage, not a work order. Keep state, inference and delivery on the platform where the data lives.
Fifth, fine-tuning away the foundation. CloudCast v2 gets its edge from the reanalysis phase, because observations show state, not dynamics. Reanalysis teaches formation, dissipation, deformation. Adapt to your domain, but if the fine-tune suppresses what the pretraining learned, you have bought a memorizer of last month.
One thing to remember
One thing to remember: the model is the smallest part of the system. The wins in these deployments come from integration, the stateful per-series context, the governance on the topics, the delivery into alerting and optimization loops. A great forecast that lands Friday for a decision that had to happen Wednesday is a dataset, not a decision.
Sources
- CloudCast v2 paper, arXiv: http://arxiv.org/abs/2609.03763v1
- IBM Research, "Real-Time Intelligence with IBM Time Series Models on Confluent", Hugging Face blog: https://huggingface.co/blog/ibm-research/real-time-intelligence
- Google, "google/timesfm-3.0-pytorch" model card, Hugging Face: https://huggingface.co/models/google/timesfm-3.0-pytorch
The bottom line
Three takeaways.
If you run planning across a long tail of series, adopt the shared-model play: one foundation model, distributions instead of point forecasts, no per-SKU modeling. That is the 5-10x productivity story, and it is the fastest win in this entire cluster.
If you need real-time decisions on a stream, keep forecasting inside the stream. Flink SQL with AI_FORECAST and AI_DETECT_ANOMALIES, state keyed per series, results on replayable topics. Extracting to batch to score it re-creates the decay the model was supposed to fix.
If you operate where labeled data is scarce, follow the CloudCast route: pretrain on synthesis or reanalysis, adapt lightly to observations. It just pushed weather forecasting past the 1-3 hour wall to 12 hours, and the same pattern is how TimesFM generalizes. Expect adapted-generator pretraining to spread into other industrial domains within a year.