TServe – Open-source inference server for time-series foundation models Sktime released TServe, an open-source inference server that loads time-series foundation models including Chronos, TimesFM, Moirai, TTM, and TiRex once into memory and serves forecast requests over HTTP via a single POST /predict endpoint. The server ships over 100 checkpoints, each with its own Docker tag or pip extra for CPU or GPU, and wraps model families through sktime's common forecaster interface so switching models requires changing one field. TServe runs on user hardware rather than as a hosted API, with a Python client that accepts dict, pandas, polars, or pyarrow tables and returns predictions in the same type. | | Documentation https://tserve.readthedocs.io/en/latest/ · Quick start https://tserve.readthedocs.io/en/latest/quick-start/ · Models https://tserve.readthedocs.io/en/latest/models/ · API https://tserve.readthedocs.io/en/latest/reference/http/ | |---|---| | Project | | | Status | | Time series serving for foundation models. TServe loads models such as Chronos, TimesFM, Moirai, TTM, and TiRex once, keeps them in memory, and answers forecast requests over HTTP. Each model family ships its own package, input format, and loading code. sktime https://www.sktime.net/ wraps them as forecasters with one common interface, and TServe runs those forecasters as a server behind a single request: a table of past values and a horizon. Trying another model means changing one field, not rewriting your pipeline. - Over 100 checkpoints. Each family has its own Docker tag or pip extra, for CPU or GPU. Catalog https://tserve.readthedocs.io/en/latest/models/ · Capabilities https://tserve.readthedocs.io/en/latest/models/ capabilities - Loaded once, kept warm. Weights download and load at startup, so each request pays only for inference. - JSON from anywhere. POST /predict works from curl or any language. Send a prediction https://tserve.readthedocs.io/en/latest/client/http/ send-a-prediction - Native tables in Python. The Client https://tserve.readthedocs.io/en/latest/client/python/ connect takes a dict, pandas, polars, or pyarrow table and returns predictions in the same type. - A dashboard in the browser. GET / plots a forecast from a sample series or your own CSV. What you can do https://tserve.readthedocs.io/en/latest/server/dashboard/ what-you-can-do - Your own sktime models. Serve a configured forecaster https://tserve.readthedocs.io/en/latest/server/live-objects/ , a saved .zip https://tserve.readthedocs.io/en/latest/server/models-dir/ , or a craft spec https://tserve.readthedocs.io/en/latest/server/craft-specs/ next to the catalog models. TServe is a server you run on your own hardware, not a hosted API. How it works https://tserve.readthedocs.io/en/latest/overview/ · Docker Hub https://hub.docker.com/r/sktime/tserve Docker is the short path. This image can load Chronos Bolt, Chronos T5, TTM, and TimesFM 2.x. The first start downloads the weights you name. docker run --rm -p 8000:8000 sktime/tserve:hub chronos bolt ttm r3 When the log prints the local URLs, the models are warm. Five days of sales, three steps ahead: curl -s http://127.0.0.1:8000/predict -H "Content-Type: application/json" -d '{ "past": { "timestamp": "2024-01-01", "2024-01-02", "2024-01-03", "2024-01-04", "2024-01-05" , "sales": 120, 135, 128, 142, 138 }, "fh": 3, "model": "chronos bolt" }' { "predictions": { "timestamp": "2024-01-06T00:00:00", "2024-01-07T00:00:00", "2024-01-08T00:00:00" , "sales": 139.96, 138.93, 138.26 }, "quantiles": null, "model": "chronos bolt", "request id": "…" } Open http://127.0.0.1:8000/ http://127.0.0.1:8000/ , pick chronos bolt , and plot the same series. The page can also take a pasted or dropped CSV. What you can do https://tserve.readthedocs.io/en/latest/server/dashboard/ what-you-can-do The same call from Python. The client posts Arrow, and predictions comes back as the same kind of table you sent: pip install "tserve client " python from tserve.client import Client past = { "timestamp": "2024-01-01", "2024-01-02", "2024-01-03", "2024-01-04", "2024-01-05" , "sales": 120, 135, 128, 142, 138 , } with Client "http://127.0.0.1:8000" as client: result = client.predict past=past, fh=3, model="chronos bolt" print result.predictions The walkthrough, including GET /models and PowerShell: Quick start https://tserve.readthedocs.io/en/latest/quick-start/ . A GPU host adds --gpus all and uses sktime/tserve:hub-gpu . GPU images https://tserve.readthedocs.io/en/latest/server/docker/ gpu-images The running server serves a browser console at GET / . The model list is whatever this process loaded. You set a horizon, optionally a prediction interval, and a series a built-in sample, pasted CSV, or a dropped file, parsed in the browser , then the page posts POST /predict and plots the result. Health and runtime stats sit on the right. What you can do https://tserve.readthedocs.io/en/latest/server/dashboard/ what-you-can-do Below, timesfm 3 https://tserve.readthedocs.io/en/latest/models/timesfm3/ forecasts retail sales with 90% prediction interval. 117 checkpoints. The extra name is the image tag, sktime/tserve: