# T – Reproducible Pipelines for Polyglot Data Science

> Source: <https://tstats-project.org/>
> Published: 2026-10-05 23:58:00+00:00

**Use Julia, Python, and R for what they’re really good at —
whatever that is for you. T orchestrates them.**

Simulations in Julia, ML in Python, statistics in R — or the exact opposite. It doesn’t matter how you divide the labor: the hard part of polyglot data science was never the languages, it was the fragile seam between them.

A language for the LLM era, T is designed to be piloted by both humans and AI models. It gives you one hermetic dependency graph where your tools communicate without glue and execute consistently through space and time: on your laptop today, on a cluster tomorrow, and five years from now without bitrot.

**Status:** Version 0.55.5 “L’Ultime combat”.

**[Install Nix](nix-installation.html)**
(installs Nix and configures the `rstats-on-nix` cache in one
step):

```
curl --proto '=https' --tlsv1.2 -sSf -L https://install.determinate.systems/nix | \
  sh -s -- install --no-confirm --extra-conf "
trusted-users = root $USER
substituters = https://cache.nixos.org https://rstats-on-nix.cachix.org
trusted-public-keys = cache.nixos.org-1:6NCHdD59X431o0gWypbMrAURkbJ16ZPMQFGspcDShjY= rstats-on-nix.cachix.org-1:vdiiVgocg6WeJrODIqdprZRUrhi1JzhBnXv7aWI6+F0="
```

**Try T immediately** in an ephemeral shell:

```
nix shell --accept-flake-config github:b-rodrigues/tlang
```

**Scaffold a project** and enter its pinned
environment:

```
t init --project my_project && cd my_project && nix develop
```

*(See the [Nix Installation
Guide](nix-installation.html) and [Getting Started
Tutorial](getting-started.html) for full platform instructions).*

Run `t demo` right in your terminal to see pipeline
introspection, hermetic Nix builds, Arrow in-memory inspection, caching,
and first-class error handling in action. The demo builds in a scratch
directory under your current directory (so it resolves your project’s
flake) and removes it on exit:

A complete analysis that simulates non-linear data in Julia, fits a gradient-boosted regressor in Python, plots ground truth vs predictions in R, and compiles a Quarto report:

```
p = pipeline {
  -- 1. Simulate non-linear DGP in Julia (seeded)
  sim_data = jln(
    command = <{
      using Random, DataFrames
      Random.seed!(42)

      t = 1:500
      shock = cumsum(randn(500))
      DataFrame(time = t, shock = shock, signal = sin.(t ./ 20) .+ shock .* 0.2)
    }>,
    serializer = ^ipc
  )

  -- 2. Train non-linear model & predict in Python (scikit-learn)
  predictions = pyn(
    command = <{
from sklearn.ensemble import HistGradientBoostingRegressor

X = sim_data[['time', 'shock']]
y = sim_data['signal']
model = HistGradientBoostingRegressor(random_state=42).fit(X, y)
sim_data['pred'] = model.predict(X)
sim_data
    }>,
    deserializer = [sim_data: ^ipc],
    serializer = ^ipc
  )

  -- 3. Publication figure in R (ggplot2)
  plot = rn(
    command = <{
      library(ggplot2)

      ggplot(predictions, aes(x = time)) +
        geom_point(aes(y = signal), alpha = 0.3, color = "#7f8c8d") +
        geom_line(aes(y = pred), color = "#e74c3c", linewidth = 1) +
        labs(title = "Julia Simulation + Python ML Predictions", y = "Value") +
        theme_minimal()
    }>,
    deserializer = [predictions: ^ipc]
  )

  -- 4. Render reproducible Quarto report
  report = node(script = "src/report.qmd", runtime = Quarto)
}

build_pipeline(p)
```

`ggplot`
object directly; T’s runner automatically renders and caches the visual
artifact without `ggsave()`. DataFrames pass between nodes
via Apache Arrow IPC (`^ipc`) without `read.csv()`
or `to_csv()` glue.`<{ ... }>` blocks. Nodes
accept external script files directly
(`jln(script = "sim.jl")`,
`pyn(script = "train.py")`,
`rn(script = "plot.R")`). Your Julia, Python, and R scripts
remain ordinary standalone files that your team can run or reuse
anywhere with standard tooling.`raise`), R (` stop()`), Julia
(`error()`), or T (` VError` artifact, and allows
independent branches to complete. Downstream nodes can inspect the error
with `read_node()` or `explain()`, or recover
programmatically.`src/report.qmd` into an HTML or PDF report inside the Nix
sandbox, directly embedding upstream metrics and figures.
Most modern quantitative projects in research, central banks,
official statistics, and regulated industries are polyglot by necessity:
- **Julia** is unmatched for raw numerical simulation,
ODEs, and heavy optimization loops. - **Python** is the
standard for modern machine learning and deep learning tooling. -
**R** remains the gold standard for survey statistics,
econometrics, and publication-ready reporting.

Connecting them today forces you to choose between three bad options:

| The Status Quo | The Failure Mode | 
|---|---|
| **In-process FFI (`reticulate`, `PyCall`, `RCall`)** | Shared memory between multiple runtimes with competing garbage collectors and conflicting OpenMP/BLAS threads causes unexplained segfaults. Upgrading one runtime breaks the other. | 
| **Ad-hoc Bash scripts & CSVs** | No caching: tweaking a title in an R ggplot re-runs your 3-hour Julia simulation. Column types and missing values silently mutate during CSV export. | 
| **Chained Docker containers** | Huge container images, slow local development, impossible for an analyst to inspect or debug interactively on a laptop. | 

`^onnx`, `^pmml`,
`^csv`). No custom serialization glue scripts.
| Feature | {targets} | {rixpress} | Snakemake | Docker (packaging only) | **T** | 
|---|---|---|---|---|---|
| **Interface & Engine** | R package ( `_targets.R` ), host environment | R package API, Nix engine | Python / CLI DSL, Conda/host | Container image, Docker daemon | **Dedicated pipeline language, Nix engine** | 
| **Cross-language seam** | R-native (polyglot is bolted on) | R-native (Python nodes via `rixpress` helpers) | Shell scripts & CLI wrappers | Manual entrypoints & volume mounts | **Process-isolated IPC across R, Python, and Julia** | 
| **Intermediate I/O** | Automatic | Automatic | Manual file paths | Manual volumes & files | **Automatic (zero-boilerplate boundary transfer)** | 
| **Node caching** | Content-addressed (R) | Content-addressed (R) | Timestamp / file hash | Docker build layer cache | **Content-addressed (all nodes)** | 
| **System library locking** | ❌ (Delegates to host) | ✅ (Hermetic Nix) | ⚠️ (Optional Conda) | ✅ (Per image) | ✅ (Hermetic per-node Nix sandbox) | 
| **Interactive inspection** | ✅ ( `tar_read()` ) | ✅ ( `read_node()` ) | ⚠️ (File inspect only) | ❌ (Container attach) | ✅ ( `read_node()` ,`explain()` ) | 
| **Error resilience** | ❌ (Aborts run) | ❌ (Aborts run) | ❌ (Aborts run) | ❌ (Container exits) | ✅ (First-class polyglot soft-failures) | 

When you define a node using `node()`, `rn()`
(R), `pyn()` (Python), `jln()` (Julia), or
`shn()` (Shell), T treats the result as a first-class
**Node** object. These objects transition through two main
states:

`build_pipeline()`,
the node points to a concrete, immutable artifact in the Nix store.
When you call `read_node(p.node_name)` in the REPL, T
looks at the node’s **serializer** and attempts to
automatically load the data back into the T environment:

| Serializer | Resulting T Type | Backend | 
|---|---|---|
| `default` /`serialize` | Varies | Native T binary serialization | 
| `arrow` | `DataFrame` | Apache Arrow IPC (zero-copy) | 
| `csv` | `DataFrame` | Native CSV parser | 
| `json` | `Dict` /`List` | JSON parser | 
| `pmml` | `Model` | Native T model evaluator | 

You can use `explain()` to look inside a built node:

```
-- Example: Inspecting a built R node
> model_node = p.model_r
> explain(model_node)
{
  `kind`: "computed_node",
  `name`: "model_r",
  `runtime`: "R",
  `path`: "/nix/store/...-model_r/artifact",
  `serializer`: "pmml",
  `class`: "lm",
  `dependencies`: ["data"]
}
```

The `path` field is the escape hatch: it gives you the
absolute path to the node’s output in the Nix store. You can inspect the
artifact directly or pass it to external tools.

`t update`, and build a
hello-world pipeline`fct_*` helpers`|>`
forwarding semantics and short-circuiting`^serializer` system for data interchange and
materialization`%` shortcuts (`%cd`, `%env`, and
more)
