T – Reproducible Pipelines for Polyglot Data Science T version 0.55.5 "L'Ultime combat" has been released, offering a hermetic dependency graph that orchestrates Julia, Python, and R pipelines through Nix and passes DataFrames between nodes via Apache Arrow IPC. The tool, installable via `nix shell --accept-flake-config github:b-rodrigues/tlang`, lets users scaffold a project with `t init --project my_project` and run a demo showing pipeline introspection, caching, and first-class error handling. T's pipeline syntax supports inline code blocks and external script files, with a Quarto node rendering reproducible reports. Use Julia, Python, and R for what they’re really good at — whatever that is for you. T orchestrates them. Simulations in Julia, ML in Python, statistics in R — or the exact opposite. It doesn’t matter how you divide the labor: the hard part of polyglot data science was never the languages, it was the fragile seam between them. A language for the LLM era, T is designed to be piloted by both humans and AI models. It gives you one hermetic dependency graph where your tools communicate without glue and execute consistently through space and time: on your laptop today, on a cluster tomorrow, and five years from now without bitrot. Status: Version 0.55.5 “L’Ultime combat”. Install Nix nix-installation.html installs Nix and configures the rstats-on-nix cache in one step : curl --proto '=https' --tlsv1.2 -sSf -L https://install.determinate.systems/nix | \ sh -s -- install --no-confirm --extra-conf " trusted-users = root $USER substituters = https://cache.nixos.org https://rstats-on-nix.cachix.org trusted-public-keys = cache.nixos.org-1:6NCHdD59X431o0gWypbMrAURkbJ16ZPMQFGspcDShjY= rstats-on-nix.cachix.org-1:vdiiVgocg6WeJrODIqdprZRUrhi1JzhBnXv7aWI6+F0=" Try T immediately in an ephemeral shell: nix shell --accept-flake-config github:b-rodrigues/tlang Scaffold a project and enter its pinned environment: t init --project my project && cd my project && nix develop See the Nix Installation Guide nix-installation.html and Getting Started Tutorial getting-started.html for full platform instructions . Run t demo right in your terminal to see pipeline introspection, hermetic Nix builds, Arrow in-memory inspection, caching, and first-class error handling in action. The demo builds in a scratch directory under your current directory so it resolves your project’s flake and removes it on exit: A complete analysis that simulates non-linear data in Julia, fits a gradient-boosted regressor in Python, plots ground truth vs predictions in R, and compiles a Quarto report: p = pipeline { -- 1. Simulate non-linear DGP in Julia seeded sim data = jln command = <{ using Random, DataFrames Random.seed 42 t = 1:500 shock = cumsum randn 500 DataFrame time = t, shock = shock, signal = sin. t ./ 20 .+ shock . 0.2 } , serializer = ^ipc -- 2. Train non-linear model & predict in Python scikit-learn predictions = pyn command = <{ from sklearn.ensemble import HistGradientBoostingRegressor X = sim data 'time', 'shock' y = sim data 'signal' model = HistGradientBoostingRegressor random state=42 .fit X, y sim data 'pred' = model.predict X sim data } , deserializer = sim data: ^ipc , serializer = ^ipc -- 3. Publication figure in R ggplot2 plot = rn command = <{ library ggplot2 ggplot predictions, aes x = time + geom point aes y = signal , alpha = 0.3, color = " 7f8c8d" + geom line aes y = pred , color = " e74c3c", linewidth = 1 + labs title = "Julia Simulation + Python ML Predictions", y = "Value" + theme minimal } , deserializer = predictions: ^ipc -- 4. Render reproducible Quarto report report = node script = "src/report.qmd", runtime = Quarto } build pipeline p ggplot object directly; T’s runner automatically renders and caches the visual artifact without ggsave . DataFrames pass between nodes via Apache Arrow IPC ^ipc without read.csv or to csv glue. <{ ... } blocks. Nodes accept external script files directly jln script = "sim.jl" , pyn script = "train.py" , rn script = "plot.R" . Your Julia, Python, and R scripts remain ordinary standalone files that your team can run or reuse anywhere with standard tooling. raise , R stop , Julia error , or T VError artifact, and allows independent branches to complete. Downstream nodes can inspect the error with read node or explain , or recover programmatically. src/report.qmd into an HTML or PDF report inside the Nix sandbox, directly embedding upstream metrics and figures. Most modern quantitative projects in research, central banks, official statistics, and regulated industries are polyglot by necessity: - Julia is unmatched for raw numerical simulation, ODEs, and heavy optimization loops. - Python is the standard for modern machine learning and deep learning tooling. - R remains the gold standard for survey statistics, econometrics, and publication-ready reporting. Connecting them today forces you to choose between three bad options: | The Status Quo | The Failure Mode | |---|---| | In-process FFI reticulate , PyCall , RCall | Shared memory between multiple runtimes with competing garbage collectors and conflicting OpenMP/BLAS threads causes unexplained segfaults. Upgrading one runtime breaks the other. | | Ad-hoc Bash scripts & CSVs | No caching: tweaking a title in an R ggplot re-runs your 3-hour Julia simulation. Column types and missing values silently mutate during CSV export. | | Chained Docker containers | Huge container images, slow local development, impossible for an analyst to inspect or debug interactively on a laptop. | ^onnx , ^pmml , ^csv . No custom serialization glue scripts. | Feature | {targets} | {rixpress} | Snakemake | Docker packaging only | T | |---|---|---|---|---|---| | Interface & Engine | R package targets.R , host environment | R package API, Nix engine | Python / CLI DSL, Conda/host | Container image, Docker daemon | Dedicated pipeline language, Nix engine | | Cross-language seam | R-native polyglot is bolted on | R-native Python nodes via rixpress helpers | Shell scripts & CLI wrappers | Manual entrypoints & volume mounts | Process-isolated IPC across R, Python, and Julia | | Intermediate I/O | Automatic | Automatic | Manual file paths | Manual volumes & files | Automatic zero-boilerplate boundary transfer | | Node caching | Content-addressed R | Content-addressed R | Timestamp / file hash | Docker build layer cache | Content-addressed all nodes | | System library locking | ❌ Delegates to host | ✅ Hermetic Nix | ⚠️ Optional Conda | ✅ Per image | ✅ Hermetic per-node Nix sandbox | | Interactive inspection | ✅ tar read | ✅ read node | ⚠️ File inspect only | ❌ Container attach | ✅ read node , explain | | Error resilience | ❌ Aborts run | ❌ Aborts run | ❌ Aborts run | ❌ Container exits | ✅ First-class polyglot soft-failures | When you define a node using node , rn R , pyn Python , jln Julia , or shn Shell , T treats the result as a first-class Node object. These objects transition through two main states: build pipeline , the node points to a concrete, immutable artifact in the Nix store. When you call read node p.node name in the REPL, T looks at the node’s serializer and attempts to automatically load the data back into the T environment: | Serializer | Resulting T Type | Backend | |---|---|---| | default / serialize | Varies | Native T binary serialization | | arrow | DataFrame | Apache Arrow IPC zero-copy | | csv | DataFrame | Native CSV parser | | json | Dict / List | JSON parser | | pmml | Model | Native T model evaluator | You can use explain to look inside a built node: -- Example: Inspecting a built R node model node = p.model r explain model node { kind : "computed node", name : "model r", runtime : "R", path : "/nix/store/...-model r/artifact", serializer : "pmml", class : "lm", dependencies : "data" } The path field is the escape hatch: it gives you the absolute path to the node’s output in the Nix store. You can inspect the artifact directly or pass it to external tools. t update , and build a hello-world pipeline fct helpers | forwarding semantics and short-circuiting ^serializer system for data interchange and materialization % shortcuts %cd , %env , and more