{"slug": "t-reproducible-pipelines-for-polyglot-data-science", "title": "T – Reproducible Pipelines for Polyglot Data Science", "summary": "T version 0.55.5 \"L'Ultime combat\" has been released, offering a hermetic dependency graph that orchestrates Julia, Python, and R pipelines through Nix and passes DataFrames between nodes via Apache Arrow IPC. The tool, installable via `nix shell --accept-flake-config github:b-rodrigues/tlang`, lets users scaffold a project with `t init --project my_project` and run a demo showing pipeline introspection, caching, and first-class error handling. T's pipeline syntax supports inline code blocks and external script files, with a Quarto node rendering reproducible reports.", "body_md": "**Use Julia, Python, and R for what they’re really good at —\nwhatever that is for you. T orchestrates them.**\n\nSimulations in Julia, ML in Python, statistics in R — or the exact opposite. It doesn’t matter how you divide the labor: the hard part of polyglot data science was never the languages, it was the fragile seam between them.\n\nA language for the LLM era, T is designed to be piloted by both humans and AI models. It gives you one hermetic dependency graph where your tools communicate without glue and execute consistently through space and time: on your laptop today, on a cluster tomorrow, and five years from now without bitrot.\n\n**Status:** Version 0.55.5 “L’Ultime combat”.\n\n**[Install Nix](nix-installation.html)**\n(installs Nix and configures the `rstats-on-nix` cache in one\nstep):\n\n```\ncurl --proto '=https' --tlsv1.2 -sSf -L https://install.determinate.systems/nix | \\\n  sh -s -- install --no-confirm --extra-conf \"\ntrusted-users = root $USER\nsubstituters = https://cache.nixos.org https://rstats-on-nix.cachix.org\ntrusted-public-keys = cache.nixos.org-1:6NCHdD59X431o0gWypbMrAURkbJ16ZPMQFGspcDShjY= rstats-on-nix.cachix.org-1:vdiiVgocg6WeJrODIqdprZRUrhi1JzhBnXv7aWI6+F0=\"\n```\n\n**Try T immediately** in an ephemeral shell:\n\n```\nnix shell --accept-flake-config github:b-rodrigues/tlang\n```\n\n**Scaffold a project** and enter its pinned\nenvironment:\n\n```\nt init --project my_project && cd my_project && nix develop\n```\n\n*(See the [Nix Installation\nGuide](nix-installation.html) and [Getting Started\nTutorial](getting-started.html) for full platform instructions).*\n\nRun `t demo` right in your terminal to see pipeline\nintrospection, hermetic Nix builds, Arrow in-memory inspection, caching,\nand first-class error handling in action. The demo builds in a scratch\ndirectory under your current directory (so it resolves your project’s\nflake) and removes it on exit:\n\nA complete analysis that simulates non-linear data in Julia, fits a gradient-boosted regressor in Python, plots ground truth vs predictions in R, and compiles a Quarto report:\n\n```\np = pipeline {\n  -- 1. Simulate non-linear DGP in Julia (seeded)\n  sim_data = jln(\n    command = <{\n      using Random, DataFrames\n      Random.seed!(42)\n\n      t = 1:500\n      shock = cumsum(randn(500))\n      DataFrame(time = t, shock = shock, signal = sin.(t ./ 20) .+ shock .* 0.2)\n    }>,\n    serializer = ^ipc\n  )\n\n  -- 2. Train non-linear model & predict in Python (scikit-learn)\n  predictions = pyn(\n    command = <{\nfrom sklearn.ensemble import HistGradientBoostingRegressor\n\nX = sim_data[['time', 'shock']]\ny = sim_data['signal']\nmodel = HistGradientBoostingRegressor(random_state=42).fit(X, y)\nsim_data['pred'] = model.predict(X)\nsim_data\n    }>,\n    deserializer = [sim_data: ^ipc],\n    serializer = ^ipc\n  )\n\n  -- 3. Publication figure in R (ggplot2)\n  plot = rn(\n    command = <{\n      library(ggplot2)\n\n      ggplot(predictions, aes(x = time)) +\n        geom_point(aes(y = signal), alpha = 0.3, color = \"#7f8c8d\") +\n        geom_line(aes(y = pred), color = \"#e74c3c\", linewidth = 1) +\n        labs(title = \"Julia Simulation + Python ML Predictions\", y = \"Value\") +\n        theme_minimal()\n    }>,\n    deserializer = [predictions: ^ipc]\n  )\n\n  -- 4. Render reproducible Quarto report\n  report = node(script = \"src/report.qmd\", runtime = Quarto)\n}\n\nbuild_pipeline(p)\n```\n\n`ggplot`\nobject directly; T’s runner automatically renders and caches the visual\nartifact without `ggsave()`. DataFrames pass between nodes\nvia Apache Arrow IPC (`^ipc`) without `read.csv()`\nor `to_csv()` glue.`<{ ... }>` blocks. Nodes\naccept external script files directly\n(`jln(script = \"sim.jl\")`,\n`pyn(script = \"train.py\")`,\n`rn(script = \"plot.R\")`). Your Julia, Python, and R scripts\nremain ordinary standalone files that your team can run or reuse\nanywhere with standard tooling.`raise`), R (` stop()`), Julia\n(`error()`), or T (` VError` artifact, and allows\nindependent branches to complete. Downstream nodes can inspect the error\nwith `read_node()` or `explain()`, or recover\nprogrammatically.`src/report.qmd` into an HTML or PDF report inside the Nix\nsandbox, directly embedding upstream metrics and figures.\nMost modern quantitative projects in research, central banks,\nofficial statistics, and regulated industries are polyglot by necessity:\n- **Julia** is unmatched for raw numerical simulation,\nODEs, and heavy optimization loops. - **Python** is the\nstandard for modern machine learning and deep learning tooling. -\n**R** remains the gold standard for survey statistics,\neconometrics, and publication-ready reporting.\n\nConnecting them today forces you to choose between three bad options:\n\n| The Status Quo | The Failure Mode | \n|---|---|\n| **In-process FFI (`reticulate`, `PyCall`, `RCall`)** | Shared memory between multiple runtimes with competing garbage collectors and conflicting OpenMP/BLAS threads causes unexplained segfaults. Upgrading one runtime breaks the other. | \n| **Ad-hoc Bash scripts & CSVs** | No caching: tweaking a title in an R ggplot re-runs your 3-hour Julia simulation. Column types and missing values silently mutate during CSV export. | \n| **Chained Docker containers** | Huge container images, slow local development, impossible for an analyst to inspect or debug interactively on a laptop. | \n\n`^onnx`, `^pmml`,\n`^csv`). No custom serialization glue scripts.\n| Feature | {targets} | {rixpress} | Snakemake | Docker (packaging only) | **T** | \n|---|---|---|---|---|---|\n| **Interface & Engine** | R package ( `_targets.R` ), host environment | R package API, Nix engine | Python / CLI DSL, Conda/host | Container image, Docker daemon | **Dedicated pipeline language, Nix engine** | \n| **Cross-language seam** | R-native (polyglot is bolted on) | R-native (Python nodes via `rixpress` helpers) | Shell scripts & CLI wrappers | Manual entrypoints & volume mounts | **Process-isolated IPC across R, Python, and Julia** | \n| **Intermediate I/O** | Automatic | Automatic | Manual file paths | Manual volumes & files | **Automatic (zero-boilerplate boundary transfer)** | \n| **Node caching** | Content-addressed (R) | Content-addressed (R) | Timestamp / file hash | Docker build layer cache | **Content-addressed (all nodes)** | \n| **System library locking** | ❌ (Delegates to host) | ✅ (Hermetic Nix) | ⚠️ (Optional Conda) | ✅ (Per image) | ✅ (Hermetic per-node Nix sandbox) | \n| **Interactive inspection** | ✅ ( `tar_read()` ) | ✅ ( `read_node()` ) | ⚠️ (File inspect only) | ❌ (Container attach) | ✅ ( `read_node()` ,`explain()` ) | \n| **Error resilience** | ❌ (Aborts run) | ❌ (Aborts run) | ❌ (Aborts run) | ❌ (Container exits) | ✅ (First-class polyglot soft-failures) | \n\nWhen you define a node using `node()`, `rn()`\n(R), `pyn()` (Python), `jln()` (Julia), or\n`shn()` (Shell), T treats the result as a first-class\n**Node** object. These objects transition through two main\nstates:\n\n`build_pipeline()`,\nthe node points to a concrete, immutable artifact in the Nix store.\nWhen you call `read_node(p.node_name)` in the REPL, T\nlooks at the node’s **serializer** and attempts to\nautomatically load the data back into the T environment:\n\n| Serializer | Resulting T Type | Backend | \n|---|---|---|\n| `default` /`serialize` | Varies | Native T binary serialization | \n| `arrow` | `DataFrame` | Apache Arrow IPC (zero-copy) | \n| `csv` | `DataFrame` | Native CSV parser | \n| `json` | `Dict` /`List` | JSON parser | \n| `pmml` | `Model` | Native T model evaluator | \n\nYou can use `explain()` to look inside a built node:\n\n```\n-- Example: Inspecting a built R node\n> model_node = p.model_r\n> explain(model_node)\n{\n  `kind`: \"computed_node\",\n  `name`: \"model_r\",\n  `runtime`: \"R\",\n  `path`: \"/nix/store/...-model_r/artifact\",\n  `serializer`: \"pmml\",\n  `class`: \"lm\",\n  `dependencies`: [\"data\"]\n}\n```\n\nThe `path` field is the escape hatch: it gives you the\nabsolute path to the node’s output in the Nix store. You can inspect the\nartifact directly or pass it to external tools.\n\n`t update`, and build a\nhello-world pipeline`fct_*` helpers`|>`\nforwarding semantics and short-circuiting`^serializer` system for data interchange and\nmaterialization`%` shortcuts (`%cd`, `%env`, and\nmore)", "url": "https://wpnews.pro/news/t-reproducible-pipelines-for-polyglot-data-science", "canonical_source": "https://tstats-project.org/", "published_at": "2026-10-05 23:58:00+00:00", "updated_at": "2026-10-06 00:18:37.735378+00:00", "lang": "en", "topics": ["developer-tools", "mlops", "ai-tools"], "entities": ["T", "Julia", "Python", "R", "Nix", "Apache Arrow", "Quarto", "scikit-learn"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/t-reproducible-pipelines-for-polyglot-data-science", "markdown": "https://wpnews.pro/news/t-reproducible-pipelines-for-polyglot-data-science.md", "text": "https://wpnews.pro/news/t-reproducible-pipelines-for-polyglot-data-science.txt", "jsonld": "https://wpnews.pro/news/t-reproducible-pipelines-for-polyglot-data-science.jsonld"}}