# Stop Writing the Same Pandas Boilerplate: How We Built a Visual Pipeline Studio for ML Preprocessing

> Source: <https://dev.to/harieshkai/stop-writing-the-same-pandas-boilerplate-how-we-built-a-visual-pipeline-studio-for-ml-22ff>
> Published: 2026-10-09 02:09:08+00:00

Hey everyone,

Every time I start a new machine learning project or Kaggle competition, I end up spending the first hour doing the exact same chores:

- handling null values and missing label rows
- writing manual IQR outlier clipping formulas
- coercing corrupt object columns to numeric/datetimes
- one-hot / label encoding categories
- trimming invisible whitespace

I built DataForger to make preprocessing fast, modular, and visual:

What it does:

1. Drag-and-Drop Ingestion: Instant health check, missing value diagnostics, and distributions.
2. Visual Node Pipeline: Drag, connect, and reorder transformation stages on a React Flow graph canvas.
3. Live Diff Previews: Inspect exactly what changed before and after applying an operation.
4. Export Clean Artifacts: Download the cleaned .csv + the reproducible JSON pipeline configuration to use in your training code.
5. Multimodal Ready: Built for Tabular, NLP text cleaning, image sizing/normalization, and time-series alignment.

Early Access & Feedback:

I'm rolling this out as a Founding Member Early Access Pass for ₹299 INR (~$3.50 USD one-time lifetime) for the first 50 users before we switch to a subscription model.

Try it out here: [[https://shirogani-dataforger-beta.vercel.app](https://shirogani-dataforger-beta.vercel.app)]

I really want your raw feedback:

- What's the most annoying preprocessing step in your daily workflow that you wish was automated?
- Does the node graph feel snappy or would you prefer a linear list view?

I'll be in the comments all day to answer questions and patch any bugs you find!
