Hey everyone,
Every time I start a new machine learning project or Kaggle competition, I end up spending the first hour doing the exact same chores:
- handling null values and missing label rows
- writing manual IQR outlier clipping formulas
- coercing corrupt object columns to numeric/datetimes
- one-hot / label encoding categories
- trimming invisible whitespace
I built DataForger to make preprocessing fast, modular, and visual:
What it does:
- Drag-and-Drop Ingestion: Instant health check, missing value diagnostics, and distributions.
- Visual Node Pipeline: Drag, connect, and reorder transformation stages on a React Flow graph canvas.
- Live Diff Previews: Inspect exactly what changed before and after applying an operation.
- Export Clean Artifacts: Download the cleaned .csv + the reproducible JSON pipeline configuration to use in your training code.
- Multimodal Ready: Built for Tabular, NLP text cleaning, image sizing/normalization, and time-series alignment.
Early Access & Feedback:
I'm rolling this out as a Founding Member Early Access Pass for ₹299 INR (~$3.50 USD one-time lifetime) for the first 50 users before we switch to a subscription model.
Try it out here: [https://shirogani-dataforger-beta.vercel.app] I really want your raw feedback:
- What's the most annoying preprocessing step in your daily workflow that you wish was automated?
- Does the node graph feel snappy or would you prefer a linear list view?
I'll be in the comments all day to answer questions and patch any bugs you find!