# FolderPilot: A Local AI Folder Organizer That Is Built Never to Delete a File

> Source: <https://dev.to/tejasrawool186/folderpilot-a-local-ai-folder-organizer-that-is-built-never-to-delete-a-file-180m>
> Published: 2026-10-04 17:34:47+00:00

*This is a submission for the [Hacktoberfest Weekend Challenge: Build for a Friend](https://dev.to/challenges/hacktoberfest-weekend-2026-10-01)*

FolderPilot is a local web app that scans any folder, reads what is inside the files, and proposes a cleaner structure. It shows the result as an interactive, color-coded map. You review the plan, edit it, and approve it. Nothing changes on disk until you do.

I built it for a friend whose folders hold exactly the kind of files you do not want to upload anywhere: ID scans, marksheets, resumes, offer letters, and assignments, all mixed together with installers and screenshots. Cloud organizers want those documents on someone else's server. Cleanup scripts delete things. I wanted a tool that does neither.

Three rules shaped the whole project:

`127.0.0.1` with an open-weight model served by Ollama.

The video walks through scanning a real folder, exploring the treemap and sunburst views, reviewing the before and after plan, applying it, and rolling it back. It also shows the folder chat refusing a request to delete files.

**Organize anything. Delete nothing.**

Privacy-first, local folder organizer with interactive visual analytics,

append-only rollback journaling, and on-device AI.

[Overview](https://github.com/TejasRawool186/FolderPilot#why-this-exists) · [How it works](https://github.com/TejasRawool186/FolderPilot#how-it-works) · [Features](https://github.com/TejasRawool186/FolderPilot#features) · [Safety model](https://github.com/TejasRawool186/FolderPilot#safety-model) · [Quick start](https://github.com/TejasRawool186/FolderPilot#quick-start)

Personal directories such as Downloads, Desktop, and work folders routinely accumulate sensitive files—including tax returns, identity cards, college transcripts, and resumes. Most of these files remain disorganized because manual sorting is slow and tedious.

Typical cloud organizers require uploading private documents to external servers. Other utility scripts use aggressive delete operations that risk permanent data loss.

FolderPilot runs entirely on local hardware, inspects documents privately on `127.0.0.1`, and guarantees that no file can ever be deleted.

Running open-weight language models locally on consumer hardware changes the economics and security of desktop file management:

| Dimension | Local Open-Source AI (FolderPilot) | Cloud AI APIs | 
|---|---|---|
| **Privacy** | Local-only. File contents, extracted text, and metadata never |  | 

The repository has a FastAPI backend, a Next.js frontend, a pytest suite for the safety guarantees, and docs covering the architecture, the AI integration, and the API.

A Downloads folder is where files go to pile up. After a year it holds duplicate resumes (`resume_final`, `resume_final2`), screenshots named `IMG_2938.png`, installers you already ran, and a few documents you would be uncomfortable seeing leaked.

Existing options each fail in a different way:

Make the AI the last resort, and make the file system layer safe by construction.

Most files do not need a language model. An extension, a hash, or a keyword count settles them. So FolderPilot runs cheap checks first and only sends the genuinely ambiguous files to a small local model. That keeps it usable on a laptop with 8 GB of RAM and no GPU, which is the machine I developed it on.

On the safety side, I did not add a "confirm delete" dialog. I removed deletion from the program. The file operations module only exposes `mkdir`, `move`, and `rename`.

The scanner walks the chosen folder and streams progress to the UI. Each file then goes through four tiers, stopping at the first one that is confident.

``` php
flowchart TD
    A["Select any folder"] --> B["Background scan"]
    B --> C["Tier 1: extension and filename rules"]
    C --> D["Tier 2: SHA-256 duplicate detection"]
    D --> E["Tier 3: keyword prototypes on extracted text"]
    E --> F["Tier 4: local LLM, JSON schema output"]
    F --> G["Draft plan, unapproved"]
    G --> H["User reviews, edits, approves"]
    H --> I["Journal entry written"]
    I --> J["mkdir / move / rename"]
    J --> K["Undo: single op, batch, or full workspace"]
```

| Tier | Mechanism | Handles | 
|---|---|---|
| 1. Rules | Extension taxonomy and filename regex | Code, archives, media, installers, screenshots | 
| 2. Hashes | Lazy SHA-256 on files with equal sizes | Exact duplicates and repeated downloads | 
| 3. Prototypes | Keyword frequency in the first ~500 extracted tokens | Assignments, invoices, offer letters | 
| 4. Local LLM | Ollama with constrained JSON output | Ambiguous PDFs and poorly named files | 

Every file gets a category and a reason, so the review screen can show why it was placed where it was.

| Layer | Technology | Why | 
|---|---|---|
| Frontend | Next.js 16, React 19, Tailwind, D3.js | D3 gives full control over the treemap, sunburst, and tree comparison | 
| Backend | Python 3.11, FastAPI | Async routes and server-sent events for live scan progress | 
| Storage | SQLite in WAL mode | File index, plans, and the journal in one local file | 
| Extraction | PyMuPDF, python-docx, python-pptx | Reads text from PDFs, Word, and PowerPoint files | 
| Local AI | Ollama with Gemma 3 1B (default), Qwen 2.5 1.5B | Small enough for CPU-only machines | 

The frontend never talks to anything but the local backend, which binds to `127.0.0.1`.

The interface has these views:

The default model is Gemma 3 1B (`gemma3:1b`, Q4_K_M, roughly 1.2 GB of RAM), served through Ollama. Qwen 2.5 1.5B is supported as an alternative. The model can be switched from the navigation bar at runtime, through the `POST /api/ai/model` endpoint, or with the `FOLDERPILOT_OLLAMA_LLM_MODEL` environment variable.

The model does two jobs:

Why open weights mattered here:

|  | Local open model | Cloud API | 
|---|---|---|
| Privacy | File text and metadata stay on the machine | Excerpts of private documents leave the machine | 
| Cost | No per-file cost | Metered by volume | 
| Offline | Works once weights are cached | Needs a connection | 
| Model choice | Swap models in one setting | Tied to one provider | 

The tradeoff is accuracy. A 1B model will misclassify some ambiguous files, especially ones with little extractable text. That is why low-confidence files are surfaced in a review queue and why nothing is applied without approval. A user's manual category overrides are stored in local SQLite to guide later scans.

**Example 1: classification**

A file named `assignment_final.pdf` sits next to `assignment.pdf`. The extension alone says "document". Tier 2 notices the sizes differ, so they are not exact duplicates. Tier 3 reads the first page, finds terms common to coursework, and places both under a college category. The plan shows the reason on the card, and the user can drag them elsewhere if it is wrong.

**Example 2: a destructive request**

In the folder chat, I typed a request to delete all the duplicate files. The chat router intercepts deletion and removal requests before they reach the model. The reply explains that FolderPilot never deletes, and offers to move the duplicates into a `_Review_Later/` folder instead. A test (`test_chat_delete_refusal`) covers this behavior.

**Example 3: undo**

After applying a plan, the Journal view lists every operation. One click on a single operation restores that file. One click on the batch restores everything the plan touched. Because the journal entry is written before each operation, an interrupted run leaves a consistent record.

Most organizers confirm before they delete. This one cannot delete.

`safe_ops.py` has no delete, unlink, remove, or truncate path. `test_no_delete_functions_exist` checks that this stays true.`Document (1).pdf`.` C:\Windows` and `C:\Program Files`, and has tests for path traversal and symlink escapes.
I started by writing a requirements document before any code, because the safety rules had to be requirements and not afterthoughts. The repo's `docs/` folder keeps that specification along with architecture notes and a development log.

The backend came first: scanner, safe file operations, journal, and undo. I built the visual layer after that, once the plan and journal data was reliable. The default model started as Qwen 2.5 1.5B and moved to Gemma 3 1B after I tuned for memory use and speed on an 8 GB machine, with Qwen kept as a switchable option.

The tests were written around the guarantees, not the features. If a safety claim appears in the README, there is a named test behind it.

Requirements: Python 3.11+, Node.js 20+, and Ollama.

```
ollama pull gemma3:1b
cd backend && python -m pip install -r requirements.txt
python -m uvicorn app.main:app --host 127.0.0.1 --port 8000
```

In a second terminal:

```
cd frontend && npm install && npm run dev -- -p 3000
```

Then open `http://127.0.0.1:3000`. On Windows, `run.bat` starts both.

I wanted a tool I would be comfortable pointing at a folder full of someone's private documents. Keeping the model local made the privacy story simple, and removing deletion from the code made the safety story simple. A small open model turned out to be enough once it was the last step in the pipeline instead of the first.
