cd /news/ai-tools/prepflow-from-study-material-to-an-a… · home › topics › ai-tools › article
[ARTICLE · art-145230] src=dev.to ↗ pub= topic=ai-tools verified=true sentiment=↑ positive

PrepFlow — From Study Material to an Actual Study Plan

A developer built PrepFlow, a study-planning tool that ingests PDFs, notes and syllabi, uses Google's Gemma model to extract topics, subtopics, difficulty and evidence, then applies deterministic software to prioritize and schedule study tasks within a student's available time. The architecture deliberately separates AI interpretation from deterministic scheduling so the system never over-allocates time, following the rule 'AI interprets. Deterministic software decides.'

by read12 min views1 publishedOct 5, 2026

What if AI didn't just explain your notes, but actually figured out what you should study next?

This is my submission for the Hacktoberfest Weekend Challenge: Build for a Friend.

I built PrepFlow because students don't usually have a shortage of study material.

They have the opposite problem.

Too much of it.

PDFs. Notes. Syllabi. Previous-year questions. Multiple subjects. Weak topics. Different confidence levels. And a deadline that keeps getting closer.

The difficult question isn't:

“What does this PDF say?”

It's:

“Given everything I need to learn, how should I spend the limited time I have left?”

That's the problem PrepFlow tries to solve.

PrepFlow takes this:

                    BEFORE PREPFLOW

       ┌─────────┐   ┌─────────┐   ┌─────────┐
       │  PDFs   │   │  Notes  │   │  Syllabus│
       └────┬────┘   └────┬────┘   └────┬────┘
            │             │             │
            └─────────────┼─────────────┘
                          │
                          ▼
                  ┌───────────────┐
                  │     STUDENT   │
                  │               │
                  │ "What do I    │
                  │ study today?" │
                  └───────────────┘

and turns it into:

                     WITH PREPFLOW

 PDF / TEXT
     │
     ▼
┌──────────────┐
│   EXTRACT    │
│ page-aware   │
│ text         │
└──────┬───────┘
       ▼
┌──────────────┐
│    CHUNK     │
│ deterministic│
│ boundaries   │
└──────┬───────┘
       ▼
┌──────────────┐
│    GEMMA     │
│ understand   │
│ material     │
└──────┬───────┘
       ▼
┌──────────────┐
│   VERIFY     │
│ evidence +   │
│ structured   │
│ output       │
└──────┬───────┘
       ▼
┌──────────────┐
│   PRIORITIZE │
│ importance   │
│ weakness     │
│ urgency      │
│ difficulty   │
└──────┬───────┘
       ▼
┌──────────────┐
│    PLAN      │
│ capacity +   │
│ prerequisites│
│ revision     │
└──────┬───────┘
       ▼
┌──────────────────────────┐
│ TODAY                    │
│                          │
│ 1. Dynamic Programming  │
│ 2. Graph Traversal      │
│ 3. Revise Greedy        │
│                          │
│ 145 / 160 min planned   │
└──────────────────────────┘

The important part is that AI doesn't generate the final timetable.

That distinction shaped almost the entire architecture.

I split the system into two responsibilities:

AI should do Software should do
Understand unstructured material Calculate available time
Identify topics Calculate priority
Identify subtopics Enforce capacity
Suggest prerequisites Validate evidence
Estimate semantic difficulty Persist state
Connect topics to source evidence Schedule tasks
Interpret weak-topic hints Schedule revision
Produce structured analysis Handle progress
Recover missed work

Because I don't want an LLM deciding whether:

“You have 120 minutes available, so I'll give you 157 minutes of work.”

That's not intelligence.

That's a bug.

So PrepFlow follows a simple rule:

AI interprets. Deterministic software decides.

The repository currently implements the complete path from material ingestion through planning, daily execution, revision, and progress.

┌─────────────────────┐
│  STUDY SETUP        │
│                     │
│ Exam date           │
│ Hours/day            │
│ Study days           │
│ Weak topics          │
│ Confidence           │
│ Buffer               │
└──────────┬──────────┘
           │
           ▼
┌─────────────────────┐
│ MATERIAL INGESTION   │
│                     │
│ PDF / pasted text   │
│ Page preservation   │
└──────────┬──────────┘
           │
           ▼
┌─────────────────────┐
│ DETERMINISTIC       │
│ CHUNKING             │
└──────────┬──────────┘
           │
           ▼
┌─────────────────────┐
│ GEMMA               │
│                     │
│ Topics              │
│ Subtopics           │
│ Difficulty          │
│ Importance          │
│ Prerequisites       │
│ Evidence            │
└──────────┬──────────┘
           │
           ▼
┌─────────────────────┐
│ VALIDATION          │
│                     │
│ JSON → Zod          │
│ Evidence matching   │
│ Page derivation     │
└──────────┬──────────┘
           │
           ▼
┌─────────────────────┐
│ KNOWLEDGE MODEL     │
│                     │
│ Subjects            │
│ Topics              │
│ Prerequisites       │
│ Evidence            │
└──────────┬──────────┘
           │
           ▼
┌─────────────────────┐
│ PRIORITY ENGINE     │
│                     │
│ Importance          │
│ Weakness            │
│ Foundation          │
│ Difficulty          │
│ Urgency             │
│ Emphasis            │
│ Evidence            │
│ Revision            │
└──────────┬──────────┘
           │
           ▼
┌─────────────────────┐
│ CAPACITY ENGINE     │
│                     │
│ Available minutes   │
│ Buffer              │
│ Revision reserve    │
│ Exam deadline       │
└──────────┬──────────┘
           │
           ▼
┌─────────────────────┐
│ PLANNER             │
│                     │
│ MUST                │
│ SHOULD              │
│ IF TIME             │
│ DEFER               │
└──────────┬──────────┘
           │
           ▼
┌─────────────────────┐
│ DAILY EXECUTION     │
│                     │
│ Start               │
│ Complete            │
│ Skip                │
│ Progress            │
└──────────┬──────────┘
           │
           ▼
┌─────────────────────┐
│ REVISION + RECOVERY │
│                     │
│ +1 / +3 / +7        │
│ Missed tasks        │
│ Rebalance           │
│ Replan              │
└─────────────────────┘

There are thousands of projects that can answer:

“Summarize chapter 4.”

That wasn't the product I wanted to build.

PrepFlow's model produces structured knowledge, not the final answer the student follows.

The application then transforms that knowledge into an executable workflow.

                 GENERIC PDF CHAT

PDF ──────► LLM ──────► Answer

                 PREPFLOW

PDF
 │
 ▼
Extraction
 │
 ▼
Chunking
 │
 ▼
Gemma
 │
 ▼
Structured Topics
 │
 ▼
Evidence Verification
 │
 ▼
Priority Engine
 │
 ▼
Capacity Engine
 │
 ▼
Planner
 │
 ▼
Daily Tasks
 │
 ▼
Progress
 │
 ▼
Revision

The difference is subtle but important:

The output isn't text.

The output is a decision system.

One of the parts I cared about most was preventing the model from inventing study topics.

If Gemma says:

“Dynamic Programming is highly important.”

PrepFlow doesn't simply trust it.

The model must provide source evidence.

The system then checks that evidence against the extracted source text and derives the PDF page programmatically.

Gemma says:

Topic:
Dynamic Programming

Evidence:
"Dynamic programming solves problems
by combining solutions to overlapping
subproblems."

             │
             ▼

       SOURCE MATERIAL

             │
             ▼

     Does this exact evidence
       exist in the source?

        ┌────┴────┐
       YES        NO
        │          │
        ▼          ▼
    ACCEPT       REJECT
        │
        ▼
   derive PDF page

This gives each topic a traceable relationship back to the material.

The repository documents the same architecture: model output is parsed, schema-validated, then evidence-verified before topics are accepted.

Not every topic deserves equal time.

PrepFlow combines multiple bounded factors:

Factor What it represents
Importance How important the topic is
Weakness How weak the student is
Foundation Whether other topics depend on it
Difficulty How demanding it is
Urgency How close the exam is
Emphasis Explicit emphasis in the source
Evidence Strength of supporting evidence
Revision Need for later review

This produces a bounded priority score rather than asking the model to invent one.

                  TOPIC PRIORITY

Importance ────────┐
Weakness ──────────┤
Foundation ────────┤
Difficulty ────────┤
Urgency ───────────┤
Emphasis ──────────┼──► PRIORITY SCORE ──► PLAN
Evidence ──────────┤
Revision ──────────┘

This was one of the most important product decisions.

Suppose the student has:

40.5 hours of estimated work

but only:

24 hours available

A typical AI timetable might confidently distribute everything across the available days.

PrepFlow doesn't.

It tells the truth.

ESTIMATED WORK
████████████████████████████████████████ 40.5h

AVAILABLE TIME
████████████████████████ 24.0h

                         ───────────────
                         16.5h OVERLOAD

Then it prioritizes the work:

Tier Meaning
🔴 MUST Highest-value work to protect
🟠 SHOULD Important if capacity permits
🟡 IF TIME Useful but lower priority
⚪ DEFER Cannot realistically fit

The planner therefore answers two questions:

That second question is surprisingly important.

A study plan shouldn't tell someone to learn:

“Advanced Dynamic Programming”

before:

“Dynamic Programming fundamentals.”

PrepFlow models prerequisite relationships.

Arrays
  │
  ▼
Recursion
  │
  ▼
Dynamic Programming
  │
  ├──────────────► Knapsack
  │
  └──────────────► Longest Common Subsequence

The planner can therefore account for foundational topics instead of treating every topic as an independent checkbox.

Real students miss tasks.

That's normal.

So PrepFlow treats planning as a feedback loop:

          ┌──────────────┐
          │ INITIAL PLAN │
          └──────┬───────┘
                 │
                 ▼
          ┌──────────────┐
          │     TODAY    │
          └──────┬───────┘
                 │
        ┌────────┼────────┐
        ▼        ▼        ▼
     COMPLETE   SKIP    MISS
        │        │        │
        └────────┴────────┘
                 │
                 ▼
        ┌─────────────────┐
        │ RECOVERY /      │
        │ REBALANCE       │
        └────────┬────────┘
                 │
                 ▼
          NEW PLAN VERSION

Completed work is preserved.

Unfinished work can be recovered or re-planned.

This is much closer to how real studying works than generating one static timetable.

Studying a topic once isn't enough.

PrepFlow schedules revision sessions after learning using spaced intervals.

Conceptually:

LEARN
  │
  ├──── +1 study day ────► REVISION 1
  │
  ├──── +3 study days ────► REVISION 2
  │
  └──── +7 study days ────► REVISION 3

The planner also respects the exam boundary and available study days.

The repository is a TypeScript monorepo containing the React frontend, Express/TypeScript API, and shared schemas/types.

PrepFlow/
│
├── apps/
│   ├── api/
│   │   ├── domain/
│   │   │   ├── planning/
│   │   │   └── analysis/
│   │   │
│   │   ├── infrastructure/
│   │   │   ├── prisma/
│   │   │   ├── PDF extraction
│   │   │   └── AI providers
│   │   │
│   │   └── HTTP API
│   │
│   └── web/
│       ├── React
│       ├── TypeScript
│       ├── Vite
│       └── Tailwind
│
├── packages/
│   └── shared/
│       └── Zod schemas + shared types
│
├── docs/
│   ├── analysis.md
│   ├── planning.md
│   └── decisions/
│
└── PostgreSQL + Prisma

I didn't want the entire application to become dependent on one model.

The architecture therefore separates:

                APPLICATION

                    │
                    ▼
             ┌─────────────┐
             │ AI PROVIDER │
             │ INTERFACE   │
             └──────┬──────┘
                    │
          ┌─────────┼─────────┐
          ▼         ▼         ▼
       Ollama     Gemini     Fake
       + Gemma     + Gemma   Provider

The model can change without rewriting the planning engine, database layer, or frontend.

The repository currently exposes ollama, gemini, and fake provider modes.

The system treats uploaded study material as untrusted input.

USER MATERIAL
     │
     │ untrusted
     ▼
┌──────────────┐
│ nonce +      │
│ delimiters   │
└──────┬───────┘
       ▼
     GEMMA
       │
       ▼
┌──────────────┐
│ tolerant     │
│ JSON parsing │
└──────┬───────┘
       ▼
┌──────────────┐
│ Zod schema   │
│ validation   │
└──────┬───────┘
       ▼
┌──────────────┐
│ source       │
│ evidence     │
│ verification │
└──────┬───────┘
       ▼
     ACCEPT

The model receives no application secrets or tools.

Its output is never treated as trusted application state.

The repository also documents bounded retries, analysis concurrency limits, upload limits, CORS controls, and environment-based secrets.

I didn't want the project to only work in a happy-path demo.

The current test suite covers multiple layers:

Layer What is tested
Shared Schemas and shared logic
API Business logic and routes
Web Frontend behavior
HTTP End-to-end API routing
Planning Priority, capacity, scheduling
Randomized invariants Planner safety properties
PostgreSQL Real persistence
Concurrency Rebalance/version behavior
PDF extraction Real document parsing

Current validation:

and

The integration tests include persistence and concurrent plan-rebalancing scenarios.

AUTOMATED TESTS

Shared       ████████████████████  22
API          ████████████████████ 358
Web          ████████████████████ 43
                                  ───
TOTAL                              423

Real database integration:

PostgreSQL integration

Passed     ████████████████████ 20
Failed                           0

I also validated the PDF extraction pipeline against a real academic PDF and compared the extracted page structure against the source.

Layer Technology Why
Frontend React + TypeScript Component-based UI
Build Vite Fast development/build
Styling Tailwind CSS Consistent UI system
Backend Node.js + Express Lightweight API
Language TypeScript Shared type safety
Validation Zod Runtime contracts
Database PostgreSQL Relational persistence
ORM Prisma Type-safe DB access
PDF.js Page-aware extraction
AI Gemma Open-weight material analysis
Local AI Ollama Local model execution
Testing Vitest Fast automated testing
CI GitHub Actions Automated verification

Study material can be surprisingly sensitive.

It can contain:

That's why I wanted PrepFlow's architecture to support local/open-weight AI, rather than making a cloud LLM the only possible path.

Gemma sits behind an AI-provider abstraction instead of being hardcoded into the application's business logic.

That gives the project a useful property:

The intelligence can evolve without rebuilding the product around a single AI provider.

PrepFlow is PrepFlow isn't
AI-assisted planning A generic chatbot
Evidence-backed analysis Blind LLM output
Deterministic scheduling LLM-generated timetable
Capacity-aware “Everything fits” fantasy
Revision-aware One-time checklist
Explainable Black-box prioritization
Open-weight AI compatible Locked to one provider
Built around execution Just another note summarizer

The biggest thing I learned wasn't how to connect an AI model to a backend.

It was learning where not to use AI.

A tempting architecture would have been:

PDF
 │
 ▼
LLM
 │
 ▼
"Here is your study plan."

It's easy.

It's also difficult to trust.

PrepFlow instead looks more like:

                 AI
                  │
                  ▼
        UNDERSTAND THE MATERIAL
                  │
                  ▼
             STRUCTURED DATA
                  │
                  ▼
        ┌────────────────────┐
        │ DETERMINISTIC CODE │
        │                    │
        │ Validate           │
        │ Prioritize         │
        │ Calculate          │
        │ Schedule           │
        │ Revise             │
        │ Recover            │
        └─────────┬──────────┘
                  │
                  ▼
             REAL PLAN

That separation gives me much more confidence in the system.

PrepFlow is intentionally an MVP.

The next areas I'd explore are:

Area Possible improvement
Documents OCR for scanned PDFs
AI More provider/model evaluation
Knowledge Better prerequisite inference
Planning Calendar-aware scheduling
Progress More sophisticated learning models
AI runtime Additional local/open-weight providers
UX Topic editing and plan refinement
Validation Larger real-student testing

The current repository roadmap already separates the completed repository/database, ingestion, AI analysis, and planning milestones from later topic-review, polish, deployment, and real-user testing.

github.com/piyusshhjangid/PrepFlow

npm install
npm run db:migrate
npm run db:seed

npm run dev:api
npm run dev:web

Then open the web application.

The repository also documents the local Gemma/Ollama and hosted Gemma provider setup.

Building for a friend changes the question.

Instead of:

“What AI feature can I add?”

you start asking:

“What is actually making this person's life harder?”

For PrepFlow, the answer was:

Students don't necessarily need more study material.

They need help turning the material they already have into a realistic sequence of actions.

So that's what I built.

Not another chatbot.

Not another PDF summarizer.

A system that tries to answer one deceptively difficult question:

Built for the Hacktoberfest Weekend Challenge — Build for a Friend.

── more in #ai-tools 4 stories · sorted by recency
── more on @prepflow 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/prepflow-from-study-…] indexed:0 read:12min 2026-10-05 · —