What if AI didn't just explain your notes, but actually figured out what you should study next?
This is my submission for the Hacktoberfest Weekend Challenge: Build for a Friend.
I built PrepFlow because students don't usually have a shortage of study material.
They have the opposite problem.
Too much of it.
PDFs. Notes. Syllabi. Previous-year questions. Multiple subjects. Weak topics. Different confidence levels. And a deadline that keeps getting closer.
The difficult question isn't:
“What does this PDF say?”
It's:
“Given everything I need to learn, how should I spend the limited time I have left?”
That's the problem PrepFlow tries to solve.
PrepFlow takes this:
BEFORE PREPFLOW
┌─────────┐ ┌─────────┐ ┌─────────┐
│ PDFs │ │ Notes │ │ Syllabus│
└────┬────┘ └────┬────┘ └────┬────┘
│ │ │
└─────────────┼─────────────┘
│
▼
┌───────────────┐
│ STUDENT │
│ │
│ "What do I │
│ study today?" │
└───────────────┘
and turns it into:
WITH PREPFLOW
PDF / TEXT
│
▼
┌──────────────┐
│ EXTRACT │
│ page-aware │
│ text │
└──────┬───────┘
▼
┌──────────────┐
│ CHUNK │
│ deterministic│
│ boundaries │
└──────┬───────┘
▼
┌──────────────┐
│ GEMMA │
│ understand │
│ material │
└──────┬───────┘
▼
┌──────────────┐
│ VERIFY │
│ evidence + │
│ structured │
│ output │
└──────┬───────┘
▼
┌──────────────┐
│ PRIORITIZE │
│ importance │
│ weakness │
│ urgency │
│ difficulty │
└──────┬───────┘
▼
┌──────────────┐
│ PLAN │
│ capacity + │
│ prerequisites│
│ revision │
└──────┬───────┘
▼
┌──────────────────────────┐
│ TODAY │
│ │
│ 1. Dynamic Programming │
│ 2. Graph Traversal │
│ 3. Revise Greedy │
│ │
│ 145 / 160 min planned │
└──────────────────────────┘
The important part is that AI doesn't generate the final timetable.
That distinction shaped almost the entire architecture.
I split the system into two responsibilities:
| AI should do | Software should do |
|---|---|
| Understand unstructured material | Calculate available time |
| Identify topics | Calculate priority |
| Identify subtopics | Enforce capacity |
| Suggest prerequisites | Validate evidence |
| Estimate semantic difficulty | Persist state |
| Connect topics to source evidence | Schedule tasks |
| Interpret weak-topic hints | Schedule revision |
| Produce structured analysis | Handle progress |
| Recover missed work |
Because I don't want an LLM deciding whether:
“You have 120 minutes available, so I'll give you 157 minutes of work.”
That's not intelligence.
That's a bug.
So PrepFlow follows a simple rule:
AI interprets. Deterministic software decides.
The repository currently implements the complete path from material ingestion through planning, daily execution, revision, and progress.
┌─────────────────────┐
│ STUDY SETUP │
│ │
│ Exam date │
│ Hours/day │
│ Study days │
│ Weak topics │
│ Confidence │
│ Buffer │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ MATERIAL INGESTION │
│ │
│ PDF / pasted text │
│ Page preservation │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ DETERMINISTIC │
│ CHUNKING │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ GEMMA │
│ │
│ Topics │
│ Subtopics │
│ Difficulty │
│ Importance │
│ Prerequisites │
│ Evidence │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ VALIDATION │
│ │
│ JSON → Zod │
│ Evidence matching │
│ Page derivation │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ KNOWLEDGE MODEL │
│ │
│ Subjects │
│ Topics │
│ Prerequisites │
│ Evidence │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ PRIORITY ENGINE │
│ │
│ Importance │
│ Weakness │
│ Foundation │
│ Difficulty │
│ Urgency │
│ Emphasis │
│ Evidence │
│ Revision │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ CAPACITY ENGINE │
│ │
│ Available minutes │
│ Buffer │
│ Revision reserve │
│ Exam deadline │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ PLANNER │
│ │
│ MUST │
│ SHOULD │
│ IF TIME │
│ DEFER │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ DAILY EXECUTION │
│ │
│ Start │
│ Complete │
│ Skip │
│ Progress │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ REVISION + RECOVERY │
│ │
│ +1 / +3 / +7 │
│ Missed tasks │
│ Rebalance │
│ Replan │
└─────────────────────┘
There are thousands of projects that can answer:
“Summarize chapter 4.”
That wasn't the product I wanted to build.
PrepFlow's model produces structured knowledge, not the final answer the student follows.
The application then transforms that knowledge into an executable workflow.
GENERIC PDF CHAT
PDF ──────► LLM ──────► Answer
PREPFLOW
PDF
│
▼
Extraction
│
▼
Chunking
│
▼
Gemma
│
▼
Structured Topics
│
▼
Evidence Verification
│
▼
Priority Engine
│
▼
Capacity Engine
│
▼
Planner
│
▼
Daily Tasks
│
▼
Progress
│
▼
Revision
The difference is subtle but important:
The output isn't text.
The output is a decision system.
One of the parts I cared about most was preventing the model from inventing study topics.
If Gemma says:
“Dynamic Programming is highly important.”
PrepFlow doesn't simply trust it.
The model must provide source evidence.
The system then checks that evidence against the extracted source text and derives the PDF page programmatically.
Gemma says:
Topic:
Dynamic Programming
Evidence:
"Dynamic programming solves problems
by combining solutions to overlapping
subproblems."
│
▼
SOURCE MATERIAL
│
▼
Does this exact evidence
exist in the source?
┌────┴────┐
YES NO
│ │
▼ ▼
ACCEPT REJECT
│
▼
derive PDF page
This gives each topic a traceable relationship back to the material.
The repository documents the same architecture: model output is parsed, schema-validated, then evidence-verified before topics are accepted.
Not every topic deserves equal time.
PrepFlow combines multiple bounded factors:
| Factor | What it represents |
|---|---|
| Importance | How important the topic is |
| Weakness | How weak the student is |
| Foundation | Whether other topics depend on it |
| Difficulty | How demanding it is |
| Urgency | How close the exam is |
| Emphasis | Explicit emphasis in the source |
| Evidence | Strength of supporting evidence |
| Revision | Need for later review |
This produces a bounded priority score rather than asking the model to invent one.
TOPIC PRIORITY
Importance ────────┐
Weakness ──────────┤
Foundation ────────┤
Difficulty ────────┤
Urgency ───────────┤
Emphasis ──────────┼──► PRIORITY SCORE ──► PLAN
Evidence ──────────┤
Revision ──────────┘
This was one of the most important product decisions.
Suppose the student has:
40.5 hours of estimated work
but only:
24 hours available
A typical AI timetable might confidently distribute everything across the available days.
PrepFlow doesn't.
It tells the truth.
ESTIMATED WORK
████████████████████████████████████████ 40.5h
AVAILABLE TIME
████████████████████████ 24.0h
───────────────
16.5h OVERLOAD
Then it prioritizes the work:
| Tier | Meaning |
|---|---|
| 🔴 MUST | Highest-value work to protect |
| 🟠 SHOULD | Important if capacity permits |
| 🟡 IF TIME | Useful but lower priority |
| ⚪ DEFER | Cannot realistically fit |
The planner therefore answers two questions:
That second question is surprisingly important.
A study plan shouldn't tell someone to learn:
“Advanced Dynamic Programming”
before:
“Dynamic Programming fundamentals.”
PrepFlow models prerequisite relationships.
Arrays
│
▼
Recursion
│
▼
Dynamic Programming
│
├──────────────► Knapsack
│
└──────────────► Longest Common Subsequence
The planner can therefore account for foundational topics instead of treating every topic as an independent checkbox.
Real students miss tasks.
That's normal.
So PrepFlow treats planning as a feedback loop:
┌──────────────┐
│ INITIAL PLAN │
└──────┬───────┘
│
▼
┌──────────────┐
│ TODAY │
└──────┬───────┘
│
┌────────┼────────┐
▼ ▼ ▼
COMPLETE SKIP MISS
│ │ │
└────────┴────────┘
│
▼
┌─────────────────┐
│ RECOVERY / │
│ REBALANCE │
└────────┬────────┘
│
▼
NEW PLAN VERSION
Completed work is preserved.
Unfinished work can be recovered or re-planned.
This is much closer to how real studying works than generating one static timetable.
Studying a topic once isn't enough.
PrepFlow schedules revision sessions after learning using spaced intervals.
Conceptually:
LEARN
│
├──── +1 study day ────► REVISION 1
│
├──── +3 study days ────► REVISION 2
│
└──── +7 study days ────► REVISION 3
The planner also respects the exam boundary and available study days.
The repository is a TypeScript monorepo containing the React frontend, Express/TypeScript API, and shared schemas/types.
PrepFlow/
│
├── apps/
│ ├── api/
│ │ ├── domain/
│ │ │ ├── planning/
│ │ │ └── analysis/
│ │ │
│ │ ├── infrastructure/
│ │ │ ├── prisma/
│ │ │ ├── PDF extraction
│ │ │ └── AI providers
│ │ │
│ │ └── HTTP API
│ │
│ └── web/
│ ├── React
│ ├── TypeScript
│ ├── Vite
│ └── Tailwind
│
├── packages/
│ └── shared/
│ └── Zod schemas + shared types
│
├── docs/
│ ├── analysis.md
│ ├── planning.md
│ └── decisions/
│
└── PostgreSQL + Prisma
I didn't want the entire application to become dependent on one model.
The architecture therefore separates:
APPLICATION
│
▼
┌─────────────┐
│ AI PROVIDER │
│ INTERFACE │
└──────┬──────┘
│
┌─────────┼─────────┐
▼ ▼ ▼
Ollama Gemini Fake
+ Gemma + Gemma Provider
The model can change without rewriting the planning engine, database layer, or frontend.
The repository currently exposes ollama, gemini, and fake provider modes.
The system treats uploaded study material as untrusted input.
USER MATERIAL
│
│ untrusted
▼
┌──────────────┐
│ nonce + │
│ delimiters │
└──────┬───────┘
▼
GEMMA
│
▼
┌──────────────┐
│ tolerant │
│ JSON parsing │
└──────┬───────┘
▼
┌──────────────┐
│ Zod schema │
│ validation │
└──────┬───────┘
▼
┌──────────────┐
│ source │
│ evidence │
│ verification │
└──────┬───────┘
▼
ACCEPT
The model receives no application secrets or tools.
Its output is never treated as trusted application state.
The repository also documents bounded retries, analysis concurrency limits, upload limits, CORS controls, and environment-based secrets.
I didn't want the project to only work in a happy-path demo.
The current test suite covers multiple layers:
| Layer | What is tested |
|---|---|
| Shared | Schemas and shared logic |
| API | Business logic and routes |
| Web | Frontend behavior |
| HTTP | End-to-end API routing |
| Planning | Priority, capacity, scheduling |
| Randomized invariants | Planner safety properties |
| PostgreSQL | Real persistence |
| Concurrency | Rebalance/version behavior |
| PDF extraction | Real document parsing |
Current validation:
and
The integration tests include persistence and concurrent plan-rebalancing scenarios.
AUTOMATED TESTS
Shared ████████████████████ 22
API ████████████████████ 358
Web ████████████████████ 43
───
TOTAL 423
Real database integration:
PostgreSQL integration
Passed ████████████████████ 20
Failed 0
I also validated the PDF extraction pipeline against a real academic PDF and compared the extracted page structure against the source.
| Layer | Technology | Why |
|---|---|---|
| Frontend | React + TypeScript | Component-based UI |
| Build | Vite | Fast development/build |
| Styling | Tailwind CSS | Consistent UI system |
| Backend | Node.js + Express | Lightweight API |
| Language | TypeScript | Shared type safety |
| Validation | Zod | Runtime contracts |
| Database | PostgreSQL | Relational persistence |
| ORM | Prisma | Type-safe DB access |
| PDF.js | Page-aware extraction | |
| AI | Gemma | Open-weight material analysis |
| Local AI | Ollama | Local model execution |
| Testing | Vitest | Fast automated testing |
| CI | GitHub Actions | Automated verification |
Study material can be surprisingly sensitive.
It can contain:
That's why I wanted PrepFlow's architecture to support local/open-weight AI, rather than making a cloud LLM the only possible path.
Gemma sits behind an AI-provider abstraction instead of being hardcoded into the application's business logic.
That gives the project a useful property:
The intelligence can evolve without rebuilding the product around a single AI provider.
| PrepFlow is | PrepFlow isn't |
|---|---|
| AI-assisted planning | A generic chatbot |
| Evidence-backed analysis | Blind LLM output |
| Deterministic scheduling | LLM-generated timetable |
| Capacity-aware | “Everything fits” fantasy |
| Revision-aware | One-time checklist |
| Explainable | Black-box prioritization |
| Open-weight AI compatible | Locked to one provider |
| Built around execution | Just another note summarizer |
The biggest thing I learned wasn't how to connect an AI model to a backend.
It was learning where not to use AI.
A tempting architecture would have been:
PDF
│
▼
LLM
│
▼
"Here is your study plan."
It's easy.
It's also difficult to trust.
PrepFlow instead looks more like:
AI
│
▼
UNDERSTAND THE MATERIAL
│
▼
STRUCTURED DATA
│
▼
┌────────────────────┐
│ DETERMINISTIC CODE │
│ │
│ Validate │
│ Prioritize │
│ Calculate │
│ Schedule │
│ Revise │
│ Recover │
└─────────┬──────────┘
│
▼
REAL PLAN
That separation gives me much more confidence in the system.
PrepFlow is intentionally an MVP.
The next areas I'd explore are:
| Area | Possible improvement |
|---|---|
| Documents | OCR for scanned PDFs |
| AI | More provider/model evaluation |
| Knowledge | Better prerequisite inference |
| Planning | Calendar-aware scheduling |
| Progress | More sophisticated learning models |
| AI runtime | Additional local/open-weight providers |
| UX | Topic editing and plan refinement |
| Validation | Larger real-student testing |
The current repository roadmap already separates the completed repository/database, ingestion, AI analysis, and planning milestones from later topic-review, polish, deployment, and real-user testing.
github.com/piyusshhjangid/PrepFlow
npm install
npm run db:migrate
npm run db:seed
npm run dev:api
npm run dev:web
Then open the web application.
The repository also documents the local Gemma/Ollama and hosted Gemma provider setup.
Building for a friend changes the question.
Instead of:
“What AI feature can I add?”
you start asking:
“What is actually making this person's life harder?”
For PrepFlow, the answer was:
Students don't necessarily need more study material.
They need help turning the material they already have into a realistic sequence of actions.
So that's what I built.
Not another chatbot.
Not another PDF summarizer.
A system that tries to answer one deceptively difficult question:
Built for the Hacktoberfest Weekend Challenge — Build for a Friend.