{"slug": "prepflow-from-study-material-to-an-actual-study-plan", "title": "PrepFlow — From Study Material to an Actual Study Plan", "summary": "A developer built PrepFlow, a study-planning tool that ingests PDFs, notes and syllabi, uses Google's Gemma model to extract topics, subtopics, difficulty and evidence, then applies deterministic software to prioritize and schedule study tasks within a student's available time. The architecture deliberately separates AI interpretation from deterministic scheduling so the system never over-allocates time, following the rule 'AI interprets. Deterministic software decides.'", "body_md": "**What if AI didn't just explain your notes, but actually figured out what you should study next?**\n\nThis is my submission for the **Hacktoberfest Weekend Challenge: Build for a Friend**.\n\nI built **PrepFlow** because students don't usually have a shortage of study material.\n\nThey have the opposite problem.\n\n**Too much of it.**\n\nPDFs. Notes. Syllabi. Previous-year questions. Multiple subjects. Weak topics. Different confidence levels. And a deadline that keeps getting closer.\n\nThe difficult question isn't:\n\n“What does this PDF say?”\n\nIt's:\n\n**“Given everything I need to learn, how should I spend the limited time I have left?”**\n\nThat's the problem PrepFlow tries to solve.\n\nPrepFlow takes this:\n\n```\n                    BEFORE PREPFLOW\n\n       ┌─────────┐   ┌─────────┐   ┌─────────┐\n       │  PDFs   │   │  Notes  │   │  Syllabus│\n       └────┬────┘   └────┬────┘   └────┬────┘\n            │             │             │\n            └─────────────┼─────────────┘\n                          │\n                          ▼\n                  ┌───────────────┐\n                  │     STUDENT   │\n                  │               │\n                  │ \"What do I    │\n                  │ study today?\" │\n                  └───────────────┘\n```\n\nand turns it into:\n\n```\n                     WITH PREPFLOW\n\n PDF / TEXT\n     │\n     ▼\n┌──────────────┐\n│   EXTRACT    │\n│ page-aware   │\n│ text         │\n└──────┬───────┘\n       ▼\n┌──────────────┐\n│    CHUNK     │\n│ deterministic│\n│ boundaries   │\n└──────┬───────┘\n       ▼\n┌──────────────┐\n│    GEMMA     │\n│ understand   │\n│ material     │\n└──────┬───────┘\n       ▼\n┌──────────────┐\n│   VERIFY     │\n│ evidence +   │\n│ structured   │\n│ output       │\n└──────┬───────┘\n       ▼\n┌──────────────┐\n│   PRIORITIZE │\n│ importance   │\n│ weakness     │\n│ urgency      │\n│ difficulty   │\n└──────┬───────┘\n       ▼\n┌──────────────┐\n│    PLAN      │\n│ capacity +   │\n│ prerequisites│\n│ revision     │\n└──────┬───────┘\n       ▼\n┌──────────────────────────┐\n│ TODAY                    │\n│                          │\n│ 1. Dynamic Programming  │\n│ 2. Graph Traversal      │\n│ 3. Revise Greedy        │\n│                          │\n│ 145 / 160 min planned   │\n└──────────────────────────┘\n```\n\nThe important part is that **AI doesn't generate the final timetable**.\n\nThat distinction shaped almost the entire architecture.\n\nI split the system into two responsibilities:\n\n| AI should do | Software should do | \n|---|---|\n| Understand unstructured material | Calculate available time | \n| Identify topics | Calculate priority | \n| Identify subtopics | Enforce capacity | \n| Suggest prerequisites | Validate evidence | \n| Estimate semantic difficulty | Persist state | \n| Connect topics to source evidence | Schedule tasks | \n| Interpret weak-topic hints | Schedule revision | \n| Produce structured analysis | Handle progress | \n|  | Recover missed work | \n\nBecause I don't want an LLM deciding whether:\n\n“You have 120 minutes available, so I'll give you 157 minutes of work.”\n\nThat's not intelligence.\n\nThat's a bug.\n\nSo PrepFlow follows a simple rule:\n\n**AI interprets. Deterministic software decides.**\n\nThe repository currently implements the complete path from material ingestion through planning, daily execution, revision, and progress.\n\n```\n┌─────────────────────┐\n│  STUDY SETUP        │\n│                     │\n│ Exam date           │\n│ Hours/day            │\n│ Study days           │\n│ Weak topics          │\n│ Confidence           │\n│ Buffer               │\n└──────────┬──────────┘\n           │\n           ▼\n┌─────────────────────┐\n│ MATERIAL INGESTION   │\n│                     │\n│ PDF / pasted text   │\n│ Page preservation   │\n└──────────┬──────────┘\n           │\n           ▼\n┌─────────────────────┐\n│ DETERMINISTIC       │\n│ CHUNKING             │\n└──────────┬──────────┘\n           │\n           ▼\n┌─────────────────────┐\n│ GEMMA               │\n│                     │\n│ Topics              │\n│ Subtopics           │\n│ Difficulty          │\n│ Importance          │\n│ Prerequisites       │\n│ Evidence            │\n└──────────┬──────────┘\n           │\n           ▼\n┌─────────────────────┐\n│ VALIDATION          │\n│                     │\n│ JSON → Zod          │\n│ Evidence matching   │\n│ Page derivation     │\n└──────────┬──────────┘\n           │\n           ▼\n┌─────────────────────┐\n│ KNOWLEDGE MODEL     │\n│                     │\n│ Subjects            │\n│ Topics              │\n│ Prerequisites       │\n│ Evidence            │\n└──────────┬──────────┘\n           │\n           ▼\n┌─────────────────────┐\n│ PRIORITY ENGINE     │\n│                     │\n│ Importance          │\n│ Weakness            │\n│ Foundation          │\n│ Difficulty          │\n│ Urgency             │\n│ Emphasis            │\n│ Evidence            │\n│ Revision            │\n└──────────┬──────────┘\n           │\n           ▼\n┌─────────────────────┐\n│ CAPACITY ENGINE     │\n│                     │\n│ Available minutes   │\n│ Buffer              │\n│ Revision reserve    │\n│ Exam deadline       │\n└──────────┬──────────┘\n           │\n           ▼\n┌─────────────────────┐\n│ PLANNER             │\n│                     │\n│ MUST                │\n│ SHOULD              │\n│ IF TIME             │\n│ DEFER               │\n└──────────┬──────────┘\n           │\n           ▼\n┌─────────────────────┐\n│ DAILY EXECUTION     │\n│                     │\n│ Start               │\n│ Complete            │\n│ Skip                │\n│ Progress            │\n└──────────┬──────────┘\n           │\n           ▼\n┌─────────────────────┐\n│ REVISION + RECOVERY │\n│                     │\n│ +1 / +3 / +7        │\n│ Missed tasks        │\n│ Rebalance           │\n│ Replan              │\n└─────────────────────┘\n```\n\nThere are thousands of projects that can answer:\n\n“Summarize chapter 4.”\n\nThat wasn't the product I wanted to build.\n\nPrepFlow's model produces **structured knowledge**, not the final answer the student follows.\n\nThe application then transforms that knowledge into an executable workflow.\n\n```\n                 GENERIC PDF CHAT\n\nPDF ──────► LLM ──────► Answer\n\n                 PREPFLOW\n\nPDF\n │\n ▼\nExtraction\n │\n ▼\nChunking\n │\n ▼\nGemma\n │\n ▼\nStructured Topics\n │\n ▼\nEvidence Verification\n │\n ▼\nPriority Engine\n │\n ▼\nCapacity Engine\n │\n ▼\nPlanner\n │\n ▼\nDaily Tasks\n │\n ▼\nProgress\n │\n ▼\nRevision\n```\n\nThe difference is subtle but important:\n\n**The output isn't text.**\n\n**The output is a decision system.**\n\nOne of the parts I cared about most was preventing the model from inventing study topics.\n\nIf Gemma says:\n\n“Dynamic Programming is highly important.”\n\nPrepFlow doesn't simply trust it.\n\nThe model must provide source evidence.\n\nThe system then checks that evidence against the extracted source text and derives the PDF page programmatically.\n\n```\nGemma says:\n\nTopic:\nDynamic Programming\n\nEvidence:\n\"Dynamic programming solves problems\nby combining solutions to overlapping\nsubproblems.\"\n\n             │\n             ▼\n\n       SOURCE MATERIAL\n\n             │\n             ▼\n\n     Does this exact evidence\n       exist in the source?\n\n        ┌────┴────┐\n       YES        NO\n        │          │\n        ▼          ▼\n    ACCEPT       REJECT\n        │\n        ▼\n   derive PDF page\n```\n\nThis gives each topic a traceable relationship back to the material.\n\nThe repository documents the same architecture: model output is parsed, schema-validated, then evidence-verified before topics are accepted.\n\nNot every topic deserves equal time.\n\nPrepFlow combines multiple bounded factors:\n\n| Factor | What it represents | \n|---|---|\n| Importance | How important the topic is | \n| Weakness | How weak the student is | \n| Foundation | Whether other topics depend on it | \n| Difficulty | How demanding it is | \n| Urgency | How close the exam is | \n| Emphasis | Explicit emphasis in the source | \n| Evidence | Strength of supporting evidence | \n| Revision | Need for later review | \n\nThis produces a bounded priority score rather than asking the model to invent one.\n\n```\n                  TOPIC PRIORITY\n\nImportance ────────┐\nWeakness ──────────┤\nFoundation ────────┤\nDifficulty ────────┤\nUrgency ───────────┤\nEmphasis ──────────┼──► PRIORITY SCORE ──► PLAN\nEvidence ──────────┤\nRevision ──────────┘\n```\n\nThis was one of the most important product decisions.\n\nSuppose the student has:\n\n**40.5 hours of estimated work**\n\nbut only:\n\n**24 hours available**\n\nA typical AI timetable might confidently distribute everything across the available days.\n\nPrepFlow doesn't.\n\nIt tells the truth.\n\n```\nESTIMATED WORK\n████████████████████████████████████████ 40.5h\n\nAVAILABLE TIME\n████████████████████████ 24.0h\n\n                         ───────────────\n                         16.5h OVERLOAD\n```\n\nThen it prioritizes the work:\n\n| Tier | Meaning | \n|---|---|\n| 🔴 MUST | Highest-value work to protect | \n| 🟠 SHOULD | Important if capacity permits | \n| 🟡 IF TIME | Useful but lower priority | \n| ⚪ DEFER | Cannot realistically fit | \n\nThe planner therefore answers two questions:\n\nThat second question is surprisingly important.\n\nA study plan shouldn't tell someone to learn:\n\n“Advanced Dynamic Programming”\n\nbefore:\n\n“Dynamic Programming fundamentals.”\n\nPrepFlow models prerequisite relationships.\n\n```\nArrays\n  │\n  ▼\nRecursion\n  │\n  ▼\nDynamic Programming\n  │\n  ├──────────────► Knapsack\n  │\n  └──────────────► Longest Common Subsequence\n```\n\nThe planner can therefore account for foundational topics instead of treating every topic as an independent checkbox.\n\nReal students miss tasks.\n\nThat's normal.\n\nSo PrepFlow treats planning as a feedback loop:\n\n```\n          ┌──────────────┐\n          │ INITIAL PLAN │\n          └──────┬───────┘\n                 │\n                 ▼\n          ┌──────────────┐\n          │     TODAY    │\n          └──────┬───────┘\n                 │\n        ┌────────┼────────┐\n        ▼        ▼        ▼\n     COMPLETE   SKIP    MISS\n        │        │        │\n        └────────┴────────┘\n                 │\n                 ▼\n        ┌─────────────────┐\n        │ RECOVERY /      │\n        │ REBALANCE       │\n        └────────┬────────┘\n                 │\n                 ▼\n          NEW PLAN VERSION\n```\n\nCompleted work is preserved.\n\nUnfinished work can be recovered or re-planned.\n\nThis is much closer to how real studying works than generating one static timetable.\n\nStudying a topic once isn't enough.\n\nPrepFlow schedules revision sessions after learning using spaced intervals.\n\nConceptually:\n\n```\nLEARN\n  │\n  ├──── +1 study day ────► REVISION 1\n  │\n  ├──── +3 study days ────► REVISION 2\n  │\n  └──── +7 study days ────► REVISION 3\n```\n\nThe planner also respects the exam boundary and available study days.\n\nThe repository is a TypeScript monorepo containing the React frontend, Express/TypeScript API, and shared schemas/types.\n\n```\nPrepFlow/\n│\n├── apps/\n│   ├── api/\n│   │   ├── domain/\n│   │   │   ├── planning/\n│   │   │   └── analysis/\n│   │   │\n│   │   ├── infrastructure/\n│   │   │   ├── prisma/\n│   │   │   ├── PDF extraction\n│   │   │   └── AI providers\n│   │   │\n│   │   └── HTTP API\n│   │\n│   └── web/\n│       ├── React\n│       ├── TypeScript\n│       ├── Vite\n│       └── Tailwind\n│\n├── packages/\n│   └── shared/\n│       └── Zod schemas + shared types\n│\n├── docs/\n│   ├── analysis.md\n│   ├── planning.md\n│   └── decisions/\n│\n└── PostgreSQL + Prisma\n```\n\nI didn't want the entire application to become dependent on one model.\n\nThe architecture therefore separates:\n\n```\n                APPLICATION\n\n                    │\n                    ▼\n             ┌─────────────┐\n             │ AI PROVIDER │\n             │ INTERFACE   │\n             └──────┬──────┘\n                    │\n          ┌─────────┼─────────┐\n          ▼         ▼         ▼\n       Ollama     Gemini     Fake\n       + Gemma     + Gemma   Provider\n```\n\nThe model can change without rewriting the planning engine, database layer, or frontend.\n\nThe repository currently exposes `ollama`, `gemini`, and `fake` provider modes.\n\nThe system treats uploaded study material as **untrusted input**.\n\n```\nUSER MATERIAL\n     │\n     │ untrusted\n     ▼\n┌──────────────┐\n│ nonce +      │\n│ delimiters   │\n└──────┬───────┘\n       ▼\n     GEMMA\n       │\n       ▼\n┌──────────────┐\n│ tolerant     │\n│ JSON parsing │\n└──────┬───────┘\n       ▼\n┌──────────────┐\n│ Zod schema   │\n│ validation   │\n└──────┬───────┘\n       ▼\n┌──────────────┐\n│ source       │\n│ evidence     │\n│ verification │\n└──────┬───────┘\n       ▼\n     ACCEPT\n```\n\nThe model receives no application secrets or tools.\n\nIts output is never treated as trusted application state.\n\nThe repository also documents bounded retries, analysis concurrency limits, upload limits, CORS controls, and environment-based secrets.\n\nI didn't want the project to only work in a happy-path demo.\n\nThe current test suite covers multiple layers:\n\n| Layer | What is tested | \n|---|---|\n| Shared | Schemas and shared logic | \n| API | Business logic and routes | \n| Web | Frontend behavior | \n| HTTP | End-to-end API routing | \n| Planning | Priority, capacity, scheduling | \n| Randomized invariants | Planner safety properties | \n| PostgreSQL | Real persistence | \n| Concurrency | Rebalance/version behavior | \n| PDF extraction | Real document parsing | \n\nCurrent validation:\n\nand\n\nThe integration tests include persistence and concurrent plan-rebalancing scenarios.\n\n```\nAUTOMATED TESTS\n\nShared       ████████████████████  22\nAPI          ████████████████████ 358\nWeb          ████████████████████ 43\n                                  ───\nTOTAL                              423\n```\n\nReal database integration:\n\n```\nPostgreSQL integration\n\nPassed     ████████████████████ 20\nFailed                           0\n```\n\nI also validated the PDF extraction pipeline against a real academic PDF and compared the extracted page structure against the source.\n\n| Layer | Technology | Why | \n|---|---|---|\n| Frontend | React + TypeScript | Component-based UI | \n| Build | Vite | Fast development/build | \n| Styling | Tailwind CSS | Consistent UI system | \n| Backend | Node.js + Express | Lightweight API | \n| Language | TypeScript | Shared type safety | \n| Validation | Zod | Runtime contracts | \n| Database | PostgreSQL | Relational persistence | \n| ORM | Prisma | Type-safe DB access | \n|  | PDF.js | Page-aware extraction | \n| AI | Gemma | Open-weight material analysis | \n| Local AI | Ollama | Local model execution | \n| Testing | Vitest | Fast automated testing | \n| CI | GitHub Actions | Automated verification | \n\nStudy material can be surprisingly sensitive.\n\nIt can contain:\n\nThat's why I wanted PrepFlow's architecture to support **local/open-weight AI**, rather than making a cloud LLM the only possible path.\n\nGemma sits behind an AI-provider abstraction instead of being hardcoded into the application's business logic.\n\nThat gives the project a useful property:\n\n**The intelligence can evolve without rebuilding the product around a single AI provider.**\n\n| PrepFlow is | PrepFlow isn't | \n|---|---|\n| AI-assisted planning | A generic chatbot | \n| Evidence-backed analysis | Blind LLM output | \n| Deterministic scheduling | LLM-generated timetable | \n| Capacity-aware | “Everything fits” fantasy | \n| Revision-aware | One-time checklist | \n| Explainable | Black-box prioritization | \n| Open-weight AI compatible | Locked to one provider | \n| Built around execution | Just another note summarizer | \n\nThe biggest thing I learned wasn't how to connect an AI model to a backend.\n\nIt was learning **where not to use AI**.\n\nA tempting architecture would have been:\n\n```\nPDF\n │\n ▼\nLLM\n │\n ▼\n\"Here is your study plan.\"\n```\n\nIt's easy.\n\nIt's also difficult to trust.\n\nPrepFlow instead looks more like:\n\n```\n                 AI\n                  │\n                  ▼\n        UNDERSTAND THE MATERIAL\n                  │\n                  ▼\n             STRUCTURED DATA\n                  │\n                  ▼\n        ┌────────────────────┐\n        │ DETERMINISTIC CODE │\n        │                    │\n        │ Validate           │\n        │ Prioritize         │\n        │ Calculate          │\n        │ Schedule           │\n        │ Revise             │\n        │ Recover            │\n        └─────────┬──────────┘\n                  │\n                  ▼\n             REAL PLAN\n```\n\nThat separation gives me much more confidence in the system.\n\nPrepFlow is intentionally an MVP.\n\nThe next areas I'd explore are:\n\n| Area | Possible improvement | \n|---|---|\n| Documents | OCR for scanned PDFs | \n| AI | More provider/model evaluation | \n| Knowledge | Better prerequisite inference | \n| Planning | Calendar-aware scheduling | \n| Progress | More sophisticated learning models | \n| AI runtime | Additional local/open-weight providers | \n| UX | Topic editing and plan refinement | \n| Validation | Larger real-student testing | \n\nThe current repository roadmap already separates the completed repository/database, ingestion, AI analysis, and planning milestones from later topic-review, polish, deployment, and real-user testing.\n\n[github.com/piyusshhjangid/PrepFlow](https://github.com/piyusshhjangid/PrepFlow?utm_source=chatgpt.com)\n\n```\nnpm install\nnpm run db:migrate\nnpm run db:seed\n\nnpm run dev:api\nnpm run dev:web\n```\n\nThen open the web application.\n\nThe repository also documents the local Gemma/Ollama and hosted Gemma provider setup.\n\nBuilding for a friend changes the question.\n\nInstead of:\n\n“What AI feature can I add?”\n\nyou start asking:\n\n**“What is actually making this person's life harder?”**\n\nFor PrepFlow, the answer was:\n\n**Students don't necessarily need more study material.**\n\nThey need help turning the material they already have into a realistic sequence of actions.\n\nSo that's what I built.\n\nNot another chatbot.\n\nNot another PDF summarizer.\n\nA system that tries to answer one deceptively difficult question:\n\nBuilt for the **Hacktoberfest Weekend Challenge — Build for a Friend**.", "url": "https://wpnews.pro/news/prepflow-from-study-material-to-an-actual-study-plan", "canonical_source": "https://dev.to/piyusshhjangid/prepflow-from-study-material-to-an-actual-study-plan-58e3", "published_at": "2026-10-05 06:42:03+00:00", "updated_at": "2026-10-05 06:43:12.421699+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "generative-ai", "ai-products"], "entities": ["PrepFlow", "Gemma", "Hacktoberfest"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/prepflow-from-study-material-to-an-actual-study-plan", "markdown": "https://wpnews.pro/news/prepflow-from-study-material-to-an-actual-study-plan.md", "text": "https://wpnews.pro/news/prepflow-from-study-material-to-an-actual-study-plan.txt", "jsonld": "https://wpnews.pro/news/prepflow-from-study-material-to-an-actual-study-plan.jsonld"}}