Most study advice is a rehash of the same 140-year-old lab experiment. We had something different sitting in our own database: 18,217 real flashcard review events from 667 students studying everything from anatomy to the bar exam. So instead of citing Ebbinghaus again, we ran the numbers ourselves.
The dataset covers every graded flashcard review logged on StudyCards AI between September 21, 2025 and July 28, 2026, aggregated and stripped of anything user-identifiable before analysis. Here is what the data actually shows.
StudyCards AI lets students upload a PDF (lecture slides, a textbook chapter, a study guide) and turns it into a deck of flashcards. Every time someone answers a card during a study session, we log whether they got it right, how long they took, and how they rated the card's difficulty afterward. That log is the same data that powers spaced repetition scheduling, and over ten months we collected enough data to make some interesting discoveries.
Before running any of the numbers below, we checked for the two things that ruin this kind of analysis: bots and outliers. No single account accounted for more than about 5% of all reviews, and we excluded one internal test account. We also excluded any reviews where the user took more than 10 minutes to answer a card as it's likely they just left the tab open.
Methodology note: this is usage data from one flashcard app rather than a controlled experiment. We can show what correlates with what, but we can't prove causation the way a lab study could.
Students get a flashcard right 69.0% of the time on their first attempt (n=15,104). By the second time they see that same card, accuracy climbs to 80.5% (n=2,228). By the third attempt it's 89.7% (n=505), and by the fourth, 97.7% (n=215). Sample size drops off fast after that, so we're not going to pretend the curve keeps climbing forever, but the first four attempts show a clean, consistent trend.
This isn't a new discovery. "The testing effect", the idea that practice improves recall, is well established in cognitive science. What's notable is seeing it this cleanly in real app usage rather than a lab setting with volunteers who know they're being studied. Real students, using a real app, on their own schedule, still show the same trend.
Caveat: we did not find a clean relationship between the length of the gap between reviews and how well people remembered the card. The spacing effect idea that longer, well-timed gaps between reviews boost retention more than short ones, is a much harder thing to isolate from real-world usage data, where people review cards whenever they happen to open the app rather than on a controlled schedule. We looked for it and the signal was too noisy to publish with confidence, so we're not claiming it here. As we collect more data, we'll revisit this.
Across 1,593 study sessions, completion rate falls steadily as deck size grows. Sessions on decks of 10 cards or fewer are completed 25.0% of the time. That drops to 23.1% for 11-20 card decks, 16.8% for 21-30, 13.9% for 31-50, and just 7.8% for decks of 50 or more cards.
It's easy to read this as "smaller decks are better," and directionally that's probably true, but big decks and small decks aren't a controlled comparison. A 50-card deck might cover a denser topic, get built for a harder class, or take longer to finish in one sitting regardless of format. The data suggests if you're building a deck that's growing past 30-40 cards, split it into smaller chunks rather than one long session. People are far more likely to actually finish that way.
After answering a card, students can rate it easy, medium, or hard. Those self-ratings line up closely with what actually happens next. Cards rated "easy" are answered correctly 95.1% of the time (n=429). Cards rated "medium" come in at 72.3% (n=16,456). Cards rated "hard" drop to 47.4% (n=365).
Response time follows the same pattern with "easy" cards taking 5.7 seconds, "medium" cards 4.5 seconds, and "hard" cards 8.9 seconds. On this data, self-assessment while studying is a reasonably reliable signal.
Three things this data indicates:
If you're turning lecture notes or a textbook chapter into flashcards yourself, StudyCards AI generates them automatically from a PDF upload, in decks sized to actually get finished rather than abandoned halfway through. 18,217 individual flashcard review events from 667 students, covering 15,104 unique flashcards, logged between September 21, 2025 and July 28, 2026 on StudyCards AI.
Yes. In this dataset, accuracy on the same card rose from 69.0% on the first attempt to 97.7% by the fourth attempt, a clean, monotonic climb across four repeat exposures.
Session completion rate drops from 25.0% for decks of 10 cards or fewer to 7.8% for decks of 50+ cards. Keeping decks smaller, roughly 30-40 cards or less, correlates with people actually finishing their sessions.
Largely, yes. Self-rated "easy" cards were answered correctly 95.1% of the time versus 47.4% for self-rated "hard" cards, showing that gut-level difficulty ratings track real performance closely.
No. All figures are aggregate counts and percentages computed directly against the production database. No individual answers, user names, or deck contents were extracted or reviewed.
Generate Anki flashcards from PDFs