Between Us is a private, evidence-first memory reconstruction system for small friend groups. It is designed for the problem that most social memory tools ignore: people share the same experience, but each person captures only fragments of it, and those fragments are scattered across photos, notes, chats, voice recordings, screenshots, and timestamps.
Instead of building a public media feed or generic AI recap engine, Between Us focuses on reconstructing likely shared Moments from multiple contributors. The app treats memory as a structured inference problem:
Each user uploads fragments
The system extracts observations from them
Related fragments are retrieved based on time, semantics, and context
A model evaluates whether those fragments likely belong to a common real-world moment
The system surfaces a candidate Moment with evidence and uncertainty
Users can review, correct, and confirm it
This is not “an AI album.” It is a memory system that helps groups answer:
“Did this happen?”
“What happened?”
“What evidence supports that?”
“What are we still uncertain about?”
It is built for:
classmates
coworkers with recurring events
roommates
travel groups
close friends and tight social circles
The user problem is simple but painful: shared memories get fragmented, forgotten, or remembered inconsistently. People remember the same event differently, and key details disappear because they weren’t captured in the same way by everyone. Traditional social platforms are optimized for public sharing, not reconstructing private group memory with provenance and human correction.
Between Us solves that by turning fragmented evidence into structured memory candidates while preserving uncertainty and group boundaries.
Live demo: https://betweenus-w1h5.onrender.com/
GitHub repo: https://github.com/srikar-naidu/BetweenUS
I built Between Us as a full-stack Next.js + TypeScript application centered around a self-hosted AI memory pipeline.
The core architecture is:
Frontend + API layer: Next.js App Router
Authentication: Better Auth + Google OAuth
Canonical structured state: MongoDB Atlas
Memory retrieval: Tiger Data
Background job processing: MongoDB-backed worker pipeline
AI reasoning: Gemma 4 E2B via Ollama locally, and a private Render-hosted runtime in production
Observability: Sentry
Group semantic context: Backboard
Optional voice evidence: ElevenLabs
The project is designed as a coherent memory system rather than a generic chatbot or dataset dump.
The app follows a structured event pipeline:
Fragment
→ Observation extraction
→ Retrieval of related evidence
→ Candidate moment construction
→ Evidence-backed reasoning
→ Moment / Story proposal
→ Human review + correction
This flow is intentionally different from a generic “AI summary” product. The model is not asked to invent a story out of thin air. It is given constrained context from relevant, authorized fragments and asked to reason about a possible shared event.
Gemma is used as the multimodal reasoning layer for:
image understanding
screenshot and caption interpretation
video-derived frame analysis
text fragment summarization
entity extraction
temporal and event-level inference
evidence-backed moment reconstruction
Gemma does not generate a “final memory” as an ungrounded narrative. The app treats it as a reasoning component operating over retrieved evidence, with validation and provenance layering around it.
Between Us is not a simple vector-search app. It uses Tiger Data to retrieve context based on:
temporal proximity
semantic similarity
shared people
shared places
repeated entities
event adjacency
existing memory relationships
This is important because the app needs to answer questions like:
“What happened around this moment?”
“Which other fragments are likely related?”
“Are these two observations likely from the same event?”
“Do we have enough evidence to reconstruct a shared moment?”
Instead of sending full raw storage to the model, the system assembles a compact evidence packet made from the most relevant fragments. That reduces noise and keeps the reasoning process closer to the actual memory problem.
The product is built around a memory graph:
User
Group
Observation
Evidence
Moment
Story
Person / Place / Entity
Correction / rejection / approval
Provenance metadata
The active objects are:
A single piece of evidence uploaded by a user, such as:
photo
screenshot
text note
short video
voice note
timestamped caption
geotagged location
metadata-rich upload
A structured understanding of a fragment:
entities
people
objects
places
timestamps
text content
confidence
uncertainty
source provenance
A possible real-world event reconstructed from several related fragments.
A higher-order grouping of repeated or related moments over time.
This is very different from a standard “gallery app”—the primary unit is not a photo, but a memory candidate.
Between Us is explicitly privacy-first. The application is designed to avoid the pattern of “publicly exposing everything because the AI thinks it’s related.”
Important principles:
group boundaries are enforced
media stays private by default
AI inference must be evidence-backed
unsupported conclusions fail closed
uncertainty stays visible
corrections become part of the memory system
only authorized evidence can inform a reconstructed Moment
This is critical because memory reconstruction is not just an inference task—it is a trust and consent problem.
The system is built around this flow:
Plain text
User uploads fragment
↓
Validate upload + permissions
↓
Store canonical fragment metadata
↓
Queue for analysis
↓
Extract structured observation
↓
Retrieve relevant nearby/related fragments
↓
Assemble constrained context packet
↓
Gemma reasons over evidence
↓
Candidate Moment generated
↓
Evidence + uncertainty displayed
↓
Human review / correction / confirmation
↓
Confirmed memory enters shared state
The model is prompted with only relevant authorized fragments, not the entire database. This is important because the goal is not to generate generic stories from broad memory context; it is to reason over bounded evidence.
The system also enforces:
source attribution
confidence limits
extraction validation
output schema constraints
rejection of unsupported or speculative claims
A memory system without correction is dangerous. Between Us treats corrections as first-class memory input.
Examples:
reject a candidate Moment
merge two related fragments into a better explanation
correct the inferred time or location
reclassify evidence as weak or irrelevant
confirm a candidate as a real group memory
This makes the system better over time and keeps the product aligned with how actual human memory works: imperfect, revisionary, uncertain, and social.
Diagram
And in plain text:
User Fragment
↓
Auth + permissions
↓
MongoDB storage
↓
Worker queue
↓
Gemma observation extraction
↓
Tiger Data retrieval
↓
Evidence packet assembly
↓
Gemma reasoning
↓
Candidate Moment / Story
↓
User review & correction
↓
Confirmed group memory
Open innovation matters because this project is fundamentally about privacy, accountability, and experimentation—not just model output quality.
What it made possible:
self-hosted Gemma inference instead of vendor lock-in
local prototyping without depending on a closed external model API
private deployment patterns aligned with sensitive memory data
transparent evidence-based reasoning rather than black-box summarization
architecture flexibility to swap retrieval, storage, or inference providers without rewriting the whole product
building a system where uncertainty is not hidden behind “confident” model output
A closed API would not fit as well because Between Us is not just a “prompt and response” app. It requires:
permission-aware memory reconstruction
privacy-safe group boundaries
provenance-aware evidence joining
correction loops
constrained reasoning on small evidence packets
traceable model behavior during debugging and evaluation
Open-weight models and open tooling allowed this to be built as a real system rather than a brittle demo wrapper around a hosted chatbot.
I’m entering the following partner categories:
Use: Core multimodal AI and memory reconstruction.
Gemma 4 analyzes user-submitted media and evidence and extracts structured observations such as:
activities
locations
text
temporal clues
It then reasons across multiple fragments to reconstruct Moments and connect them into Stories.
This is the product’s core intelligence layer.
Use: AI-assisted development.
GitHub Copilot was used throughout the project to help with:
system design
implementation planning
repetitive engineering tasks
tests
data model scaffolding
API design
debugging large implementation surfaces
It accelerated development without being the runtime product itself.
Use: AI observability and debugging.
Sentry helps monitor:
inference failures
retrieval failures
latency
pipeline errors
abnormal behavior in moment reconstruction
This is especially important because memory systems are sensitive to subtle failure modes such as poor retrieval or weak evidence.
Use: Deployment + runtime infrastructure.
Render hosts the deployed application and background processing needed to support the asynchronous memory pipeline.
The architecture intentionally supports running the model in a private runtime environment rather than depending on developer-only local machines.
Use: Temporal and semantic memory retrieval.
Tiger Data helps find relevant fragments by combining:
shared entities
shared locations
event relationships
This is essential for memory reconstruction because the key question is not “what is similar in general?” but “what is relevant to this event and this group?”
Use: Canonical structured memory graph and app state.
MongoDB Atlas stores:
users and groups
fragments
observations
moments
stories
evidence
provenance
corrections
group memory state
This is the canonical system of record for all the structured memory objects.
Use: Persistent semantic group context.
Backboard stores:
nicknames
aliases
inside jokes
recurring references
group-specific meanings
This helps the system understand the group’s shared language, which is essential because memory is shaped by social context.
Use: Voice evidence and narrated memories.
ElevenLabs supports:
speech-to-text for voice evidence
optional narration of grounded moments
This makes the memory system more complete by allowing voice notes to contribute to the evidence graph and optionally converting grounded memories into a narrated experience.
Between Us is designed to do one thing very well:
Turn fragmented evidence from a group into a credible shared memory candidate while keeping uncertainty visible and preserving privacy.
It is not a generic AI social app.
It is not a memory dump.
It is not a public feed.
It is a system that says:
here is what may have happened
here is the evidence
here is what remains uncertain
here is what the group can review and correct
That is the real product value.