Most meeting software is very good at producing a transcript and surprisingly bad at answering the question I actually care about: what did we agree to last time? I built RecallixAI around that gap, using a live meeting pipeline for immediate context and Hindsight for the part that has to survive long after the browser tab is closed.
What RecallixAI actually does
RecallixAI is a full-stack meeting intelligence system built around four pieces:
A Chrome companion that observes Google Meet closed captions without putting another bot into the meeting.
A FastAPI backend that owns meeting sessions, transcripts, notes, analysis, and integrations.
Google Gemini for turning a raw transcript plus private scratchpad into structured commitments, open questions, discrepancies, and a summary.
Vectorize Hindsight for persistent, per-user meeting memory, followed by automation through n8n, Google Calendar, Gmail, and Google Sheets.
The important distinction is that I don't treat the transcript as the product.
The transcript is an input. The useful state is the set of commitments, decisions, unresolved questions, people involved, and the history surrounding those things.
The repository reflects that separation. The backend exposes a streaming path for captions and notes, a completion path for post-meeting synthesis, and dedicated endpoints for meeting preparation and contact dossiers. The browser extension feeds the first path, while the dashboard consumes the resulting state.
At the top level, FastAPI mounts the API router and serves the dashboard:
app = FastAPI(
title="Meeting Intelligence Agent API & Dashboard",
description="Autonomous Meeting Intelligence & Prep Agent powered by Vectorize Hindsight & Google Calendar",
version="2.0.0"
)
app.include_router(api_router, prefix="/api")
That sounds ordinary, and intentionally so. I wanted the orchestration to live behind a small HTTP surface instead of spreading meeting state across browser code and third-party services.
The interesting part starts before the LLM
I made a deliberate choice not to build the capture layer around a meeting bot.
The Chrome extension watches the closed-caption DOM inside Google Meet. It uses MutationObserver to detect changes, extracts the speaker and text, and sends structured caption events to the backend.
The core loop is small:
async function pushCaptionToBackend(speaker, text) {
const meetingId = getMeetingId();
const payload = {
meeting_id: meetingId,
speaker: speaker || "Participant",
text: text,
timestamp: new Date().toISOString()
};
await fetch(${API_BASE_URL}/api/stream/caption, {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify(payload)
});
}
The extension is deliberately dumb. It doesn't try to summarize the meeting, infer commitments, or maintain a second copy of the application's business logic.
That matters because browser integrations are brittle enough already. Google Meet's DOM can change. I would rather have one small responsibility fail than have meeting capture, inference, memory, and workflow automation coupled to selectors in a content script.
There is also a practical detail that ended up being important: captions can be emitted incrementally and sometimes repeat. The extension tracks the previous caption and speaker, while the backend also avoids consecutive duplicate caption records.
The backend receives those events through a simple endpoint:
@router.post("/stream/caption")
def ingest_stream_caption(req: StreamCaptionRequest):
caption = session_service.add_caption(
meeting_id=req.meeting_id,
speaker=req.speaker,
text=req.text,
timestamp=req.timestamp
)
return {
"status": "ingested",
"meeting_id": req.meeting_id,
"caption": caption
}
That gives me a clean boundary: capture produces events; the backend owns meeting state.
I kept live state and long-term memory separate
This was probably the most important architectural decision.
During a call, I need low-friction state. I need the latest captions and the user's scratchpad immediately. I don't need a sophisticated memory system for every keystroke.
The session service therefore keeps an active MeetingSession with captions, notes, attendees, timestamps, and status.
class MeetingSession:
def **init**(self, meeting_id: str, title: Optional[str] = None):
self.meeting_id = meeting_id
self.title = title or f"Meeting {meeting_id}"
self.start_time = datetime.datetime.now(datetime.timezone.utc).isoformat()
self.captions = []
self.user_notes = ""
self.attendees = []
self.status = "active"
Hindsight enters at a different boundary.
When a meeting is complete, Gemini reconciles the spoken transcript with the user's scratchpad. The resulting structured analysis is retained in Hindsight as durable memory.
That gives me two very different stores:
Session state: what is happening right now.
Hindsight memory: what should still matter weeks later.
I don't want a database of raw events masquerading as memory. I want the system to be able to answer a future question in terms of the relationship and its history.
That is where Hindsight on GitHub fits naturally into the design. Hindsight provides retain, recall, and reflection-oriented memory primitives rather than forcing me to build a retrieval layer from scratch. Its model is specifically aimed at persistent agent memory rather than simply storing conversation history.
Hindsight is useful because the question is contextual
A conventional search over old transcripts can answer "find meetings containing pricing."
It is much less useful for:
"I'm meeting Sarah again. What did we promise each other, and what was still unresolved?"
RecallixAI sends Hindsight a query shaped around that actual task:
query = f"What was discussed, promised, or left unresolved with {attendee_email}?"
recall_response = self.client.recall(
bank_id=bank_id,
query=query,
budget="mid"
)
The memory bank is scoped to the user, and the bank gets a mission describing what matters:
MISSION_STATEMENT = (
"Track all meeting dialogues, commitments made by both parties, "
"agreed deadlines, unresolved questions, and user note preferences."
)
I also configure Hindsight's disposition for this workload:
DEFAULT_DISPOSITION = {
"literalism": 4,
"skepticism": 2,
"empathy": 3
}
I like this because memory is not completely generic. A meeting assistant has a different definition of useful memory from a coding assistant or a personal chatbot.
The Hindsight documentation describes this broader model through memory banks, retain, recall, reflect, observations, and consolidated knowledge. For RecallixAI, the practical benefit is that I can treat Hindsight as a dedicated long-term memory layer instead of inventing an application-specific combination of embeddings, metadata filters, and transcript search.
This is also why the distinction between agent memory and ordinary retrieval matters. The system isn't merely asking, "Which old text looks similar to this query?" It is trying to preserve useful knowledge about an ongoing relationship.
The LLM isn't allowed to define the application's state
Another choice I made was to force the meeting analysis into a typed structure.
Gemini receives the transcript and scratchpad, but the output has to conform to MeetingAnalysis:
class MeetingAnalysis(BaseModel):
summary: str
promises_by_us: List[str] = Field(default_factory=list)
promises_by_them: List[str] = Field(default_factory=list)
missed_or_pending_followups: List[str] = Field(default_factory=list)
note_discrepancies: List[str] = Field(default_factory=list)
suggested_followup_date: Optional[str] = None
The model call requests JSON with that schema:
response = self.gemini_client.models.generate_content(
model=settings.GEMINI_MODEL,
contents=prompt,
config={
"response_mime_type": "application/json",
"response_schema": MeetingAnalysis
}
)
That gives the rest of the application something much more reliable than a generated paragraph.
After synthesis, the backend builds a durable memory record containing the summary, commitments from both sides, pending follow-ups, discrepancies, and the original notes. It then calls Hindsight's retain() with meeting metadata.
self.client.retain(
bank_id=bank_id,
content=content,
context=context,
document_id=doc_id,
metadata=metadata
)
The interesting part is what happens next. The same memory is available when preparing for another meeting:
result = hindsight_service.recall_prep_context(
user_id=req.user_id,
attendee_email=req.attendee_email
)
So the workflow becomes:
capture → reconcile → retain → recall → act
That loop is the actual system.
A concrete interaction
Imagine a meeting where I say:
I'll send the enterprise pricing breakdown by Thursday.
The client says they'll send their security questionnaire the next morning.
During the call, I write:
[Promise: I will deliver the updated enterprise pricing breakdown by Friday afternoon]
[Action: Sarah needs to send over technical security questionnaire]
The transcript says Thursday, while my note says Friday.
RecallixAI doesn't simply preserve both pieces of text and call the job done. Gemini is instructed to compare the two sources and produce a discrepancy.
The analysis can therefore contain:
My commitment: send the pricing breakdown.
Their commitment: send the security questionnaire.
An unresolved item: clarify the remaining security/SLA questions.
A discrepancy: my scratchpad says Friday while the spoken agreement says Thursday.
A proposed follow-up date.
That structured result is retained in Hindsight.
When the next meeting appears on the calendar, the preparation endpoint asks Hindsight what was previously discussed, promised, or left unresolved. The dashboard can then present the relationship context before I join the call.
After the meeting is completed, action items can also be sent through n8n. The backend turns extracted commitments into task records and posts them to the configured workflow. A separate calendar endpoint can create the confirmed follow-up event.
This separation is useful because n8n is handling integration plumbing, not business reasoning. Gemini determines what the meeting produced. Hindsight preserves why it matters later. n8n handles the mechanical work of moving those results into external systems.
What I learned building it
My first instinct with meeting software would have been to store everything and add search later.
That is backwards.
The useful unit is not "a transcript from September 29." It is "what matters when I talk to this person again?" Defining that question early makes the memory mission, metadata, retention content, and recall query much easier to reason about.
A live session wants simple mutable state.
Long-term memory wants durable, searchable knowledge.
Trying to make one mechanism handle both creates unnecessary complexity. Keeping the session service lightweight and using Hindsight at the retention boundary made the architecture easier to reason about.
I don't want downstream code parsing prose from a model.
The MeetingAnalysis schema gives the model freedom to infer the content while keeping the application's state predictable. That makes persistence, automation, UI rendering, and testing much easier.
Sending every caption independently into long-term memory would produce a lot of noise.
RecallixAI instead synthesizes the meeting first and retains a compact record containing the things I actually expect to matter later: commitments, unresolved work, discrepancies, decisions, and context.
Hindsight then has a better-quality memory to recall from.
Scheduling a follow-up or writing a spreadsheet row is easy.
Knowing why a follow-up is necessary is harder.
I kept that ordering explicit: first capture the conversation, then analyze it, then retain the durable result, and only then trigger external actions.
The bigger lesson
The interesting part of RecallixAI isn't transcription, summarization, or calendar automation individually. None of those are particularly new problems.
The hard part is continuity.
A meeting produces decisions today. Those decisions need to influence what I do tomorrow. The next conversation should not begin with a blank context window and a pile of old transcripts.
That is why Hindsight became a central architectural component rather than an optional search feature. It gives RecallixAI a place to retain the parts of a meeting that should survive the meeting itself, and a way to retrieve that knowledge in the context of the next interaction.
The resulting architecture is intentionally straightforward:
Google Meet captions → FastAPI session state → Gemini reconciliation → Hindsight memory → pre-meeting recall → external actions.
That is the pattern I would reuse elsewhere: keep the real-time path simple, make the LLM output structured, and treat long-term memory as its own system with a clear purpose.
The goal isn't for the application to remember everything.
It's for it to remember the things that will change what happens next.