{"slug": "i-built-my-friend-a-voice-first-ai-for-unfinished-thoughts", "title": "I Built My Friend a Voice-First AI for Unfinished Thoughts", "summary": "A developer built Thread, a voice-first AI companion that captures unfinished thoughts by recording speech, transcribing and organizing it, and using Google's Gemma model to surface connections to earlier recordings. The app keeps the original transcript visible alongside the AI-generated title, summary and suggestions, and can return no connection when the relationship is too weak. It is implemented as a fixed three-stage pipeline rather than an autonomous agent.", "body_md": "*This is a submission for the [Hacktoberfest Weekend Challenge: Build for a Friend](https://dev.to/challenges/hacktoberfest-weekend-2026-10-01)*\n\nI built Thread for my friend, who often gets ideas while walking, reading or watching a series.\n\nOne afternoon, an idea surfaced while he was already concentrating on something else. He decided to write it down after finishing rather than interrupt the work in front of him. When he finally opened his notes, he could remember that the idea had felt useful, but not enough of it to recover the idea itself.\n\nHe normally uses a notes app or a notebook. Both work when he has enough time to stop and write. The difficulty is that an unfinished thought rarely arrives as a clear paragraph. Opening an app, deciding what to call the note and arranging the idea into proper sentences can create enough friction for it to disappear.\n\nThread gives that unfinished thought somewhere to wait.\n\nThe experience has three steps:\n\nPress record and speak naturally.\n\nLet Thread transcribe, organise and save the thought.\n\nReturn later to see whether it relates to something recorded earlier.\n\nSaving the thought is only the beginning. When a new one is captured, Thread retrieves a small set of relevant earlier thoughts. Gemma then examines whether one of them offers a useful connection, implication or question worth exploring.\n\nA person may mention difficulty concentrating during a walk and, several days later, consider a voice-based space for capturing ideas. Neither recording needs to refer deliberately to the other. Thread can bring the two together and suggest that the later idea may address the earlier problem.\n\nThread keeps the person’s original words separate from the AI interpretation. The transcript remains visible beside the generated title, summary and suggestions, so he can judge whether the interpretation is faithful. Gemma can also return no connection when the relationship is too weak.\n\nThe AI does not decide what he meant or what he should pursue. It offers a possibility. He makes that decision.\n\nThread helps my friend capture an idea before it disappears, return to his original words and notice when something said on another day gives an unfinished thought a possible direction.\n\nThe demonstration begins with several sample thoughts captured at separate moments. None of them is a complete idea by itself, and the final recording does not deliberately refer to the earlier ones.\n\nAfter Thread saves the new thought, it retrieves a relevant earlier recording. Gemma explains the possible relationship and offers a question that could help the person continue exploring it.\n\nThe screenshot below shows the result of that flow.\n\nThe original words remain visible beside Gemma’s interpretation.\n\nSeveral recordings later, Thread retrieved a related earlier thought and examined whether the relationship was useful.\n\n*Sample data showing how Thread brings back an earlier thought and presents Gemma’s connection as a possibility rather than a conclusion.*\n\nThe complete journey from the original voice recordings to the connection produced by Gemma is shown in the demo below.\n\nA voice-first thought companion for ideas that arrive before the words are in order.\n\nThread helps someone capture an unfinished thought without stopping to write and organise it. A person speaks naturally, and the application preserves the transcript, creates a structured interpretation, and saves both to a timeline.\n\nWhen a later thought is recorded, Thread retrieves relevant earlier thoughts and asks Gemma whether they reveal a useful relationship. It may suggest a connection, an implication, and a question worth exploring—or return no connection when the evidence is weak.\n\nThread follows a thought through three stages: capture what the person said, find earlier thoughts that may be related, and ask whether the relationship is useful enough to show.\n\nI built this as a fixed pipeline rather than an autonomous agent. Each stage has one responsibility, which makes the result easier to inspect when something goes wrong.\n\n```\nVoice → transcript → structured thought → PostgreSQL\n                    ↓\n               embedding → pgvector\n                    ↓\n        relevant earlier thoughts\n                    ↓\n          Gemma connection analysis\n```\n\nThe browser records the voice note and sends the audio to ElevenLabs Scribe for transcription.\n\nThe transcript then goes to Gemma, which returns a title, a short summary, categories and, when appropriate, a possible action or a question to explore. Zod validates the structure before the result is saved.\n\nThat validation catches responses with missing fields or an unexpected format. It cannot prove that the model interpreted the thought correctly.\n\nFor that reason, Thread stores the original transcript separately from every AI-generated field. The person can always return to what he actually said instead of seeing only the model’s summary.\n\nThe audio preview is temporary and disappears after the page is refreshed. Thread keeps the transcript as the permanent record, rather than becoming an audio archive.\n\nKeyword search works when two notes use the same words. Unfinished ideas often do not.\n\nA person might record one thought about losing concentration while studying and, several days later, another about remembering an idea after explaining it aloud. The wording is different, but there may still be a useful relationship between them.\n\nThread represents each transcript as an embedding: a list of numbers that captures aspects of its meaning. PostgreSQL stores these values using the pgvector extension.\n\nWhen a new thought is opened, Thread compares its embedding with earlier thoughts from the same workspace. It retrieves at most five candidates that pass the configured similarity threshold.\n\nIn simplified form, the retrieval behaves like this:\n\n``` js\nSELECT\n  earlier.title,\n  1 - (earlier.embedding <=> current.embedding) AS similarity\nFROM thoughts AS current\nJOIN thoughts AS earlier\n  ON earlier.id <> current.id\n  AND earlier.created_at < current.created_at\nWHERE current.id = $1\n  AND earlier.workspace_id = current.workspace_id\n  AND 1 - (earlier.embedding <=> current.embedding) >= $2\nORDER BY similarity DESC\nLIMIT 5;\n```\n\nThe real query also checks that the thoughts were created with a compatible embedding model and version. This matters because a score produced by one embedding model cannot safely be treated as equivalent to a score from another.\n\nRetrieval narrows the search. It does not decide that two thoughts form a meaningful idea.\n\nThread currently uses pgvector’s exact cosine-distance search. Since each workspace contains a small number of thoughts, I kept retrieval simple; an HNSW index would become useful as the collection grows.\n\nThe current thought and the retrieved candidates then go to Gemma for a second kind of reasoning.\n\nThe prompt does not ask, “Are these notes about the same subject?” Two thoughts can both mention learning, work or concentration without helping each other.\n\nInstead, Gemma looks for a more specific relationship. One thought should help address, test, support or challenge a goal, action, constraint or uncertainty in another.\n\nFor this step, Thread sends only the structured fields needed for comparison. It does not send the complete timeline, database identifiers, embeddings or raw transcripts.\n\nGemma can also return no connection. That is part of the design. If the evidence is weak, displaying nothing is more useful than forcing two thoughts into a convincing-sounding story.\n\nWhen a connection is found, Thread presents it as You Were Onto Something, followed by an explanation and a question the person may want to explore. The wording remains an AI interpretation. The original thoughts stay visible, and the person decides whether the suggestion is worth pursuing.\n\nConnection explanations are generated when the detail page runs the analysis. They are not saved as permanent conclusions, so a later analysis may produce different wording or decide that there is no useful connection.\n\nThread separates the application logic from the model provider.\n\nFor local reasoning, it can use Gemma 3 4B through Ollama. Local embeddings use EmbeddingGemma. This allows the reasoning and retrieval pipeline to run on hardware the user controls.\n\nThe public Render deployment uses Gemma 4 through Google AI Studio and Google’s hosted embedding model. I kept this path so judges and other users can try the application without installing models on their computers.\n\nThese paths share the same application interfaces, but they do not promise identical results. The local and hosted configurations use different reasoning and embedding models, so each configuration has its own similarity threshold and must be evaluated separately.\n\nSpeech transcription currently uses ElevenLabs in both configurations. The local option therefore moves the Gemma reasoning and embedding stages onto the user’s machine; it does not make the complete voice workflow offline.\n\nA successful API response only proves that a service answered. It does not prove that Thread found the right earlier thought or produced a useful connection.\n\nI tested the stages separately:\n\n| Evaluation | Observed result | \n|---|---|\n| Related thoughts with local embeddings | Similarity `0.748` ; included above the local`0.70` threshold | \n| Related thoughts in the hosted deployment | Similarity `0.8545` ; included above the hosted`0.80` threshold | \n| Unrelated thought | Excluded from retrieval | \n| Two passive observations sharing a topic | Gemma returned no useful connection | \n| No qualifying earlier thoughts | Connection inference was skipped | \n\nThe local and hosted similarity numbers should not be compared as a model ranking because they come from different embedding models. These checks verify several expected cases; they are not a measurement of general accuracy.\n\nI also added Sentry tracing around transcription, embedding, retrieval and connection reasoning. In one production verification trace, retrieval took about 25 ms and connection inference took about 918 ms. That is one request rather than a performance benchmark, but it helped show where the time was spent.\n\n*One production trace of Thread's connection workflow. Sentry separates retrieval from Gemma inference and records timing and token usage without including the thought text.*\n\nThe tracing export is filtered so thought content, transcripts, generated text, embeddings and record identifiers are not included.\n\nThe main engineering decision in Thread is the separation between finding something similar and deciding whether it is useful. Retrieval gives Gemma a small amount of relevant context. Gemma can then suggest a possible direction—or leave the thought alone when there is not enough evidence.\n\nA closed API could summarise a thought or return the same structured fields. What it cannot provide is a model whose weights I can run on the user’s own computer.\n\nThat matters for Thread because one recording is only a note, but many recordings gradually become a history of what someone has been thinking about. In the local configuration, Gemma 3 4B structures thoughts and examines their connections through Ollama. EmbeddingGemma creates the vectors used for retrieval, and PostgreSQL stores the timeline on the same computer.\n\nElevenLabs receives one recording for transcription, but it does not receive the timeline or the related thoughts. After the transcript returns, the accumulated memory, semantic search and connection analysis can remain local.\n\nThe public demo uses hosted Gemma, hosted embeddings and Render PostgreSQL so anyone can try Thread without installing models. That is a deployment choice rather than a requirement of the application. The reasoning and embedding layers sit behind provider interfaces, so the local and hosted versions follow the same flow even though their models may produce different results.\n\nOpen weights also mean that Gemma can be evaluated, replaced or eventually customised without rebuilding Thread around a different proprietary API. I have not fine-tuned the model or established a cost advantage. The benefit demonstrated here is more direct: the part of the application that reasons across a person’s growing thought history can run on hardware they control.\n\nI built Thread over the hackathon weekend with Codex. I used the session to turn the original idea into a working application, examine weak assumptions, trace failures and prepare the project for someone other than me to try.", "url": "https://wpnews.pro/news/i-built-my-friend-a-voice-first-ai-for-unfinished-thoughts", "canonical_source": "https://dev.to/kaustubh_05/i-built-my-friend-a-voice-first-ai-for-unfinished-thoughts-2dbn", "published_at": "2026-10-05 01:20:27+00:00", "updated_at": "2026-10-05 01:42:18.357121+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-products", "ai-tools", "large-language-models", "natural-language-processing"], "entities": ["Thread", "Gemma", "Google", "Hacktoberfest"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/i-built-my-friend-a-voice-first-ai-for-unfinished-thoughts", "markdown": "https://wpnews.pro/news/i-built-my-friend-a-voice-first-ai-for-unfinished-thoughts.md", "text": "https://wpnews.pro/news/i-built-my-friend-a-voice-first-ai-for-unfinished-thoughts.txt", "jsonld": "https://wpnews.pro/news/i-built-my-friend-a-voice-first-ai-for-unfinished-thoughts.jsonld"}}