cd /news/generative-ai/please-let-me-phone-my-podcasts-app · home topics generative-ai article
[ARTICLE · art-84871] src=interconnected.org ↗ pub= topic=generative-ai verified=true sentiment=· neutral

Please let me phone my podcasts app

In a blog post, technology writer and designer Matt Webb proposed a voice-first podcast app that lets listeners talk back to the show to capture notes, citing a need for a minimal Granola-like interface to timestamp moments. Webb referenced Ethan Jucovy's Voice Claude project, which wired Claude into a car via Twilio, and OpenAI's 1-800-ChatGPT, highlighting the potential for ambient voice AI assistants.

read5 min views1 publishedAug 3, 2026
Please let me phone my podcasts app
Image: source

Please let me phone my podcasts app #

14.55, Friday 24 Jul 2026

Link to this post My perfect podcast app would be a long phone call in which I listen to the show and can also talk back to remember things. (Someone build this for me.)

It would immeasurably improve my running. Long story short, I added slow runs into the training mix, and had to switch from high tempo music to podcasts because I kept going too fast.

BUT genuinely interesting podcasts are a problem because I keep wanting to stop running to take notes.

So I end up listening to shows that only scrape past the threshold of keeping my interest, nothing more mind-fizzy. My Goldilocks zone podcast genre is premium mediocre.

Sad.

Designer Kate Pincott suggested (when we were out at a designers dinner last week) that I need a minimal Granola-like interface. A way to take notes simply by marking a moment.

Like: perhaps I would double-tap my watch, and it would timestamp the moment in the podcast episode to capture it, then later it could pull out a few relevant sentences and drop it into my notes.

Like dog-earing pages when you’re reading a book, only for audio. I love it, you rarely need more than that. And then I could start listening to all the best podcasts again.

But it got me thinking about the whole podcast experience…

Podcasts are audio-first.

If my podcast app was a person in the faves in my phone app or in WhatsApp, I could call them up and talk back to them. hey what new episodes do we have since last time?

(I can subscribe to podcasts either with voice or using the regular app, because the graphical UI is best for most tasks)oh actually the Rest is History had a series about Julius Caesar I was listening to, let’s have the next episode of that

(AI is good at deciphering this kind of intent)- But it might speak back to me: the series you were listening to most recently or the one that just came out?

(And then we could have a convo, and I would start listening) thanks, good morning,

I say to someone who held back their dog to let me run by (and the app intelligently ignores me)oh that’s interesting, I wonder about the

(Then it transcribes the audio at that moment, in a semantically meaningful chunk, appends my own commentary, and adds it to my notes.)relationship between Augustus and Caesarion

Proper note-taking… on the run.

Wouldn’t that be so great?

(I’m not interested in talking back to the hosts via generative AI.)

The last time I suggested a voice-first app it was Roadtrip, a voice-mode AI travel buddy that lives in the dash of my car, and I can ask questions to as I drive. Occasionally it pops up with an observation about something we’re passing.

Then Ethan Jucovy sent me their project where they’d put Claude at the end of a phone line and wired it into their car: Voice Claude.

Which is just perfect. Phoning Claude is the perfect interface.

I hope Ethan doesn’t mind me quoting their email:

I was driving the hour+ home from a Medieval music concert and wanted urgently to talk to Claude about the surprisingly-Islamic-sounding touches I had heard in a 13th century English secular winter song (would influences from Spain have reached England to cross-pollinate by then? or was it some shared lineage? or just a coincidence?).

…so they built Voice Claude, wired it up via Twilio, and had been using it daily since. 2024!

Around the same time, OpenAI launched 1-800-ChatGPT. All that is missing from these is the ability for a voice buddy to be ambient. Someone chilling in the shotgun seat, piping up only occasionally. Most phone calls don’t work like that, so they would need some kind of twist.

So I’m into transcription too, and I semi-regularly kick off projects by talking into my watch for 30 minutes.

It works best when my transcriber has a name for control plane instructions: hello Diane.

i.e. here I am with my imaginary posse of voice buddies.

BUT: why shouldn’t this all run via Siri?

Why not Hey Siri [podcast intent here] and Hey Siri [I’m write a lecture now so listen] and Hey Siri [etc]? That is to say, given all of this could run via future Siri, why do I keep coming back to the idea of these individualised, domain-specific voice modes?

For me, a lot of it comes down to “app level” vs “system level” interactions… In typical conversations, both parties are aware of the mutal common ground of the topic (the “app”) – and both parties are aware that the other is aware of the topic. This implicature allows for higher bandwidth. There is still of course of need to multiplex on-topic statements and out-of-band statements, but look when people talk: we tend to use intonation and body language to signal a shift. (And voice AI doesn’t pick on that yet.)

In a GUI, an application window - or, better, full screen - is used to explicitly focus the human and the machine on a topic. Out-of-band/control plan/system interactions (whatever you call it) is a swipe from the bottom of the screen or a move of the cursor up to the menu bar.

There is no full screen button for voice mode.

So the best solution I can come up with is that each different voice capability is a character in my phone app. It would help with discoverability too.

Maybe, when you install an app with voice mode, it appears in a special section of your phone contacts.

Honestly I can’t say I’m in love with this, as an approach. It’s a little twee and likely cognitively overwhelming.

But it’s nicely composable (you could invite your podcast app and your Wikipedia app to the same group call?).

And we’re going to need to figure this out at some point, given the upcoming proliferation of voice apps and devices.

All of that said, I still want a podcast app I can phone and talk back to.

If you build this I would like a free subscription or I’ll wait 3 months and paste this post into Fable 2 myself no hassle kthxbye

── more in #generative-ai 4 stories · sorted by recency
── more on @matt webb 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/please-let-me-phone-…] indexed:0 read:5min 2026-08-03 ·