{"slug": "flet-agent-got-eyes-and-ears", "title": "Flet agent got eyes and ears", "summary": "Flet Studio's latest release gives its AI agent the ability to run a user's app, take a screenshot of it and read its console output while working, plus support for attaching pictures, PDFs and code files to messages. The update addresses the previous limitation where the agent could only manipulate files in the app workspace, forcing users to describe visual and runtime problems in text. Flet says early adopters moving from the free plan to the Creator plan funded the work.", "body_md": "# Flet agent got eyes and ears\n\nBuilding an AI agent is exciting. Seeing people actually use it and get the results they wanted is even more exciting. And seeing early adopters go beyond the free plan and subscribe to the Creator plan is the best part! Thank you for your support! :)\n\nToday's release of [Flet Studio](https://studio.flet.dev) gives the agent eyes and ears:\nit can now **run your app, take a screenshot of it and read its console** while it's working. You can also attach pictures, PDFs and code files to your\nmessages, or just paste a screenshot.\n\n## A two-minute primer on AI agents\n\nYou can find an infinite number of articles about \"what an AI agent is\" or \"how to build your own agent\", but I'd like to quickly recap it here and establish some terminology, so we speak the same language here and in other places of the Flet docs.\n\n### The agent loop\n\nIn short, an agent is a loop with an LLM (or simply \"model\") in the middle, surrounded by\nprompts, skills, MCP servers and tools. Everything around the model is often called a\n**harness**: it gives the model knowledge it doesn't have and keeps its \"creativity\" (read:\nhallucinations) in check.\n\n### Turns\n\nYou send a prompt and a new **turn** starts. The prompt, together with the system prompt and\navailable tools, goes to the model. The model either answers right away or asks the agent to\ncall a tool - list files, read the Flet API docs, write `main.py`. The agent calls the tool and\nsends the result back to the model, which decides what to do next. This repeats until the model\nhas nothing more to do and gives you its final answer. That's the end of the turn.\n\n### Conversations and the context window\n\nSend another prompt and all previous turns - your prompts, the model's tool calls, their results and answers - are sent along with it. The model has no memory of its own, so the whole history goes back to it every time.\n\nThe sum of all turns is a **conversation**, or **context**, which the agent UI usually calls a\n\"chat\".\n\nHow long a conversation can be is limited by the model's maximum context window size. That's also why every extra turn costs more than the previous one - more tokens are sent each time.\n\n## Building UI blindfolded\n\nDeveloping an agent is a challenge. An LLM is non-deterministic, so if you ask it 10 times to \"build a counter app\", you get 10 different, but plausible, results. Prompts and skills narrow it down, and the tools define what the agent can actually do.\n\nThe first iteration of the Flet agent had only tools for manipulating files in the app workspace: list directories, read and write files. The information flowed in one direction - from the model to the app workspace.\n\nOK, let's give the agent a brain (the model) and a pen (the workspace tools) and see what happens!\n\n🙂 Build a counter app with plus and minus buttons.\n\n🤖 *(checks the Flet API, thinks for a minute, writes `main.py`)* Here, I built a simple\ncounter app with plus and minus buttons.\n\n😐 There is no minus button.\n\n🤖 *(checks the Flet API again, thinks for a few seconds, updates `main.py`)* I added the\nmissing minus button.\n\n😕 There is a Python error about a missing import.\n\n🤖 *(changes `main.py`)* My apologies, I used an incorrect import. It's fixed now.\n\n🙂 OK. Now, put the app content at the center of the page.\n\n🤖 *(thinks, changes `main.py`)* I centered the app content on the page.\n\n😠 It's not at the center.\n\n🤖 *(changes `main.py`)* I fixed the app and now it's centered.\n\n🤬 It's **still** not at the center!\n\n🤖 *(thinks harder, changes `main.py` again)* I restructured the app to make sure its content\nis at the center of the page.\n\n🤬 It's not centered, again!!!\n\n🤖 Could you give me a screenshot of what you see to help me figure out the solution?\n\n😤 I can't give you a screenshot!!!\n\n*(🚪 user left the chat)*\n\nDialogs like this could go on for hours. If you are persistent, you can probably finish the job in a dozen turns. But it's still really annoying, right?\n\nFor an agent building UI with nothing but file tools, not being able to see the result is the\nbiggest gap. The only feedback it gets is you describing, in text, what *you* see. Or it asks\nyou: \"I added some prints to the code - run the app, click here and there, and paste me the\ncontents of the Console\". Doable, but it's unnecessary friction. The human is still constantly\nneeded in the loop - you are the agent's eyes and ears.\n\n## Closing the loop: eyes and ears\n\nThis release fixes both sides of that conversation.\n\n### Attach files and screenshots\n\nYou can now attach files to your messages: images, PDFs, and text or code files. Use the paperclip button, or paste a screenshot from the clipboard with Cmd+V (Ctrl+V on Windows and Linux). So now you can tell the agent \"build a form like the one in the attached picture\" or \"this button doesn't look like I described - see attached\".\n\nAttachments work on the New app page too, so you can start a new app straight from a mockup. And the agent can put an attached file into your app - for example, \"use this image as the logo\".\n\n### Three new tools\n\nThe agent itself gets three new tools, and they take it to the next level:\n\n- **Run the app** - previously, the app was run once, automatically, after each agent turn. Now\nthe agent can re-run it any number of times in the middle of a turn. You'll see the Preview\nrestart when it does.\n- **Take a screenshot** - the agent's eyes. It looks at the running app and sees whether that\navatar is really in the corner.\n- **Read the console** - the agent's ears. Python errors, tracebacks and log messages no longer\nneed to be copy-pasted by you.\n\n## The new agent in action\n\nWith the new tools, conversations get shorter and reach the result without your intervention.\n\n### Catching Python errors\n\n🙂 Build a simple dashboard.\n\n🤖\n\n- Checks the Flet API\n- Writes\n`main.py`\n**Runs the app**\n**Reads the console** - there's an error!- Fixes\n`main.py`\n**Reads the console** - no errors- Done!\n😄\n\n### Fixing the layout\n\n🙂 Build a page with a user avatar in the bottom right corner of the screen.\n\n🤖\n\n- Writes\n`main.py`\n**Takes a screenshot** - the avatar is not in the corner- Checks the Flet API\n- Fixes\n`main.py`\n**Takes a screenshot** - the avatar is in the right place now- Done!\n😄\n\n### Debugging together\n\nSome bugs only show up when you click through the app. The agent can't click (yet), but it can set a trap and read what it caught:\n\n🙂 I have issues with my app...\n\n🤖\n\n- Adds logging messages to\n`main.py`\n**Runs the app**- \"I've added logging messages in various places of your program. Please click through the app and let me know when you're done.\"\n🙂 OK, done.\n\n🤖\n\n**Reads the console**- Fixes the issues in\n`main.py` and removes the logging messages\n**Runs the app**- Done!\n😄\n\nWith proper feedback, the agent finishes the job in one turn instead of ten - fewer turns means fewer tokens, fewer credits spent, and no babysitting.\n\n## Also in this release\n\n- **Better-organized apps.** New apps keep their code in a`src` folder, Flet's standard\nlayout, and split it into files by what they do. Small apps stay in a single file. New apps\nalso come with logging set up, which the agent uses for debugging.\n- The agent is better at **layout** , such as centering content or keeping a footer at the\nbottom of the page.\n- When your app **calls a web API** , the agent writes code that works both in the Preview and in\nthe app you build with`flet build` .\n- The **Expert** agent now runs on a newer model, and Flet Studio uses**Flet 1.0.3** .\n- Long tasks no longer stop partway because the agent made too many steps.\n\nSee [What's new in Flet Studio](https://flet.dev/docs/studio/whats-new) for the full list.\n\n## What's next\n\nIt would be great to give the agent an ability to tap, swipe, drag and do other interactions in a running app - with some guardrails, of course: we can't let the agent delete something in your production database.\n\nA self-learning agent with memory is another thing we could tackle in future releases.\n\nThat's all for today! Open [Flet Studio](https://studio.flet.dev) and give the new agent a try.\nIf you have an interesting story to share - what you built or tried to build with the Flet\nagent, challenges, obstacles - tell us in\n[GitHub Discussions](https://github.com/flet-dev/flet/discussions) or on\n[Discord](https://discord.gg/dzWXP8SHG8).\n\nHappy Flet-ing!", "url": "https://wpnews.pro/news/flet-agent-got-eyes-and-ears", "canonical_source": "https://flet.dev/blog/flet-agent-got-eyes-and-ears/", "published_at": "2026-10-01 03:37:01+00:00", "updated_at": "2026-10-01 03:48:38.281961+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools", "ai-products"], "entities": ["Flet", "Flet Studio", "Creator plan"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/flet-agent-got-eyes-and-ears", "markdown": "https://wpnews.pro/news/flet-agent-got-eyes-and-ears.md", "text": "https://wpnews.pro/news/flet-agent-got-eyes-and-ears.txt", "jsonld": "https://wpnews.pro/news/flet-agent-got-eyes-and-ears.jsonld"}}