cd /news/artificial-intelligence/selective-agent-guidance-via-entropy… · home topics artificial-intelligence article
[ARTICLE · art-118633] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers

Researchers propose SAGE (Selective Agent Guidance via Entropy), a framework that queries a Vision-Language Model (VLM) only when a learner is uncertain and distills the guidance into a lightweight Reinforcement Learning (RL) policy, reducing VLM usage and improving performance over unguided RL in several sparse-reward visual reasoning and navigation tasks. The study, released on arXiv (2609.01567v1), shows that selective guidance is most beneficial when the VLM helps discover high-reward trajectories, and that learned policies can even exceed their VLM teacher.

read1 min views1 publishedSep 2, 2026

arXiv:2609.01567v1 Announce Type: new Abstract: Vision-Language Models (VLMs) provide useful priors for interactive decision-making, but using them directly as policies is expensive and brittle: they must be queried at every step, do not improve from environment interaction, and can repeat systematic errors. We study how to learn a cheap autonomous policy from an online, expensive, and imperfect but informative VLM teacher. We propose SAGE (Selective Agent Guidance via Entropy), a framework that queries a VLM only when the learner is uncertain, executes the suggested action during training, and distills guidance into a lightweight Reinforcement Learning (RL) policy. Because VLM advice is not always reliable, SAGE can weight teacher-action distillation using environment-derived advantages rather than treating all suggestions as equally useful. Across sparse-reward visual reasoning and navigation tasks, SAGE learns policies that act without VLM guidance at evaluation time and improves over unguided RL in several environments, including settings where the learned policy exceeds its VLM teacher. The results show that selective guidance is most beneficial when the VLM can help the agent discover high-reward trajectories, and less useful when unguided exploration already succeeds or teacher actions do not lead to informative experience. SAGE also reduces VLM usage by prompting the teacher only on a fraction of training steps and requiring no VLM calls at deployment. Overall, our results suggest that VLMs don't need to be used as fixed policies to be useful; they can instead act as temporary, imperfect sources of guidance whose value is tested and internalized through interaction.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @sage 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/selective-agent-guid…] indexed:0 read:1min 2026-09-02 ·