{"slug": "geosteer-geodesic-optimization-for-activation-steering-in-large-language-models", "title": "GEOSTEER: Geodesic Optimization for Activation Steering in Large Language Models", "summary": "Researchers introduced GeoSteer, an optimization-based method for norm-preserving activation steering in large language models that formulates steering as a Riemannian optimization problem and updates activations through a sequence of small geodesic steps on the representation manifold. GeoSteer learns a nonlinear activation-space objective to adaptively guide each steering step instead of relying on fixed steering directions, and it consistently improved over state-of-the-art activation steering baselines across the TruthfulQA, RealToxicityPrompts, and UltraFeedback benchmarks. The authors state the results suggest norm-preserving steering can be made more effective by replacing predefined one-step edits with adaptive, geometry-aware optimization.", "body_md": "arXiv:2609.10658v1 Announce Type: new \nAbstract: Activation steering provides a lightweight way to control large language models (LLMs) by modifying their hidden activations at inference time. Among these approaches, norm-preserving steering aims to change model behavior without altering the activation norm, reducing the risk of representation collapse and degradation. However, existing norm-preserving methods are limited by predefined steering trajectories and by their reliance on one-step updates, which may fail to capture the complex structure of activation distributions. We propose GeoSteer, an optimization-based method for norm-preserving activation steering. GeoSteer formulates steering as a Riemannian optimization problem and updates activations through a sequence of small geodesic steps on the representation manifold. To avoid fixed steering directions, GeoSteer learns a nonlinear activation-space objective that distinguishes desired from undesired activations, and uses this function to adaptively guide each steering step. This multistep formulation yields smoother, more stable, and more consistent steering behavior while preserving the activation norm. Across TruthfulQA, RealToxicityPrompts, and UltraFeedback benchmarks, GeoSteer consistently improves over state-of-the-art activation steering baselines. These results suggest that norm-preserving steering can be made more effective by replacing predefined one-step edits with adaptive, geometry-aware optimization.", "url": "https://wpnews.pro/news/geosteer-geodesic-optimization-for-activation-steering-in-large-language-models", "canonical_source": "https://arxiv.org/abs/2609.10658", "published_at": "2026-09-11 04:00:00+00:00", "updated_at": "2026-09-11 04:28:32.201894+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "machine-learning", "artificial-intelligence"], "entities": ["GeoSteer", "TruthfulQA", "RealToxicityPrompts", "UltraFeedback"], "alternates": {"html": "https://wpnews.pro/news/geosteer-geodesic-optimization-for-activation-steering-in-large-language-models", "markdown": "https://wpnews.pro/news/geosteer-geodesic-optimization-for-activation-steering-in-large-language-models.md", "text": "https://wpnews.pro/news/geosteer-geodesic-optimization-for-activation-steering-in-large-language-models.txt", "jsonld": "https://wpnews.pro/news/geosteer-geodesic-optimization-for-activation-steering-in-large-language-models.jsonld"}}