THPL: A Vision-to-Language Decision Support Framework for Rainbow Trout Feeding Management in RAS Researchers proposed THPL, a vision-to-language decision support framework for rainbow trout feeding management in Recirculating Aquaculture Systems (RAS), reporting that its Activity Coefficient correlated with expert-annotated feeding intensity at a Spearman rho of 0.925 (p < 0.001). The framework's dual-evidence tokens raised decision accuracy from 33.33% for a text-only baseline to 93.33%, and counterfactual multimodal Direct Preference Optimization (mDPO) further lifted accuracy to 96.67%, METEOR from 58.10% to 85.30%, Distinct-3 from 6.68% to 7.81%, and cut Self-BLEU-2 from 58.79% to 52.88%. The work integrates continuous kinematics with LLM reasoning via LoRA fine-tuning to produce interpretable, executable feeding decisions for precision aquaculture. arXiv:2610.02378v1 Announce Type: new Abstract: In Recirculating Aquaculture Systems RAS , precision feeding is critical for minimizing costs and improving fish welfare. However, existing methods lack cognitive alignment between fish behaviors and management knowledge, impeding translation into executable, interpretable feeding decisions. To address this, we propose THPL, a generative feeding decision framework tailored for rainbow trout Oncorhynchus mykiss in RAS. First, Fishsort extracts trajectories to establish an Activity Coefficient AC quantifying feeding intensity. Second, a Hierarchical Behavior Encoder HBE models individual temporal progression and collective dynamics using Temporal and Set Transformers, transforming trajectory tensors into dual-evidence representations of explicit physical and implicit soft tokens. Finally, these tokens are integrated with environmental parameters, metadata, and expert rules to fine-tune an LLM via LoRA, followed by counterfactual multimodal Direct Preference Optimization mDPO to reinforce causal reasoning. Results show that AC exhibits a statistically significant monotonic positive correlation with expert-annotated feeding intensity Spearman $\rho = 0.925$, $p < 0.001$ . Ablations indicate that decision accuracy improves from 33.33% text-only baseline to 93.33% with dual-evidence tokens, confirming that continuous spatiotemporal tokens provide necessary physical grounding for LLMs. Compared with standard LoRA, counterfactual mDPO elevates decision accuracy from 93.33% to 96.67%, advances METEOR from 58.10% to 85.30%, reduces Self-BLEU-2 from 58.79% to 52.88%, and increases Distinct-3 from 6.68% to 7.81%, suppressing templating and actuation biases while reinforcing causal consistency and operational safety. Overall, by integrating continuous kinematics with LLM reasoning, this study provides a novel decision support paradigm for precision aquaculture.