{"slug": "longcat-deepresearch-technical-report", "title": "LongCat-DeepResearch Technical Report", "summary": "LongCat released LongCat-DeepResearch, a deep research system pairing an enhanced LongCat model with a multi-agent workflow that separates global planning from section-level investigation, scoring 55.25 on DeepResearchBench, 51.35 on DeepResearchBench II, and 79.83 on ResearchRubrics. On an in-house benchmark the system scored 76.04, ranking second among four compared systems, and the workflow also generates research tasks and trajectories used for mid-training and post-training of LongCat's general-purpose models. Development-set analyses found combining planning perspectives helped while further planning refinement had mixed effects, and additional editing improved average automatic readability preference across two benchmarks with differing trends on each.", "body_md": "arXiv:2609.36071v1 Announce Type: new \nAbstract: We present LongCat-DeepResearch, a deep research system that combines an enhanced LongCat model with a multi-agent workflow for producing comprehensive, evidence-grounded reports. The workflow separates global planning from detailed investigation and coordinates revision at the section level. Multiple planning agents first explore external sources and refine an actionable research plan, termed ResearchSpec. Research agents then investigate and draft their assigned sections in parallel, gathering additional evidence in separate contexts as their analyses develop. Once the sections are assembled, global review guides targeted local revisions, reducing reliance on repeated full-report rewriting. This workflow also supports the construction of research tasks and trajectories for the mid-training and post-training of LongCat's general-purpose models. LongCat-DeepResearch achieves 55.25 on DeepResearchBench, 51.35 on DeepResearchBench II, and 79.83 on ResearchRubrics. On an in-house benchmark, it scores 76.04, ranking second among four compared systems. Development-set analyses show benefits from combining planning perspectives, while further planning refinement has mixed effects. Additional editing improves average automatic readability preference across two benchmarks, with different trends on each.", "url": "https://wpnews.pro/news/longcat-deepresearch-technical-report", "canonical_source": "https://arxiv.org/abs/2609.36071", "published_at": "2026-09-30 04:00:00+00:00", "updated_at": "2026-09-30 04:18:16.271431+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-agents", "ai-research"], "entities": ["LongCat", "LongCat-DeepResearch", "DeepResearchBench", "DeepResearchBench II", "ResearchRubrics", "ResearchSpec"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/longcat-deepresearch-technical-report", "markdown": "https://wpnews.pro/news/longcat-deepresearch-technical-report.md", "text": "https://wpnews.pro/news/longcat-deepresearch-technical-report.txt", "jsonld": "https://wpnews.pro/news/longcat-deepresearch-technical-report.jsonld"}}