{"slug": "cordisbench-can-language-models-reason-about-component-lifecycles-in-dynamic", "title": "CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harnesses?", "summary": "Researchers introduced CordisBench, a 1,200-question benchmark testing whether language models can reason about component lifecycles in dynamic agent harnesses, finding that models handle small systems well but become less reliable as interactions grow, with GPT-5.6 Luna using nearly 3,000 reasoning tokens per question at medium effort on the 16-interaction subset. The study, submitted to arXiv on 1 Sep 2026, evaluated three efficiency-oriented models at low reasoning effort with 2 to 32 relevant interactions, and noted that an independent finite reference semantics agreed with Cordis execution on all 528 executable questions, suggesting the cost is avoidable.", "body_md": "# Computer Science > Computation and Language\n\n[Submitted on 1 Sep 2026]\n\n# Title:CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harnesses?\n\n[View PDF](/pdf/2609.01600v1)\n\n[HTML (experimental)](https://arxiv.org/html/2609.01600v1)\n\nAbstract:Dynamic agent harnesses let language models change the software that shapes their own execution. This flexibility brings a new reasoning burden: a local plugin change can propagate through dependencies and cleanup. We introduce CordisBench, a 1,200-question benchmark of this lifecycle reasoning. It combines a controlled formal setting with programs executed against Cordis, a runtime that manages component dependencies and cleanup, and asks models to identify affected components, predict state after a specified teardown order, determine which conditions hold under all or some orders, and choose reconfigurations that succeed when executed. Across these tasks, we evaluate three efficiency-oriented models at low reasoning effort with 2, 4, 8, 16, 24, or 32 relevant interactions, using deterministic task-specific scoring. Models usually handle small systems well but grow less reliable as more interactions become relevant, especially when predicting final state and when reasoning across teardown orders. Additional inference effort recovers marked gains for some models. The cost is nontrivial: on our 16-interaction subset, GPT-5.6 Luna uses nearly 3,000 reasoning tokens per question at medium effort. For these controlled instances, that cost is avoidable: an independent finite reference semantics agrees with Cordis execution on every observation and action outcome used for scoring across all 528 executable questions.\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer\n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers\n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps\n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations\n\n*(*[What are Smart Citations?](https://www.scite.ai/))# Code, Data and Media Associated with this Article\n\nalphaXiv\n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers\n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub\n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub\n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face\n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast\n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower\n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender\n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [ Learn more about arXivLabs](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/cordisbench-can-language-models-reason-about-component-lifecycles-in-dynamic", "canonical_source": "http://arxiv.org/abs/2609.01600v1", "published_at": "2026-09-02 21:10:21+00:00", "updated_at": "2026-09-02 21:52:51.761417+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-agents"], "entities": ["CordisBench", "Cordis", "GPT-5.6 Luna", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/cordisbench-can-language-models-reason-about-component-lifecycles-in-dynamic", "markdown": "https://wpnews.pro/news/cordisbench-can-language-models-reason-about-component-lifecycles-in-dynamic.md", "text": "https://wpnews.pro/news/cordisbench-can-language-models-reason-about-component-lifecycles-in-dynamic.txt", "jsonld": "https://wpnews.pro/news/cordisbench-can-language-models-reason-about-component-lifecycles-in-dynamic.jsonld"}}