cd /news/artificial-intelligence/diagnosing-compositional-generalizat… · home topics artificial-intelligence article
[ARTICLE · art-86799] src=aclanthology.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Diagnosing Compositional Generalization in Transformers on ReCOGS with Compositional Graph Similarity

Researchers Bruno Franco and Edson Scalabrin introduced Compositional Graph Similarity (CGS), a graph-based metric for evaluating Transformer models on the ReCOGS benchmark, and found that the lowest-scoring categories are cp recursion (45.0%), obj pp to subj pp (65.4%), and prim to inf arg (66.7%). Follow-up experiments showed 0% Semantic Exact Match under depth extrapolation and constituent-role relocation, but 99.9% for prim to inf arg in isolation, indicating that Transformer limitations are partly structural and partly due to dataset distribution.

read2 min views13 publishedJul 31, 2026
Diagnosing Compositional Generalization in Transformers on ReCOGS with Compositional Graph Similarity
Image: Aclanthology (auto-discovered)
Abstract

This paper investigates the evaluation of compositional generalization in Transformer models on the ReCOGS benchmark. The problem addressed is that ReCOGS relies on Semantic Exact Match, a binary metric that assigns the same penalty to minor local mismatches and severe structural errors, limiting diagnostic interpretation. To address this, the study introduces Compositional Graph Similarity (CGS), a graph-based metric that compares predicted and reference semantic structures through explicit edit operations, providing graded and interpretable structural evaluation. The work also uses controlled synthetic datasets to test whether low-scoring ReCOGS categories reflect true model limitations or weaknesses in dataset coverage. Empirical results show that CGS satisfies all seven quality criteria adopted for graph similarity and identifies the lowest-scoring ReCOGS categories as cp recursion (45.0%), obj pp to subj pp (65.4%), and prim to inf arg (66.7%). Follow-up experiments showed 0% Semantic Exact Match under depth extrapolation and constituent-role relocation, but 99.9% Semantic Exact Match for prim to inf arg in isolation. These findings support the conclusion that CGS is more informative than Semantic Exact Match and that Transformer limitations in ReCOGS are partly structural and partly induced by dataset distribution.- Anthology ID:

- 2026.brigap-1.4
- Volume:
[Proceedings of the Third Workshop on the Bridges and Gaps between Formal and Computational Linguistics (BriGap-3)](/volumes/2026.brigap-1/)- Month:
  • July
  • Year:
  • 2026
  • Address:
  • Paris, France
- Editors:
[Timothée Bernard](/people/timothee-bernard/),[Emmanuele Chersoni](/people/emmanuele-chersoni/),[Giulia Rambelli](/people/giulia-rambelli/unverified/)- Venues:
[BriGap](/venues/brigap/)|[WS](/venues/ws/)- SIG:
- Publisher:
  • Association for Computational Linguistics
- Note:
- Pages:
  • 31–39
- Language:
- URL:
[https://aclanthology.org/2026.brigap-1.4/](https://aclanthology.org/2026.brigap-1.4/)- DOI:
- Cite (ACL):
[Diagnosing Compositional Generalization in Transformers on ReCOGS with Compositional Graph Similarity](https://aclanthology.org/2026.brigap-1.4/)(Franco & Scalabrin, BriGap 2026)- PDF:
[https://aclanthology.org/2026.brigap-1.4.pdf](https://aclanthology.org/2026.brigap-1.4.pdf)
── more in #artificial-intelligence 4 stories · sorted by recency
── more on @bruno franco 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/diagnosing-compositi…] indexed:0 read:2min 2026-07-31 ·