arXiv:2610.00713v1 Announce Type: new Abstract: When a graph neural network (GNN) explainer produces an unexpected attribution on a molecule, the attribution alone cannot reveal whether the explainer has failed or the model has learned a shortcut. We introduce WOMBAT, a benchmark of 14 whitebox GNNs, each with message-passing weights set by hand to detect a specific SMARTS motif. Each model's decision rule is known by construction, providing attribution ground truth against which explainer errors can be identified and studied. We validate the models on millions of PubChem molecules and evaluate post-hoc explainers including GNNExplainer, PGExplainer, and Integrated Gradients. Guided by our qualitative analysis, we construct a model that causes Integrated Gradients to spread attribution across the graph, even though the model reliably detects the intended motif. We release the dataset, models, and evaluation code to help researchers in the development of newer XAI tools for GNNs.
WOMBAT: Whitebox Oracle for Molecular Benchmarking and Attribution Testing
Researchers introduced WOMBAT, a benchmark of 14 whitebox graph neural networks (GNNs) whose message-passing weights are hand-set to detect a specific SMARTS motif, providing attribution ground truth for evaluating explainers. Validated on millions of PubChem molecules, WOMBAT was used to test post-hoc explainers including GNNExplainer, PGExplainer, and Integrated Gradients, and the team constructed a model that causes Integrated Gradients to spread attribution across the graph even though the model reliably detects the intended motif. The dataset, models, and evaluation code were released to support development of new XAI tools for GNNs.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.