# Can an LLM Read a Fraud Network? Testing Gemma 4 31B

> Source: <https://dev.to/chef_p/can-an-llm-read-a-fraud-network-testing-gemma-4-31b-160n>
> Published: 2026-10-11 11:46:23+00:00

*This is a submission for the [Kaggle Benchmarking Challenge](https://dev.to/challenges/kaggle-2026-09-23).*

Can a language model use the relationships around a transaction to help identify fraud?

I built a benchmark using transaction data from the IEEE-CIS Fraud Detection dataset. Each example presents a target transaction alongside a small text representation of its local network neighborhood. The neighboring transactions are connected through shared attributes such as card, address, or device information. The model does not receive the ground-truth labels in its prompt.

The benchmark measures two behaviors: classifying the target transaction as fraudulent or legitimate, and identifying a potentially suspicious neighboring transaction. I was interested in whether a general-purpose language model could make useful predictions from graph context after that context had been translated into text.

I ran **Gemma 4 31B (`google/gemma-4-31b`)** on 100 samples, with both tasks evaluated for each sample. I chose it as the benchmark’s target language model to test whether a large general-purpose model could interpret transaction relationships from the text prompt.

I also compared its classification results with the project’s GraphSAGE graph-neural-network baseline. This was a single-model evaluation, not a broad ranking across language models.

The run completed with **200 non-empty responses**: 100 for transaction classification and 100 for suspicious-neighbor identification. A controlled probe returned text with no token override but an empty response when the 128-token cap was nested under `extra_body`. The completed run sent the cap as a top-level `max_tokens` parameter and produced all 200 responses.

Gemma achieved **55% classification accuracy** and **53% accuracy on the fraud-neighbor proxy**. The 100 classification labels were evenly split—50 fraudulent and 50 legitimate—so the majority-class baseline was 50%. At 55 correct out of 100, this result does not establish performance above that reference. No baseline comparison is claimed for the fraud-neighbor proxy.

On classification, the GraphSAGE baseline was correct on 67 of the 100 samples. Both systems were correct on 44; Gemma was correct when GraphSAGE was wrong on 11; and GraphSAGE was correct when Gemma was wrong on 23. Both missed 22. In this sample, the graph-based baseline outperformed Gemma.

The neighbor score needs careful interpretation. IEEE-CIS does not provide verified fraud-ring membership or ring-leader labels, so I scored the selected neighbor using its transaction fraud label. That makes 53% a **fraud-neighbor proxy score**, not evidence that the model can identify actual fraud rings or their leaders.

The main takeaway is that providing a local transaction neighborhood as text did not, by itself, produce strong results from this model on this sample. I would next test more models, increase the sample size, and run prompt ablations to measure how much the neighborhood details affect predictions.
