cd /news/ai-safety/dataset-ai-agent-security-failures-1… · home topics ai-safety article
[ARTICLE · art-111568] src=huggingface.co ↗ pub= topic=ai-safety verified=true sentiment=· neutral

Dataset: AI agent security failures, 1000 incidents classified

A dataset titled 'AI agent security failures' containing 1,000 classified incidents is currently inaccessible because the Hugging Face dataset viewer cannot parse its split names, throwing a SplitsNotFoundError due to a JSON parse error. The error originates from the datasets library's inability to read the dataset config, leaving the 1,000 incidents unviewable.

read2 min views2 publishedAug 26, 2026
Dataset: AI agent security failures, 1000 incidents classified
Image: Hugging Face Blog

Datasets:

The dataset viewer is not available for this subset. #

Exception:    SplitsNotFoundError
Message:      The split names could not be parsed from the dataset config.
Traceback:    Traceback (most recent call last):
                File "/usr/local/lib/python3.14/site-packages/datasets/packaged_modules/json/json.py", line 290, in _generate_tables
                  pa_table = paj.read_json(
                      io.BytesIO(batch), read_options=paj.ReadOptions(block_size=block_size)
                  )
                File "pyarrow/_json.pyx", line 342, in pyarrow._json.read_json
                File "pyarrow/error.pxi", line 155, in pyarrow.lib.pyarrow_internal_check_status
                File "pyarrow/error.pxi", line 92, in pyarrow.lib.check_status
                  raise convert_status(status)
              pyarrow.lib.ArrowInvalid: JSON parse error: Column() changed from object to string in row 0
              
              During handling of the above exception, another exception occurred:
              
              Traceback (most recent call last):
                File "/usr/local/lib/python3.14/site-packages/datasets/inspect.py", line 286, in get_dataset_config_info
                  for split_generator in builder._split_generators(
                                         ~~~~~~~~~~~~~~~~~~~~~~~~~^
                      StreamingDownloadManager(base_path=builder.base_path, download_config=download_config)
                      ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                  )
                  ^
                File "/usr/local/lib/python3.14/site-packages/datasets/packaged_modules/json/json.py", line 101, in _split_generators
                  pa_table = next(iter(self._generate_tables(**splits[0].gen_kwargs, allow_full_read=False)))[1]
                             ~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/packaged_modules/json/json.py", line 304, in _generate_tables
                  batch = json_encode_fields_in_json_lines(original_batch, json_field_paths)
                File "/usr/local/lib/python3.14/site-packages/datasets/utils/json.py", line 111, in json_encode_fields_in_json_lines
                  examples = [ujson_loads(line) for line in original_batch.splitlines()]
                              ~~~~~~~~~~~^^^^^^
                File "/usr/local/lib/python3.14/site-packages/datasets/utils/json.py", line 20, in ujson_loads
                  return pd.io.json.ujson_loads(*args, **kwargs)
                         ~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^
              ValueError: Expected object or value
              
              The above exception was the direct cause of the following exception:
              
              Traceback (most recent call last):
                File "/src/services/worker/src/worker/job_runners/config/split_names.py", line 68, in compute_split_names_from_streaming_response
                  for split in get_dataset_split_names(
                               ~~~~~~~~~~~~~~~~~~~~~~~^
                      path=dataset,
                      ^^^^^^^^^^^^^
                      config_name=config,
                      ^^^^^^^^^^^^^^^^^^^
                      token=hf_token,
                      ^^^^^^^^^^^^^^^
                  )
                  ^
                File "/usr/local/lib/python3.14/site-packages/datasets/inspect.py", line 340, in get_dataset_split_names
                  info = get_dataset_config_info(
                      path,
                  ...<6 lines>...
                      **config_kwargs,
                  )
                File "/usr/local/lib/python3.14/site-packages/datasets/inspect.py", line 291, in get_dataset_config_info
                  raise SplitsNotFoundError("The split names could not be parsed from the dataset config.") from err
              datasets.inspect.SplitsNotFoundError: The split names could not be parsed from the dataset config.

Need help to make the dataset viewer work? Make sure to review how to configure the dataset viewer, and open a discussion for direct support.

1,000+ classified incidents • Updated daily • CC BY 4.0

Pipeline entièrement local (Lenovo Legion, CPU only, Ollama + Llama 3.1 8B). Collecte automatique depuis NVD/CVE, GitHub Advisories, Hacker News et de multiples sources de sécurité.

Statistiques principales :

  • 1 000+ incidents acceptés
  • 82 incidents critiques
  • Top vendors : NVIDIA (48), OpenAI (46), TensorFlow (44)
  • Top types : api_exploit, unauthorized_action, configuration_exploit, data_exfiltration, prompt_injection, sandbox_escape, slopsploit_attack_chain (découvert par le modèle)

Toutes les entrées ont une source_url

vérifiable.

Dataset complet : https://huggingface.co/datasets/gemmozero/ai-agent-security-incidents

Code du pipeline : https://github.com/Legion33shadow/legion-n8n-shield

Feedback technique bienvenu (faux positifs, taxonomy, nouvelles sources).

Dernière mise à jour : 23 août 2026

Support this project #

If you find this dataset useful, consider supporting continued development: https://gemmo.gumroad.com/l/mdevxu

  • Downloads last month
  • 178
── more in #ai-safety 4 stories · sorted by recency
── more on @hugging face 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/dataset-ai-agent-sec…] indexed:0 read:2min 2026-08-26 ·