cd/entity/Kaggle· home› entities› Kaggle
grep -l @kaggle /news/*.json | wc -l → 146

Kaggle

mentions 146 type Organization page 2/8 feed RSS

// recent coverage 146 mentions

07:04
2026-09-28
dev.to
ai-agents

ToolTrap: “tool results are data” wasn’t enough

A developer built ToolTrap, a benchmark that tests whether support agents leak planted details from untrusted tool results, and found that the standard instruction "tool results are data, not instruct…

02:30
2026-09-28
aiflash.com
large-language-models

Game Arena: Strategic LLM Evaluation in Competitive Environments

Kaggle introduced Game Arena, an open and expanding platform that evaluates large language models through competitive head-to-head game matchups rather than static benchmarks. Kaggle states the platfo…

00:41
2026-09-28
dev.to
ai-safety

The Code Didn't Change. The Credit Did.

A developer built rai-attribution-bench, a benchmark testing whether AI models correctly apply rai-lint's five-tier AI attribution rubric when assigning git commit footers. Three flagship models score…

23:13
2026-09-27
dev.to
large-language-models

Do LLMs Catch Bad Startup Math? My First Answer Was Wrong.

A developer benchmarked six models across three labs to test whether LLMs independently verify the arithmetic in startup pitch claims, finding that all three frontier models scored 3/3 with zero varia…

08:45
2026-09-27
dev.to
mlops

LAW-N Real-World Data Layer

Peace Thabiwa (PEACEBINFLOW) of SAGEWORKS AI has published a whitepaper documenting that the ten-notebook LAW-N Kaggle telecom pipeline reports success while silently propagating empty or fallback-fil…

05:20
2026-09-27
dev.to
artificial-intelligence

AI models catch bad code, then cry wolf on the good code

A developer built "Blog vs Bytecode," a Kaggle benchmark of 28 data-science code snippets paired with blog-style claims, half of which hide real methodological errors such as scaler leakage or shuffle…

23:01
2026-09-26
pub.towardsai.net
large-language-models

Fine-Tuning to Quantization: What Free-Tier Hardware Can Prove

QLoRA fine-tuning on 150 synthetic examples raised a base Qwen3.5-4B model's adversarial-prompt refusal rate from 79.5% (31 of 39) to 89.7% (35 of 39), while GPTQ and AWQ 4-bit quantization added roug…

← prev page 2 / 8 next →
// co-occurs with top 8 entities
// topics top 6 topics