cd /news/large-language-models/give-an-llm-no-facts-and-ask-for-a-j… · home › topics › large-language-models › article
[ARTICLE · art-145268] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Give an LLM no facts and ask for a job ad. What does it commit to?

A developer built a Kaggle Benchmarks leaderboard that scores how often small, fast LLMs invent concrete specifics when asked to write a job ad from a fact-free brief. Claude Haiku 4.5 led at 0.50 and Gemini 3.7 Flash at 0.48, while Gemini 2.5 Flash (0.37) and GPT-5.4 mini (0.35) committed to fewer made-up details; Grok 4.5 errored out and was excluded because the task refuses to publish partial scores. The developer cautions that each score reflects a single roughly 30-claim ad, with a 95% interval of about 0.31 to 0.66 for Gemini 3.7 Flash, so the two pairs are a direction to check rather than a ranking.

by read4 min views1 publishedOct 5, 2026

This is a submission for the Kaggle Benchmarking Challenge. Job ads are long. The first one I read end to end. After that I scan. I look for the concrete parts: pay, the stack, the team size, where the work happens, what on-call looks like. Everything else is filler I have learned to skip.

I also do not apply because I match every line. I apply when something concrete catches my attention.

More and more of these ads are now drafted by LLMs. So I wanted to know: when a model writes a job ad and is given no facts at all, what does it commit to?

My main goal was to try Kaggle Benchmarks. I already had a scoring method from an earlier project, so I ported it into a Kaggle notebook. In about two days it went from an empty notebook to a public leaderboard.

The task. Every model gets the same brief: a role, a seniority level, and the topics the ad must cover. The brief on the current leaderboard is a junior backend engineer. The ad must cover the service, languages and data stores, code review, team size, location, pay, and the first six months. The brief names no employer and supplies no facts.

The model has three options for every topic:

The judge. A fixed judge model (Claude Sonnet 4.5, never a contestant) runs two stages:

The score is concrete claims divided by all claims, from 0 to 1.

The judge is not trusted on its own. Code enforces the rules after every call:

Higher is not better. The brief contains no facts, so every concrete claim is made up. The score measures how willing a model is to commit to specifics. Which behaviour you want depends on whether a human fills in the real facts afterwards.

The tier rules come from my earlier project, which scored real job postings. There I measured my own labeling consistency before trusting any model number: Before you report LLM agreement, measure the human twice.

I picked the small, fast tier from each vendor. These are the models a company would plug into an HR tool to draft ads at volume. Cost and speed matter there more than peak quality.

Model Score Cost per run Time per run
Claude Haiku 4.5 0.50 $0.037 29 s
Gemini 3.7 Flash 0.48 $0.062 44 s
Gemini 2.5 Flash 0.37 $0.074 60 s
GPT-5.4 mini 0.35 $0.069 45 s
Grok 4.5 error — —

Cost and time are as Kaggle reports them for one run: the ad plus both judge calls. Most of the token count is the judge, not the model under test.

The judge is kept out of the contest on purpose. A model grading its own family's writing is a conflict I did not want to explain.

What one ad looks like. Gemini 3.7 Flash wrote a 2,528-character ad. The judge extracted 29 claims:

No blanks at all. Every topic got either an invented specific or a soft phrase. It never wrote "[salary range]".

The leaderboard splits into two pairs. Claude Haiku 4.5 and Gemini 3.7 Flash commit to specifics in about half their claims. Gemini 2.5 Flash and GPT-5.4 mini sit around a third.

Do not read it as a ranking yet. The current leaderboard runs one brief, so each score is one ad of roughly 30 claims. For Gemini 3.7 Flash, 14 concrete out of 29 gives a 95% interval of about 0.31 to 0.66. Every model on the board sits inside that interval. The two pairs are a direction to check, not a result.

The judge is unmeasured on these ads. On real postings, an earlier version of the rules reached 86.7% agreement with my hand labels on the concrete/not-concrete boundary. That was a different judge model and older prompts. Nobody has hand-checked this judge on generated ads.

Grok 4.5 failed, and that is by design. The task refuses to publish a score unless every brief is scored. A partial average across a subset of briefs is a different benchmark, and it would sit on the same leaderboard as if it were not.

What surprised me on the Kaggle side:

<think> text before their JSON. The built-in schema parser then fails, so the judge parses JSON itself.%choose keeps only the latest run of the main task. Run it once in the notebook, then add models from the task page./. Sanitize them before using them in file names. Why model choice matters here. Ads are many and they are long. Getting from an ad to its actual claims takes effort, and that effort goes to a model on both sides. The employer's model decides what gets committed. The candidate's model decides what counts as concrete. Pick the wrong one on the writing side and you get invented salaries. Pick the wrong one on the reading side and "competitive salary" passes as a commitment.

What I would measure next:

Briefs, prompts and judge code are in the notebook. The method behind the tiers is in the repo: github.com/kargut/job-posting-specificity. When you read a job ad, what makes you apply: the concrete details, or something else entirely?

── more in #large-language-models 4 stories · sorted by recency
── more on @kaggle 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/give-an-llm-no-facts…] indexed:0 read:4min 2026-10-05 · —