At midnight, I gave Grok Bot a job.
Not a prompt.
A job.
Find software engineering roles that actually fit me, apply to the strongest matches, keep track of everything, and do it again in four hours.
Then I went to sleep.
That’s the part that got me.
For the last few years, we’ve been talking to AI.
Now we’re starting to assign work to it.
And after turning Grok Bot into an autonomous job-search agent, I think that distinction is much more interesting than another benchmark.
The idea was slightly unhinged:
Every 4 hours
Find new software engineering jobs
↓
Verify they're still open
↓
Compare them against my resume
↓
Score each opportunity
↓
Apply to the strongest matches
↓
Record what happened
↓
Repeat
Up to 10 applications per run.
But I added one important rule:
10 applications is a maximum, not a quota.
If there are three great jobs, apply to three.
If everything sucks, apply to zero.
I don’t want an AI desperately searching for application #10 and deciding that Senior Principal CUDA Kernel Wizard is probably close enough.
The objective isn't:
MAXIMIZE APPLICATIONS
It's:
MAXIMIZE INTERVIEW PROBABILITY
That completely changes how you design the agent.
Normally, building this myself would eventually become:
LLM
+
Playwright
+
Postgres
+
Cron
+
Session storage
+
Authentication
+
Retries
+
Monitoring
+
A weekend I'll never get back
Grok Bot removes a lot of that infrastructure.
It has an environment where the Bot can operate software, retain context, reuse skills, and run routines on a schedule. xAI defines a skill as reusable instructions for performing a task and a routine as the thing that tells a Bot when to run that workflow.
That means the interesting engineering problem becomes:
How do I define the job well enough for the agent to own it?
That's a very different problem.
This was probably the biggest lesson.
A weak instruction looks like:
Find 10 AI engineering jobs
and apply to them.
That's not autonomy.
That's a spam cannon.
The actual prompt I used was much closer to an operating policy.
I've removed personal details and shortened it, but the important parts looked like this:
You are my autonomous job-search and application agent.
Your goal is to find high-quality jobs matching my resume
and submit applications to strong matches.
Do not merely recommend jobs.
Use the attached resume as the source of truth.
Never fabricate experience, qualifications, employers,
education, dates, accomplishments, or technical skills.
TARGET
Prioritize roles such as:
- AI/ML Engineer
- Applied AI Engineer
- LLM Engineer
- Agentic AI Engineer
- AI Product Engineer
- Machine Learning Engineer
- Full-Stack AI Engineer
- Backend Engineer working on AI infrastructure
Prefer recently posted opportunities.
For every job:
1. Confirm the role is still open.
2. Compare the requirements against my resume.
3. Check location and compensation.
4. Evaluate technical and experience alignment.
5. Score the job from 0-100.
6. Automatically apply to strong matches.
7. Save borderline matches for review.
8. Skip weak matches.
Never apply just to hit the application target.
Avoid duplicate applications.
Do not apply to more than a small number
of roles at the same company.
GUARDRAILS
Automatically answer ordinary questions only when
the answer is already known.
for:
- uncertain information
- legal questions
- assessments
- references
- sensitive identity information
- financial information
- authentication or verification
- anything requiring information you cannot verify
Never invent an answer.
VERIFICATION
Do not mark an application as submitted until
you reach a real confirmation.
Track:
Company
Role
Match score
Application status
Confirmation
Notes
Recommended follow-up
Look at what that actually defines:
Goal
+
Context
+
Search strategy
+
Evaluation
+
Decision thresholds
+
Execution
+
Memory
+
Verification
+
Escalation
+
Guardrails
That's not really a prompt anymore.
It's closer to an agent specification.
And I think that's where prompting is heading.
We're moving from prompt engineering toward responsibility engineering.
Finding jobs isn't difficult.
LinkedIn can find jobs.
Indeed can find jobs.
Google can find jobs.
The interesting question is:
Should I actually apply to this one?
So I made Grok evaluate every opportunity before acting.
My actual weighting was roughly:
Technical alignment 35
Relevant AI/ML work 25
Experience level 15
Location / work arrangement 10
Compensation / company 10
Mission / product 5
──
100
Then:
75–100 → Strong enough to apply
65–74 → Save as secondary
<65 → Skip
Now the agent might discover 50 jobs and end up with:
50 discovered
↓
24 clearly irrelevant
↓
16 reasonable
↓
7 strong
↓
Apply to 7
That's success.
An autonomous agent without judgment is just a faster way to make bad decisions.
Without state, four hours later the agent wakes up and goes:
Wow, this role at Stripe looks great.
Four hours later:
Wow, this role at Stripe looks great.
Four hours later:
Have I told you how passionate I am about Stripe?
Not ideal.
So the agent needs an application ledger.
Company
Role
Posting
Score
Date discovered
Date applied
Status
Confirmation
Notes
Follow-up
Before every application:
Have we already seen this?
Have we already applied?
Is this the same role under another URL?
Did a previous attempt fail?
If yes:
SKIP
This stuff is boring.
It's also the stuff that makes autonomous systems useful.
State.
Idempotency.
Verification.
Retries.
Audit trails.
Everyone wants to talk about how intelligent the model is.
I increasingly care about whether the system around the model is dependable.
I'm perfectly comfortable letting an AI:
Then the form asks:
Do you have a security clearance?
Have you signed a non-compete?
Provide your Social Security number.
Complete this recorded interview.
Please electronically sign here.
Nope.
My boundary is basically:
KNOWN + LOW RISK
↓
HANDLE
UNCERTAIN
↓
ASK
SENSITIVE
↓
STOP
That's what good autonomy should look like.
Not:
Never bother me.
But:
Only bother me when human judgment is actually required.
xAI's own guidance for Grok Bot routines emphasizes approval boundaries, idempotent retries, reporting partial completion, and automating preparation before higher-impact execution.
That's the right direction.
Once the workflow works, the last step is almost boring.
Every 4 hours:
Search.
Verify.
Evaluate.
Score.
Apply.
Record.
Escalate when needed.
And now something changes.
I'm not waking up every morning and typing:
Find me some new AI engineering jobs.
Instead:
12 AM → work
4 AM → work
8 AM → work
12 PM → work
The responsibility exists independently of the prompt.
That's the part I keep thinking about.
Grok Bot isn't happening in isolation.
We're seeing the same pattern emerge across the industry.
ChatGPT Work can take a goal, operate across apps and files, stay with longer projects, produce finished work, and keep workflows moving through scheduled tasks.
Codex is pushing the same model into software engineering: give an agent a repository and a goal, let it build, refactor, test, review, and even take on scheduled background engineering work.
OpenAI is even introducing workspace agents designed around repeatable workflows that can run on schedules and act across tools with explicit permissions and approval checkpoints.
Different products.
Same direction.
OLD SOFTWARE
Human
↓
Instruction
↓
Software
↓
Result
is becoming:
AGENTIC SOFTWARE
Responsibility
+
Context
+
Memory
+
Tools
+
Schedule
+
Guardrails
↓
Continuous work
And that's a much bigger shift than "chatbots got smarter."
Would I let Grok Bot blindly submit 60 job applications every day?
Absolutely not.
That's a fantastic way to become the first person automatically rejected by every startup in America before breakfast.
But would I let it continuously:
Yes.
That's useful.
Really useful.
And the job-search experiment is almost beside the point.
The same architecture works for:
Sales prospecting
Lead qualification
Recruiting
Customer follow-up
Competitive research
Operations
Support
QA
Engineering maintenance
The pattern keeps showing up:
Context
+
Memory
+
Tools
+
A lane of responsibility
+
A schedule
+
Guardrails
We've spent years asking:
How smart can AI get?
I'm becoming more interested in another question:
How much responsibility can we safely give it?
Because once software can understand a goal, remember what happened, use tools, make bounded decisions, wake up on its own, and know when to ask for help...
We're not really talking about chatbots anymore.
We're talking about workers made of software.
That's exactly the idea behind Roster.
Instead of wiring together models, memory, tools, schedules, and guardrails yourself, Roster is built around giving an AI employee a real lane of responsibility and letting it handle recurring work while bringing you the decisions that actually need a human.
If you have work in your business that happens again and again, try handing it to an AI employee.