cd /news/ai-agents/someone-made-jev-play-atari-games · home topics ai-agents article
[ARTICLE · art-135253] src=github.com ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Someone made jev play Atari games

A developer released an open-source project that evaluates TypeSafe's Jev model, version jev-1.13.0, as a policy in Gymnasium and Atari environments, converting game observations into structured JSON so Jev selects a legal action through a single Choice question. The project supports CartPole, Pong, Breakout, and Ms. Pac-Man, with Montezuma's Revenge marked experimental and development paused, and records rewards and decisions without training or updating model weights, using a seeded random policy as a local baseline. Atari runs use the Arcade Learning Environment with Python 3.14, a committed uv.lock, and a trial wrapper defaulting to 64 decisions per game, 2 emulator frames per decision, and a 0.25 sticky-action probability.

read4 min views1 publishedSep 20, 2026
Someone made jev play Atari games
Image: Michielbdejong (auto-discovered)

Evaluate TypeSafe Jev as a policy in Gymnasium and Atari environments. Game-specific adapters convert observations into structured JSON; Jev selects a legal action through a single Choice question.

This project evaluates a fixed model. It records rewards and decisions without training or updating model weights. A seeded random policy provides a local baseline.

Three games, one shared action-selection question. Joysticks and probability bars show Jev's recorded decisions at 2× game speed. See recording details for the source runs, rendering command, and smaller gameplay-only GIF.

Environment → structured state → Jev Choice → legal action → environment
                   ↓                ↓
             RAM, frames, requests, responses, rewards → NPY / JSONL / GIF
Environment Observation adapter Status
CartPole Named position, velocity, angle, and angular velocity Supported
Pong RAM or RGB object detection Supported
Breakout RAM: ball, paddle, lives, and brick map Supported
Ms. Pac-Man RAM: maze, actors, food, and local navigation Supported
Montezuma's Revenge RAM and reference room geometry Experimental; development d

RAM and known room geometry provide privileged information compared with pixel-only Atari benchmarks. Adapter limitations are documented in observation adapters and the Montezuma prototype.

Install uv, then from the repository root:

uv sync --locked --extra atari --extra recording

uv run --locked --no-env-file --extra atari --extra recording main.py \
  --env ALE/Breakout-v5 --policy random --max-steps 500 \
  --log runs/breakout-random.jsonl --save-npy runs/breakout-random.npy

uv run --locked --no-env-file --extra recording npy_to_gif.py \
  runs/breakout-random.npy --scale 2

The project uses Python 3.14 and a committed uv.lock. Atari runs through Arcade Learning Environment. No ROM files or credentials are included in this repository. Raw experiment recordings stay local; media/ contains selected demonstration GIFs and a showcase preview. CartPole can run without the Atari extra:

uv run --locked --no-env-file main.py --env CartPole-v1 --policy random

Create .env from the blank template only if you do not already have one:

cp -n .env.example .env

Set TYPESAFE_API_KEY in your local .env or process environment. The .env file is ignored by Git. uv --env-file .env loads it at runtime; the Python scripts do not open it. CLI entry points suppress SDK/HTTP debug output and raw exception payloads. Response capture redacts runtime credentials and excludes headers/cookies.

Start with a bounded live trial:

uv run --locked --env-file .env --extra atari --extra recording run_jev_trial.py \
  --game breakout --max-steps 64 --output-dir runs/breakout-trial

Game choices: pong, breakout, mspacman, montezuma, or both (Pong and Breakout). The wrapper uses RAM observations, one episode, no HTTP retries, and exports a GIF automatically. Existing outputs are never overwritten; use a new output directory for each trial. Runs stop at game over, truncation, or the requested decision limit.

Setting Default
Model jev-1.13.0
Emulator frames per Atari decision 2
Sticky-action probability 0.25
Seed / episodes 7 / 1
General CLI decision limit 500
Trial-wrapper decision limit 64 per game
HTTP retries General CLI: 2; trial wrapper: 0

The general CLI exposes frame skipping, sticky actions, model, seed, episode count, and replay controls. Run main.py --help for details. Pong uses RGB by default in the general CLI; add --pong-state ram for RAM input.

The simulator s during each API request. Playback uses emulator time, so network latency affects runtime but not the recorded game speed. Each decision normally uses one API request; general-CLI retries can add requests. Check current TypeSafe pricing and limits before large runs, and use logged token usage to measure cost.

  • .npy : RGB frames, raw observations, exact structured states, actions, rewards, episode summaries, and redacted API responses.
  • .jsonl : transition records;.responses.jsonl : exact redacted API requests and response bodies, captured before validation.
  • .gif : game-time playback exported from the recording.

All local outputs belong under the ignored runs/ directory. NPY files contain pickled dictionaries; only load your own trusted recordings. Details and examples: recordings, replay, GIF conversion, and offline debugging.

Join GIFs side by side, stopping when the shortest finishes:

uv run --locked --no-env-file --extra recording join_gifs.py \
  runs/pong.gif runs/breakout.gif runs/mspacman.gif --output runs/combined.gif
uv sync --locked --extra atari --extra recording
uv run --locked --no-env-file ruff check .
uv run --locked --no-env-file ruff format --check .
uv run --locked --no-env-file --extra atari --extra recording python -m pytest -q

Tests use local emulators and mock HTTP transports. No live TypeSafe requests or API key are required. See CONTRIBUTING.md for adapter development and AGENTS.md for agent instructions.

The core modules are adapters.py and the game decoders (state), policies.py (action selection), runner.py (episode loop), recording.py (NPY), and responses.py (redacted capture and replay), under jev_rl/.

Built with Gymnasium, ALE, and TypeSafe. RAM decoding references and Montezuma room geometry draw on OCAtari. Its license and the installed TypeSafe skill's license are recorded in third-party notices.

── more in #ai-agents 4 stories · sorted by recency
── more on @typesafe 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/someone-made-jev-pla…] indexed:0 read:4min 2026-09-20 ·