{"slug": "5-content-lanes-one-watchdog-how-i-stopped-wondering-if-my-automation-still-runs", "title": "5 Content Lanes, One Watchdog: How I Stopped Wondering If My Automation Still Runs", "summary": "A developer built a watchdog shell script to monitor five content generation lanes and restart any that fail, replacing daily manual log checks with a done-marker pattern. The script, content-watchdog.sh, checks for completion files and relaunches incomplete tasks, ensuring automation reliability and freeing the developer from verification overhead.", "body_md": "Every morning I used to open my logs and ask the same question: did it actually run last night? The scripts always reported success. The articles were sometimes 186 bytes of an error message. This is how I replaced that daily anxiety with a single shell script.\n\nWhen I first built a content generation pipeline, the first problem I hit was this: I thought it was running, but it had actually stopped.\n\nlaunchd fires the script every morning, but it starts before the network is up and exits silently. The API times out with no response, yet a log file still exists. Claude's token budget runs dry, and the error message flows straight into the output file, leaving the body at zero bytes. All of these failures leave behind nothing but the fact that \"the script ran.\"\n\nWhat does it mean to build an **environment** rather than do work? My answer was a design principle: prove completion by the existence of a file. Not what the script wrote to the log — only whether `~/.claude/logs/.article-daily-done-20260710`\n\nexists is treated as truth. That's the essence of the done-marker pattern.\n\n`content-watchdog.sh`\n\nis what happens when you extend that idea across all five content lanes. It gets invoked multiple times a day and does one simple job: check the done-markers for every lane, and restart only the ones that are missing. It doesn't break precisely because it's simple.\n\nAutomation gets stuck for operational reasons more often than technical ones. If you design on the assumption that \"it worked yesterday, so it'll work today,\" it quietly dies on the morning the Wifi isn't connected, at midnight when the battery is at 3%, at the end of the month when the budget runs out. Having a watchdog freed me from the nagging worry that \"it should still be running today.\" The reason I can keep 10 iOS apps going in parallel and hold ¥1.2M/month in revenue is that the time spent on verification is as close to zero as it gets.\n\nHere's the relationship between the watchdog and the individual lane scripts.\n\n```\nlaunchd (複数スロット)\n    │\n    └─→ content-watchdog.sh (sweep モード)\n            │\n            ├─ acquire_lock()       # mkdir 競合ロック\n            │       └─ ~/.claude/locks/content-watchdog.lockd/\n            │\n            ├─ for lane in article note maker series ameba\n            │       │\n            │       ├─ done_lane()  # done-marker / .done ファイルを確認\n            │       │\n            │       └─ [未完了なら] run_capped 1800 bash <script> apply\n            │                │\n            │                └─ 各レーンスクリプトが自前のdone-markerを立てる\n            │                       例: ~/.claude/logs/.article-daily-done-20260710\n            │\n            └─ [全レーン完了なら] send_heartbeat_once()\n                    └─ Discord通知 + ~/.claude/logs/.content-watchdog-heartbeat-20260710\n```\n\nLet's read through the actual code in order.\n\nThe `done_lane()`\n\nfunction has different check logic for each of the five lanes.\n\n```\ndone_lane() {\n  local lane=\"$1\"\n  local hit\n  case \"$lane\" in\n    article)\n      [ -f \"$LOG_DIR/.article-daily-done-$TODAY\" ]\n      ;;\n    note)\n      hit=\"$(find \"$LOG_DIR/note-daily\" -maxdepth 1 -name \"$TODAY-*.done\" -print -quit 2>/dev/null || true)\"\n      [ -n \"$hit\" ]\n      ;;\n    maker)\n      hit=\"$(find \"$LOG_DIR/maker-daily\" -maxdepth 1 -name \"$TODAY-*.done\" -print -quit 2>/dev/null || true)\"\n      [ -n \"$hit\" ]\n      ;;\n    series)\n      [ -f \"$LOG_DIR/.series-daily-done-$TODAY\" ]\n      ;;\n    ameba)\n      local ad td f\n      ad=\"$HOME_DIR/Desktop/Article/ameba\"\n      td=\"$(date +%F)\"\n      grep -rlq \"created:.*$td\" \"$ad\" 2>/dev/null && return 0\n      for f in \"$ad\"/*.md; do\n        [ -f \"$f\" ] || continue\n        [ \"$(stat -f %Sm -t %F \"$f\" 2>/dev/null)\" = \"$td\" ] && return 0\n      done\n      return 1\n      ;;\n  esac\n}\n```\n\n`article`\n\nand `series`\n\nare managed with a single hidden file (`.article-daily-done-20260710`\n\n). `note`\n\nand `maker`\n\nuse a `20260710-*.done`\n\npattern, a design that allows multiple files to exist. Only `ameba`\n\nhas no done-marker and instead checks the actual artifact directly (the `created:`\n\nmetadata in `.md`\n\nfiles, or the mtime). This is a fallback implementation, needed because the ameba script has no done-marker spec of its own, and it functions as \"a realistic compromise for wiring an existing script that lives outside the design into the watchdog.\"\n\nLooking at the implementation of `article-daily-stock.sh`\n\n, the timing of the done-marker is clear.\n\n```\nDONE_MARKER=\"$HOME/.claude/logs/.article-daily-done-${TODAY}\"\n```\n\nThe path is fixed at the top of the script, and after the article body has passed generation, validation, stock placement, and secret scanning, the marker is set **before** the git push (line 453).\n\n```\n# 生成成功＝この時点で当日doneを確定する。以降のgit pushはbest-effort（失敗しても生成は成功扱い）\n# なので、git失敗でマーカー未設定→翌スロットで重複生成、という事故を防ぐためここで先に立てる。\ntouch \"$DONE_MARKER\"\n```\n\nThe comment says it's \"to prevent the accident of git failure → marker unset → duplicate generation in the next slot.\" As long as push is best-effort, judging done by push success causes duplicate generation. The design principle that \"generation and delivery are independent responsibilities\" shows up right here.\n\nSince the watchdog itself can be launched from multiple slots, it uses `mkdir`\n\nfor mutual exclusion.\n\n```\nacquire_lock() {\n  if mkdir \"$LOCKDIR\" 2>/dev/null; then\n    printf '%s\\n' \"$$\" > \"$LOCKDIR/pid\"\n    trap release_lock EXIT INT TERM\n    return 0\n  fi\n\n  local now mod age\n  now=\"$(date +%s)\"\n  mod=\"$(stat -f %m \"$LOCKDIR\" 2>/dev/null || printf '%s\\n' \"$now\")\"\n  age=$((now - mod))\n  if [ \"$age\" -ge 1800 ]; then\n    log \"lock stale age=${age}s; taking over\"\n    rm -rf \"$LOCKDIR\"\n    if mkdir \"$LOCKDIR\" 2>/dev/null; then\n      printf '%s\\n' \"$$\" > \"$LOCKDIR/pid\"\n      trap release_lock EXIT INT TERM\n      return 0\n    fi\n  fi\n\n  log \"lock held; skip\"\n  exit 0\n}\n```\n\n`mkdir`\n\nis an atomic operation at the POSIX level. Even if two processes call `mkdir`\n\nsimultaneously, only one succeeds. No flock needed, no Linux/macOS compatibility issues, and the simplicity of `exit 0`\n\nimmediately on failure is its strength.\n\nIf the lock has been sitting untouched for 1800 seconds (30 minutes) or more, it judges that \"the previous process died abnormally and left the lock behind\" and forcibly takes over. The same 1800 seconds is used as the timeout for each lane's script invocation, so the next watchdog won't start unless a healthy watchdog has just finished its run.\n\n`article-daily-stock.sh`\n\nalso has its own lock using the same approach.\n\n```\nLOCKDIR=\"$HOME/.claude/locks/article-daily.lock\"\nif ! /bin/mkdir \"$LOCKDIR\" 2>/dev/null; then\n  oldpid=$(cat \"$LOCKDIR/pid\" 2>/dev/null || true)\n  if [ -n \"${oldpid:-}\" ] && kill -0 \"$oldpid\" 2>/dev/null; then\n    log \"別インスタンス実行中(pid=$oldpid) — skip\"; exit 0\n  fi\n  rm -rf \"$LOCKDIR\"; /bin/mkdir \"$LOCKDIR\" 2>/dev/null || exit 0\nfi\necho $$ > \"$LOCKDIR/pid\"\ntrap 'rm -rf \"$LOCKDIR\"' EXIT INT TERM\n```\n\nThis one adds a PID liveness check. If the previous process is alive it genuinely skips; if it's dead, it force-releases the lock and runs itself. Because the watchdog's lock and each lane script's lock exist as two separate layers, \"the watchdog is trying to restart\" and \"the previous article generation process is still running\" don't interfere with each other.\n\nThe design of notifying on failure but only once a day on success is baked into `send_heartbeat_once()`\n\n.\n\n```\nsend_heartbeat_once() {\n  [ \"$MODE\" = \"sweep\" ] || return 0\n  if [ -f \"$HEARTBEAT_MARKER\" ]; then\n    log \"heartbeat skip already_sent=1\"\n    return 0\n  fi\n  notify alerts \"💓 content-watchdog heartbeat: 全レーン当日生成済み\"\n  touch \"$HEARTBEAT_MARKER\"\n  log \"heartbeat sent\"\n}\n```\n\nThe watchdog gets invoked many times a day. If all lanes are healthy, it sends a `💓`\n\nto Discord only the first time and short-circuits on the marker afterward. Failure notifications, by contrast, fire every time.\n\n```\nif [ -n \"$FAILED_LANES\" ]; then\n  log \"result=unhealthy lanes=$FAILED_LANES\"\n  if [ \"$MODE\" = \"sweep\" ]; then\n    notify alerts \"🚨 content停止: $FAILED_LANES 当日未生成(自己修復不能)\"\n  fi\n```\n\nIf the done-marker still isn't set after attempting auto-repair (restarting inside `run_lane`\n\n, then checking `done_lane()`\n\nonce more), it tells Discord that manual intervention is required.\n\nLooking at the internals of `run_lane`\n\n, the \"restart → recheck\" cycle fits into a single function.\n\n```\nrun_lane() {\n  local lane=\"$1\"\n  local script rc\n  script=\"$(script_for \"$lane\")\"\n\n  if done_lane \"$lane\"; then\n    log \"lane=$lane status=healthy action=skip\"\n    return 0\n  fi\n\n  if [ ! -f \"$script\" ]; then\n    log \"lane=$lane status=missing script=$script\"\n    return 1\n  fi\n\n  log \"lane=$lane status=missing-done action=reinvoke script=$script timeout=1800s\"\n  if [ \"$lane\" = \"ameba\" ]; then\n    run_capped 1800 bash \"$script\" >> \"$LOG\" 2>&1\n  else\n    run_capped 1800 bash \"$script\" apply >> \"$LOG\" 2>&1\n  fi\n  rc=$?\n  log \"lane=$lane reinvoke_exit=$rc\"\n\n  if done_lane \"$lane\"; then\n    log \"lane=$lane status=healthy-after-reinvoke\"\n    return 0\n  fi\n\n  log \"lane=$lane status=failed-after-reinvoke\"\n  return 1\n}\n```\n\nThe flow is: check with `done_lane`\n\n→ skip if fine → restart if not → check with `done_lane`\n\nagain. The exit code of the restart (`rc`\n\n) is written to the log, but it isn't used for the decision. The strictness of \"even with exit code 0, no done-marker means failure\" is what catches the case where a script spits an error into the output file and exits normally.\n\nI explained earlier that `run_lane()`\n\nuses the done-marker as its trust basis. So what gets checked before the done-marker is set? The validation layer in `article-daily-stock.sh`\n\nis the answer.\n\nFirst, `article_ok()`\n\nchecks the file's existence and its contents.\n\n```\nMIN_ARTICLE_BYTES=1200\n\narticle_ok() {\n  local f=\"$1\"\n  [ -s \"$f\" ] || return 1\n  [ \"$(stat -f%z \"$f\" 2>/dev/null || echo 0)\" -ge \"$MIN_ARTICLE_BYTES\" ] || return 1\n  grep -qE '^title:' \"$f\" || return 1\n  grep -qiE 'request timed out|不明な商品|TODO: *本文|\\(生成失敗\\)' \"$f\" && return 1\n  return 0\n}\n```\n\nFour guards are chained in series: **does the file exist, is it at least 1200 bytes, does the frontmatter have a title line, and is it free of timeout wording or generation-failure phrases?** Keep that last `grep -qiE`\n\nin mind — it connects directly to a story about getting stuck later.\n\nThumbnail verification is handled by a separate function, `thumb_ok()`\n\n.\n\n```\nthumb_ok() {\n  local f=\"$1\" w; [ -s \"$f\" ] || return 1\n  w=$(sips -g pixelWidth \"$f\" 2>/dev/null | awk '/pixelWidth/{print $2}')\n  [ -n \"$w\" ] && [ \"$w\" -ge 2000 ] 2>/dev/null\n}\n```\n\n`sips`\n\nis macOS's built-in image tool. Anything under 2000 pixels wide isn't accepted as a thumbnail. Even if the file exists, zero-byte or corrupted PNGs get rejected.\n\n`audit_repair()`\n\nuses these two functions to re-verify the entire stock on every launch (it's a 312-line monster of a function, so I'll quote only the key parts).\n\n```\naudit_repair() {\n  for f in \"$STOCK_ART\"/[0-9][0-9]-*.md; do\n    if [ \"$t_ok\" = \"missing\" ] && [ \"$a_ok\" = \"ok\" ]; then\n      if regen_thumb \"$slug\" \"$no\" \"$f\"; then\n        t_ok=ok; fixed=$((fixed+1)); log \"FIX: サムネ再生成 $no-$slug\"\n      fi\n    fi\n    [ \"$a_ok\" = \"broken\" ] && broken+=(\"$no-$slug\")\n  done\n  for nb in \"${broken[@]:-}\"; do\n    n=$(bump_attempt \"$slug\")\n    if [ \"$n\" -le \"$MAX_ATTEMPTS\" ]; then\n      jq --argjson t \"$topic\" '[$t] + (map(select(.slug != ($t.slug))))' \"$QUEUE\" > ...\n      log \"REQUEUE: $slug 再生成キューへ投入(試行 $n/$MAX_ATTEMPTS)\"\n    else\n      needhuman+=(\"$nb (再生成${n}回失敗)\")\n    fi\n  done\n}\n```\n\n`MAX_ATTEMPTS=2`\n\nis declared as a constant (line 136). If a thumbnail is broken it's regenerated automatically; if the body is broken it's returned to the queue for regeneration up to two times, and on the third it notifies a human and stops. This is the self-repair that keeps things from being \"left stuck.\"\n\n`audit_repair()`\n\nruns every time, **before** article generation (line 246). Even when the generated flag causes today's generation to be skipped, the audit keeps running. The core of the design is that new generation and stock quality assurance run on independent cycles.\n\n```\naudit_repair\nif [ \"$MODE\" = \"audit\" ] || [ \"$SKIP_GEN\" = \"1\" ]; then\n  log \"===== article-daily done(audit$([ \"$SKIP_GEN\" = \"1\" ] && echo '+gen-skipped')) =====\"\n  exit 0\nfi\n```\n\nEven when launchd calls the script every morning, the environment right after boot is less stable than you'd think. There are three mechanisms.\n\n**1. caffeinate: keep the Mac from sleeping**\n\n```\nif [ -z \"${CAFFEINATED:-}\" ]; then\n  exec /usr/bin/caffeinate -i -s env CAFFEINATED=1 /bin/bash \"$0\" \"$@\"\nfi\n```\n\nIt uses `exec`\n\nto relaunch itself wrapped in `caffeinate`\n\n. Passing the environment variable `CAFFEINATED=1`\n\nprevents a double exec. `-i`\n\nis the flag that prevents idle sleep while the process is alive, and `-s`\n\nprevents system sleep. Without this, the Mac falls asleep mid-generation when running on battery late at night.\n\n**2. Network wait: don't call the API before Wifi is up**\n\n```\nif [ \"$MODE\" != \"audit\" ]; then\n  for _ in $(seq 1 18); do\n    /usr/bin/nc -z -G 3 1.1.1.1 443 2>/dev/null && break; sleep 5\n  done\nfi\n```\n\nIt uses `nc`\n\nto check connectivity to port 443 on 1.1.1.1, repeating up to 18 times (90 seconds) until it connects. `-G 3`\n\nis a 3-second connect timeout. Since launchd sometimes calls the script right after the Mac boots, this handles the case where hitting Claude's API before the Wifi connection is established wipes everything out. Audit-only mode skips it because it doesn't use Claude.\n\n**3. Budget check: do nothing if tokens are exhausted**\n\n```\nBUDGET=$(~/.claude/scripts/token-budget-advisor.sh --short 2>/dev/null || echo \"n/a\")\nlog \"budget: $BUDGET\"\nif [ \"$MODE\" != \"audit\" ] && echo \"$BUDGET\" | grep -qE '🔴|critical|cap-near'; then\n  log \"ABORT: budget critical — 次スロットで再試行\"; exit 0\nfi\n```\n\n`token-budget-advisor.sh`\n\nreturns the token consumption status, and if it contains `🔴`\n\nor `critical`\n\n, the script exits without doing anything. Even if the budget runs out at the end of the month, the watchdog automatically restarts things at the next slot (after the monthly reset), so there's nothing to do by hand.\n\n`run_capped()`\n\nin `content-watchdog.sh`\n\nlooks simple at a glance, but it's an important wrapper that absorbs environment differences.\n\n```\ntimeout_bin() {\n  if [ -x /opt/homebrew/bin/gtimeout ]; then\n    printf '%s\\n' /opt/homebrew/bin/gtimeout\n  else\n    command -v gtimeout 2>/dev/null || true\n  fi\n}\n\nrun_capped() {\n  local limit=\"$1\"\n  shift\n  local tb\n  tb=\"$(timeout_bin)\"\n  if [ -n \"$tb\" ]; then\n    \"$tb\" \"$limit\" \"$@\"\n  else\n    log \"gtimeout unavailable; running without timeout: $*\"\n    \"$@\"\n  fi\n}\n```\n\nIt checks the homebrew absolute path first, and if that's not there, searches PATH with `command -v`\n\n. If neither works, it runs without a timeout while logging that fact. The key point is `|| true`\n\n, which keeps it from erroring — preventing the accident of \"the entire watchdog dies because gtimeout is missing.\" `article-daily-stock.sh`\n\nhas a `run_to`\n\nhelper built on the same philosophy.\n\n```\nTIMEOUT_BIN=\"/opt/homebrew/bin/gtimeout\"\n[ -x \"$TIMEOUT_BIN\" ] || TIMEOUT_BIN=\"/opt/homebrew/bin/timeout\"\n[ -x \"$TIMEOUT_BIN\" ] || TIMEOUT_BIN=\"\"\nrun_to() { local s=$1; shift; if [ -n \"$TIMEOUT_BIN\" ]; then \"$TIMEOUT_BIN\" --kill-after=30 \"$s\" \"$@\"; else \"$@\"; fi; }\n```\n\nThis one also adds `--kill-after=30`\n\n, giving a 30-second grace period before sending SIGKILL after the timeout. Claude's process doesn't die instantly when it receives `SIGTERM`\n\n, so this two-stage approach is necessary.\n\nArticle topics are stored in `~/.topic-queue.json`\n\n. When this queue empties, the script doesn't stop — it uses `claude -p`\n\nto automatically draft one topic and add it to the queue.\n\n```\nQLEN=$(jq 'length' \"$QUEUE\" 2>/dev/null || echo 0)\nif [ \"$QLEN\" -eq 0 ]; then\n  log \"queue 空 → 実作業からネタ自動立案\"\n  # ...長いプロンプト組み立て...\n  TOPIC_JSON=$(run_to 600 \"$CLAUDE\" -p \"$REPLENISH_PROMPT\" \\\n    --model \"${ARTICLE_MODEL:-claude-sonnet-4-6}\" --effort high \\\n    --output-format text --allowedTools \"Read,Grep,Glob,Bash\" --max-turns 20 2>>\"$LOG\")\n```\n\n`--max-turns 20`\n\nmatters. It's the cap on how many round-trip turns Claude gets to Read actual files before producing output for topic planning. Without it, Claude investigates too thoroughly and burns through tokens.\n\nIt also checks whether the drafted topic duplicates an existing slug.\n\n```\nif used_slugs | grep -qx \"$NEW_SLUG\"; then\n  log \"ABORT: 立案slug '$NEW_SLUG' が既出 → 次スロット再試行\"; exit 0\nfi\n```\n\n`used_slugs()`\n\ncollects slugs from four sources — the queue, the completed list, existing article files, and coverage.json — and checks for duplicates. Without it, you end up mass-producing articles on exactly the same topic (this actually happened; more on that below).\n\nThis was the first failure. The launchd log recorded a normal exit. The file existed. But when I opened it, this was all that was inside.\n\n```\n---\ntitle: \"Claude Codeを使った自動化\"\nemoji: \"🤖\"\ntype: \"tech\"\npublished: true\n---\n\nrequest timed out\n```\n\nThe autonomous Claude Code session was cut off by a 30-second network timeout, and that error message flowed into the output file. The script's exit code was 0. The `[ -s \"$ART\" ]`\n\ncheck passed too. Because this state was treated as \"success,\" the next morning's watchdog decided \"done-marker present, skip\" and the article was never repaired.\n\nThe fix was adding `article_ok()`\n\n. I added `grep -qiE 'request timed out'`\n\nto the validation conditions; if the phrase is present in the body, it's treated as a generation failure and the file is deleted. The script exits without setting the done-marker, so the watchdog restarts it at the next slot.\n\nLesson: **exit code 0 is proof that the script ran, not proof that the output is correct.** You can't tell whether generation succeeded without looking at the contents.\n\nIn the original implementation, a successful git push was the condition for setting the done-marker. It sounds logical — the idea that \"done includes publishing.\"\n\nBut it broke at the end of the month. I hit GitHub's API rate limit and pushes failed repeatedly. The watchdog then decided \"no done-marker, restart\" at every slot, and three articles on the same topic were generated on the same day. The contents were subtly different. The filenames had sequential numbers. It was a mess.\n\nA comment in the code records that history (line 453).\n\n```\n# 生成成功＝この時点で当日doneを確定する。以降のgit pushはbest-effort（失敗しても生成は成功扱い）\n# なので、git失敗でマーカー未設定→翌スロットで重複生成、という事故を防ぐためここで先に立てる。\ntouch \"$DONE_MARKER\"\n```\n\nNow the queue pop, the record into the done-queue, and the marker placement all happen before the git push. Push is nothing more than \"best-effort delivery to the Zenn repo,\" an independent responsibility from article generation and stock placement. Just in case, `touch \"$DONE_MARKER\"`\n\nis also called again right after the git push (line 468). It's an idempotent operation, so the cost is zero.\n\nOne night while running my MacBook Air on battery, article generation kicked off at 2 AM and the OS went to sleep. launchd started the process, but disk I/O stopped, the Claude session was interrupted, and the article ended mid-way. The file exists, it's 500 bytes, but the body is cut off. `article_ok()`\n\nrejected it for being \"under 1200 bytes\" so it was regenerated the next day — but I didn't notice until then.\n\nAfter adding `caffeinate`\n\n, the Mac no longer sleeps while article generation is running. Because of the `-s`\n\nflag (system sleep prevention), it doesn't sleep even with the lid closed. \"Automatic generation every morning\" simply didn't hold together without this mechanism.\n\nThe `generate.sh`\n\nfor ameba wasn't written by me; it was an existing project I bolted the watchdog onto afterward. That script had no done-marker mechanism.\n\nTrying to check for `.ameba-daily-done-20260710`\n\nthe same way as the other four lanes, that file would never exist. The watchdog judged \"not done\" every time and kept restarting.\n\nSo I had no choice but to give the ameba branch of `done_lane()`\n\na different check method.\n\n```\nameba)\n  local ad td f\n  ad=\"$HOME_DIR/Desktop/Article/ameba\"\n  td=\"$(date +%F)\"\n  grep -rlq \"created:.*$td\" \"$ad\" 2>/dev/null && return 0\n  for f in \"$ad\"/*.md; do\n    [ -f \"$f\" ] || continue\n    [ \"$(stat -f %Sm -t %F \"$f\" 2>/dev/null)\" = \"$td\" ] && return 0\n  done\n  return 1\n  ;;\n```\n\nIt greps for whether an md file with today's `created:`\n\nmetadata exists, and if not, uses stat to check whether there's a file with today's mtime. A two-stage fallback. I wrote in a code comment that it's \"a realistic compromise for pulling an existing script that lives outside the design into the watchdog,\" and that's genuinely what it is — the ideal would be to add a done-marker on the generate.sh side. But I compared the cost of touching an existing script against the cost of running with a fallback, and chose the fallback.\n\nThis is from before I added slug duplicate checking. The queue emptied and the automatic topic refill ran. The slug Claude produced was `claude-code-automation`\n\n. It already existed in the done-queue. With no duplicate check, it was added to the queue, and the next day the same topic was generated again. When I checked 10 days later, `01-claude-code-automation.md`\n\nthrough `10-claude-code-automation.md`\n\nwere all lined up. The contents differ slightly each time.\n\nThe `used_slugs()`\n\nfunction was added after this accident. By consolidating four sources (queue, completed, existing files, coverage.json), there's no gap no matter when the duplicate check runs.\n\nWhen a process is force-killed with `kill -9`\n\n, the `trap`\n\ndoesn't fire and the lock directory is left behind. When that happens, the next slot's watchdog decides \"lock held; skip\" and exits immediately.\n\nI noticed two days later, because the `💓`\n\nnever arrived on Discord. When I checked the log, \"lock held; skip\" was lined up more than 100 times.\n\nThe 1800-second stale detection in `acquire_lock()`\n\n(the code quoted earlier) was added in response to this accident. If the lock has been sitting for 1800 seconds (30 minutes), it forcibly takes over. The longest timeout for article generation is also 1800 seconds, so if a healthy process is running, it will always release the lock within that time. Takeover only happens when the lock is judged stale.\n\nSince putting this fix in, \"days when the Discord `💓`\n\ndoesn't arrive\" have been zero. Even when the watchdog gets stuck, a `🚨`\n\nflies to Discord, so I can grasp the health of all five lanes just by checking notifications over my morning coffee. One reason I can maintain ¥1.2M/month in revenue while juggling development of 10 iOS apps is this \"no need to check\" design.\n\nI covered six \"actually got stuck\" stories earlier. Here I'll list, in bullet form, the finer traps I noticed while reviewing the implementation. These are all the \"looks like it's working, then you realize it's broken\" variety.\n\nLine 4 at the top of `content-watchdog.sh`\n\nsays this.\n\n```\nexport PATH=\"~/.nvm/versions/node/v24.13.0/bin:/opt/homebrew/bin:/usr/bin:/bin:/usr/sbin:/sbin\"\n```\n\nWithout this, `gtimeout`\n\n, `claude`\n\n, and `jq`\n\nall go missing. The `PATH`\n\nfor a process launched by launchd is only `/usr/bin:/bin:/usr/sbin:/sbin`\n\n. Nine out of ten cases of \"it works in the terminal but not under launchd\" are this. `article-daily-stock.sh`\n\nsolves it with a different approach.\n\n```\nNODE_BIN=$(ls -d \"$HOME\"/.nvm/versions/node/*/bin 2>/dev/null | sort -V | tail -1)\nexport PATH=\"/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin:/usr/sbin:/sbin:$PATH\"\n[ -n \"$NODE_BIN\" ] && export PATH=\"${NODE_BIN}:$PATH\"\n```\n\n`sort -V`\n\nsorts by version to dynamically discover the newest nvm binary. The claude binary is resolved right after with a three-stage fallback (`command -v`\n\n→ `~/.local/bin`\n\n→ searching under nvm with `ls -t`\n\n). If none of them find it, it logs `ABORT: claude binary not found`\n\nand does `exit 0`\n\n. It's `exit 0`\n\nrather than `exit 1`\n\nbecause having launchd record it as a \"failure\" and enter a retry loop would be a problem.\n\n`article`\n\nand `series`\n\nuse hidden files with fixed names (`.article-daily-done-20260710`\n\n), but `note`\n\nand `maker`\n\nare different.\n\n```\nnote)\n  hit=\"$(find \"$LOG_DIR/note-daily\" -maxdepth 1 -name \"$TODAY-*.done\" -print -quit 2>/dev/null || true)\"\n  [ -n \"$hit\" ]\n  ;;\n```\n\nIt's a `20260710-*.done`\n\nwildcard pattern. The reason is that note and maker can generate multiple topics in a day, so the topic slug goes into the filename. With a fixed-name done-marker, you'd lose track of which topic completed. `-print -quit`\n\nexits as soon as one match is found, to keep `find`\n\nfrom scanning everything as the file count grows.\n\nIf you write this with the same pattern for all lanes during implementation, the note lane's done-marker will be \"never found\" forever. You need to understand from the start that the check logic differs per lane.\n\nClaude's process doesn't terminate immediately when it receives `SIGTERM`\n\n. If `SIGTERM`\n\narrives mid-conversation while it's processing tokens, it takes anywhere from a few seconds to a few dozen seconds to finish processing before exiting. That's why the `run_to`\n\nhelper in `article-daily-stock.sh`\n\nattaches `--kill-after=30`\n\n.\n\n```\nrun_to() { local s=$1; shift; if [ -n \"$TIMEOUT_BIN\" ]; then \"$TIMEOUT_BIN\" --kill-after=30 \"$s\" \"$@\"; else \"$@\"; fi; }\n```\n\n`gtimeout`\n\n's `--kill-after=30`\n\nis a two-stage approach: \"send SIGTERM after the timeout, and if it's still alive 30 seconds later, send SIGKILL.\" Without it, the post-timeout process lingers like a zombie and trips the double-execution lock when the watchdog starts at the next slot.\n\nQueue processing, coverage.json updates, and the duplicate check in used_slugs() all use `jq`\n\n. `article-daily-stock.sh`\n\nhas an explicit check.\n\n```\ncommand -v jq >/dev/null 2>&1 || { log \"ABORT: jq not found\"; exit 0; }\n```\n\nBefore I added this, running the script in an environment without `jq`\n\nburied `jq: command not found`\n\ndeep in the log, the queue appeared to be read correctly but was actually treated as empty, and the automatic topic refill ran every single time.\n\nThe watchdog is launched multiple times a day. Since all of Claude's output flows into the log on every article generation, the log file becomes enormous within days if you do nothing. `article-daily-stock.sh`\n\nrotates once it exceeds 5MB (line 78).\n\n```\n[ -f \"$LOG\" ] && [ \"$(stat -f%z \"$LOG\" 2>/dev/null || echo 0)\" -gt 5242880 ] && mv \"$LOG\" \"$LOG.old\"\n```\n\nRotation isn't implemented on the `content-watchdog.log`\n\nside, so long-term operation requires manual checking. This is one of the unresolved items in my current setup.\n\nThe `SKIP_GEN`\n\nflag is set right after `article-daily-stock.sh`\n\nstarts.\n\n```\nSKIP_GEN=0\n[ \"$MODE\" = \"apply\" ] && [ -f \"$DONE_MARKER\" ] && SKIP_GEN=1\n```\n\nAfter that, `audit_repair()`\n\nalways runs. It doesn't stop even with `SKIP_GEN=1`\n\n(line 246).\n\n```\naudit_repair\nif [ \"$MODE\" = \"audit\" ] || [ \"$SKIP_GEN\" = \"1\" ]; then\n  log \"===== article-daily done(audit$([ \"$SKIP_GEN\" = \"1\" ] && echo '+gen-skipped')) =====\"\n  exit 0\nfi\n```\n\nIn other words, even when the watchdog calls the script multiple times a day, every call after the first only runs \"audit and self-repair of existing stock\" and exits. Duplicate generation and stock quality degradation are controlled independently. Without understanding this separation, you won't be able to figure out \"why does the watchdog call the script every time but generation only happens once?\"\n\nFrom the experience of actually breaking things and fixing them, here are more than 10 design decisions I wish I'd made from the start.\n\n**1. Make the proof of completion the existence of a file**\n\nLog contents, process exit codes, and printf output only prove \"the fact that the script ran.\" If you design so that only whether the file `~/.claude/logs/.article-daily-done-20260710`\n\nexists is treated as truth, then any failure pattern gets restarted at the next slot.\n\n**2. Make generation and delivery independent responsibilities**\n\nIf you judge the done-marker by git push success or failure, network outages and API rate limits look like generation failures. As the comment on line 453 of `article-daily-stock.sh`\n\nshows, set the done-marker at the point generation succeeds and make push best-effort. `touch \"$DONE_MARKER\"`\n\nis called again after push (line 468), but it's an idempotent operation, so the cost is zero.\n\n**3. Validate the contents before setting the done-marker**\n\nThe done-marker isn't set unless the four guards in `article_ok()`\n\npass.\n\n```\narticle_ok() {\n  local f=\"$1\"\n  [ -s \"$f\" ] || return 1\n  [ \"$(stat -f%z \"$f\" 2>/dev/null || echo 0)\" -ge \"$MIN_ARTICLE_BYTES\" ] || return 1\n  grep -qE '^title:' \"$f\" || return 1\n  grep -qiE 'request timed out|不明な商品|TODO: *本文|\\(生成失敗\\)' \"$f\" && return 1\n  return 0\n}\n```\n\nOnly after clearing three things — 1200 bytes, a title line in the frontmatter, and the absence of timeout wording — does it judge that \"this holds together as an article.\" It's the implementation of the principle that exit code 0 proves nothing.\n\n**4. Put a cap on auto-repair attempts**\n\n```\nMAX_ATTEMPTS=2\n```\n\nThis constant is on line 136 of `article-daily-stock.sh`\n\n. It automatically retries regeneration of a broken article up to twice, and on the third it stacks it into the `needhuman`\n\nlist and notifies. Repeating regeneration infinitely wastes tokens when the queue is corrupted or Claude can't handle a certain kind of input.\n\n**5. Check duplicates against four sources**\n\n`used_slugs()`\n\nconsolidates four sources.\n\n```\nused_slugs() {\n  { jq -r '.[].slug' \"$QUEUE\" 2>/dev/null\n    jq -r '.[].slug' \"$DONEQ\" 2>/dev/null\n    ls \"$ARTICLES\" 2>/dev/null | sed 's/\\.md$//'\n    jq -r '.[].slug' \"$COVERAGE\" 2>/dev/null\n  } | sort -u\n}\n```\n\nIf any one of queue, completed, existing files, or coverage.json is missing, duplicate topics get generated. The \"10 articles lined up on the same topic\" accident happened because only the coverage check was missing.\n\n**6. Take contention locks with mkdir**\n\n`mkdir`\n\nis atomic at the POSIX level. `flock`\n\nbehaves subtly differently on Linux and macOS, but `mkdir`\n\n's atomicity is guaranteed on both. `content-watchdog.sh`\n\n's lock directory name ends in `.lockd`\n\n(trailing d) to make it explicit that it's a directory.\n\n**7. Auto-take-over stale locks at 1800 seconds**\n\nSince each lane's timeout is 1800 seconds, a healthy process will always release the lock within that time.\n\n```\nif [ \"$age\" -ge 1800 ]; then\n  log \"lock stale age=${age}s; taking over\"\n  rm -rf \"$LOCKDIR\"\n  ...\n```\n\nThis value prevents mistaken takeovers by matching \"how many seconds until takeover\" with \"the lane script's timeout.\" The point is to keep the numbers consistent.\n\n**8. Decide whether to restart based on the done-marker, not the exit code**\n\n`run_lane()`\n\nlogs the exit code after restarting (`rc`\n\n), but doesn't use it for the decision.\n\n```\nrc=$?\nlog \"lane=$lane reinvoke_exit=$rc\"\n\nif done_lane \"$lane\"; then\n  log \"lane=$lane status=healthy-after-reinvoke\"\n  return 0\nfi\n```\n\nExit code 0 but no done-marker counts as needing a restart. Exit code 1 but a done-marker present counts as success. This accurately catches the case where a script flushes an error into the output file and exits normally (timeout wording mixed into the body).\n\n**9. Prevent midnight sleep with caffeinate exec**\n\nIf article generation runs late at night on a battery-powered Mac, the OS goes to sleep.\n\n```\nif [ -z \"${CAFFEINATED:-}\" ]; then\n  exec /usr/bin/caffeinate -i -s env CAFFEINATED=1 /bin/bash \"$0\" \"$@\"\nfi\n```\n\nIt uses `exec`\n\nto relaunch itself wrapped in `caffeinate`\n\n. The `-s`\n\nflag prevents system sleep, so it doesn't sleep even with the lid closed. Passing `CAFFEINATED=1`\n\nas an environment variable prevents a double exec.\n\n**10. Always implement a network wait**\n\nlaunchd sometimes calls processes right after the Mac boots. If you hit Claude's API before the Wifi connection is established, all lanes fail.\n\n```\nfor _ in $(seq 1 18); do\n  /usr/bin/nc -z -G 3 1.1.1.1 443 2>/dev/null && break; sleep 5\ndone\n```\n\nIt waits a maximum of 18 times × 5 seconds = 90 seconds. Capping the connect timeout at 3 seconds with `-G 3`\n\nmeans it exits within a few seconds if Wifi is connected. `audit`\n\nmode skips this loop because it doesn't use Claude (line 87).\n\n**11. Do the budget check before hitting the API**\n\n```\nBUDGET=$(~/.claude/scripts/token-budget-advisor.sh --short 2>/dev/null || echo \"n/a\")\nif echo \"$BUDGET\" | grep -qE '🔴|critical|cap-near'; then\n  log \"ABORT: budget critical — 次スロットで再試行\"; exit 0\nfi\n```\n\nEven if the budget runs out at the end of the month, the watchdog automatically restarts at a slot after the next month's reset. This is a mechanism I added after experiencing \"everything stopped at the end of the month and I restarted it by hand.\"\n\n**12. Heartbeat once a day, failure notifications every time**\n\nNotifying Discord of every success causes notification fatigue. `send_heartbeat_once()`\n\nnarrows it to the first time only via `HEARTBEAT_MARKER`\n\n. Failure notifications (`🚨`\n\n) fire every time.\n\n```\nnotify alerts \"💓 content-watchdog heartbeat: 全レーン当日生成済み\"\ntouch \"$HEARTBEAT_MARKER\"\n```\n\nBy contrast, failure notification evaluates its condition every time with `if [ -n \"$FAILED_LANES\" ]`\n\n. It's a design of \"quiet success, loud failure.\" Open Discord in the morning: a `💓`\n\nmeans all lanes are healthy, a `🚨`\n\nmeans manual intervention is needed.\n\n**13. For existing scripts like ameba, check the actual artifact rather than a done-marker**\n\nWhen bolting a watchdog onto an existing script after the fact, the cost of modifying that script to add a done-marker is sometimes high. The `ameba`\n\nlane's fallback implementation (a two-stage check of `created:`\n\nmetadata and mtime) is a realistic compromise. The ideal is to add the done-marker on the script side, but which to prioritize — \"making it work\" or \"making it perfect\" — depends on the situation.\n\nIn this article I walked through the actual code of `content-watchdog.sh`\n\nand `article-daily-stock.sh`\n\nto examine the core of a design that \"repairs itself when it gets stuck.\"\n\nThree ideas sit at the center.\n\n**Prove completion by the existence of a file.** Not a log string, not an exit code — only whether `~/.claude/logs/.article-daily-done-20260710`\n\nexists is treated as truth. This catches every instance of the \"the script ran but there's no content\" failure pattern.\n\n**Make generation and delivery independent responsibilities.** If git push success is a condition for the done-marker, a network outage looks like a generation failure. Treat the moment the article is written to `~/Desktop/Article`\n\nas success, and make delivery best-effort. End-of-month API rate limits no longer affect that day's article generation.\n\n**Run the audit every time, and generation only once a day.** `audit_repair()`\n\nexecutes even with `SKIP_GEN=1`\n\n. Even on days when new generation is skipped, quality checks on existing stock and regeneration of broken articles keep running quietly. Without this separation, you're left with a hole: \"on a day with the generated flag set, nobody notices when an existing article is broken.\"\n\nThe reason I can maintain ¥1.2M/month in revenue while juggling 10 iOS apps six months after being laid off is this \"no need to check\" design. Just confirming that a `💓`\n\narrives on Discord every morning tells me all five content lanes are healthy. Before the watchdog existed, I checked logs manually while carrying the worry that \"it should still be running today.\" That time and anxiety cost dropping to zero is what creates the room to focus on other development.\n\nI've put the full picture of the system, the breakdown of the ¥1.2M/month, and the 30-day walkthrough into a paid note.\n\n📕 [Claude Code自律環境で、実際どう稼ぐか ― 仕組み・実例・始め方・サポート](https://note.com/bokuwalily/n/n849b3a07784a)\n\n*Written by **Lily** — I ship iOS apps and automate my content stack with Claude Code.\n\nFollow along: [Portfolio](https://bokuwalily.com) · [X](https://x.com/bokuwalily) · [GitHub](https://github.com/bokuwalily)*", "url": "https://wpnews.pro/news/5-content-lanes-one-watchdog-how-i-stopped-wondering-if-my-automation-still-runs", "canonical_source": "https://dev.to/bokuwalily/5-content-lanes-one-watchdog-how-i-stopped-wondering-if-my-automation-still-runs-3fh1", "published_at": "2026-08-18 05:00:07+00:00", "updated_at": "2026-08-18 05:11:53.350463+00:00", "lang": "en", "topics": ["developer-tools"], "entities": ["Claude", "launchd", "Discord"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/5-content-lanes-one-watchdog-how-i-stopped-wondering-if-my-automation-still-runs", "markdown": "https://wpnews.pro/news/5-content-lanes-one-watchdog-how-i-stopped-wondering-if-my-automation-still-runs.md", "text": "https://wpnews.pro/news/5-content-lanes-one-watchdog-how-i-stopped-wondering-if-my-automation-still-runs.txt", "jsonld": "https://wpnews.pro/news/5-content-lanes-one-watchdog-how-i-stopped-wondering-if-my-automation-still-runs.jsonld"}}