{"slug": "dotghostboard-2-1-leviathan-smart-clipboard-tagging-contextual-actions-on-linux", "title": "DotGhostBoard 2.1 'Leviathan' — Smart Clipboard Tagging & Contextual Actions on Linux (Part 1)", "summary": "A developer built DotGhostBoard 2.1 \"Leviathan,\" a Linux clipboard manager that classifies copied content locally using pre-compiled regular expressions, structural heuristics, and a pool-based keyspace entropy estimate rather than an LLM or cloud API. The engine, in core/security/auto_tagger.py, runs synchronously on every new clipboard item before storage and UI rendering, tagging items as links, JSON, code, secrets, IPs, emails, paths, or hashes. The developer reports classification completes in under 1ms for typical payloads under 5KB, scaling to about 1.4ms worst-case at the 8KB cap.", "body_md": "You copy a minified JSON blob from a terminal log. To actually read it, you paste it into an editor, run a formatter, then copy the result back. You copy an API key from a config file and your clipboard manager stores it as plain text, alongside screenshots and grocery lists, until you manually hunt it down and secure it.\n\nMost clipboard tools treat every copied item the same: inert text in a database. For anyone who spends time in terminals, editors, and browser devtools, that uniformity creates constant small frictions that accumulate fast.\n\n**DotGhostBoard 2.1 \"Leviathan\"** takes a different approach. The clipboard understands what it's holding and surfaces the right action immediately — format this JSON, open that URL, secure this credential.\n\nThis is a two-part series on how we built it:\n\n`.vault` Backups, CSPRNG Password Generator, Secret Expiry, and Encrypted Version History.\nThe obvious shortcut for text classification is an LLM or a cloud API. For a clipboard manager, that's a non-starter on three counts.\n\nYour clipboard routinely holds API keys, internal paths, private tokens, and proprietary code snippets. Sending that to any remote service — even a local on-device model — introduces a trust boundary that shouldn't exist in a tool people actually rely on for security. Beyond that, loading inference libraries adds hundreds of MB of idle RAM, and any latency between `Ctrl+C` and the item appearing in history is enough to make the tool feel broken.\n\nOur constraint was concrete: **every classification must run locally, produce deterministic results, and complete in under 1ms for typical clipboard payloads (< 5KB), scaling to ~1.4ms worst-case at the 8KB classification cap.**\n\nThe approach: pre-compiled regular expressions, structural heuristics, and a pool-based keyspace entropy estimate — no opaque models.\n\nThe engine lives in `core/security/auto_tagger.py`. It runs synchronously on every new clipboard item before storage and UI rendering.\n\n```\nCaptured Text ──► Pre-compiled Patterns & Heuristics ──► Tags ──► Storage & UI\n                  ├── URL regex                       ──► #link\n                  ├── JSON fast-path + json.loads     ──► #json\n                  ├── Code keywords + bracket density ──► #code\n                  ├── Known prefixes + entropy gate   ──► #secret\n                  ├── IPv4 regex + IPv6 regex         ──► #ip\n                  ├── RFC-like email regex            ──► #email\n                  ├── Unix / Windows path regex       ──► #path\n                  └── Hex length check {32,64}        ──► #hash\n```\n\nAll patterns compile once at module load time. The function is pure — no I/O, no Qt, no database access — which makes it fast to call and trivial to test.\n\nHere's the real implementation, abbreviated:\n\n``` python\n# core/security/auto_tagger.py\n\nimport json, re, math, string\n\n# ── Compiled patterns (module-level — compiled once, reused forever) ──────────\n\n_RE_URL = re.compile(r\"https?://[^\\s\\\"'<>]{3,}|ftp://[^\\s\\\"'<>]{3,}\", re.IGNORECASE)\n_RE_EMAIL = re.compile(r\"\\b[A-Za-z0-9._%+\\-]+@[A-Za-z0-9.\\-]+\\.[A-Za-z]{2,}\\b\")\n\n_RE_IPV4 = re.compile(                             # validates 0-255 per octet\n    r\"\\b(?:(?:25[0-5]|2[0-4]\\d|[01]?\\d\\d?)\\.){3}\"\n    r\"(?:25[0-5]|2[0-4]\\d|[01]?\\d\\d?)\\b\"\n)\n_RE_IPV6 = re.compile(r\"\\b(?:[0-9a-fA-F]{1,4}:){2,7}[0-9a-fA-F]{1,4}\\b\")\n\n# Pragmatic range matching MD5 (32), SHA-1 (40), SHA-256 (64); strict lengths planned for v2.2\n_RE_HEX_HASH = re.compile(r\"\\b[0-9a-fA-F]{32,64}\\b\")\n\n# Known credential prefixes — the first line of secret detection\n_RE_SECRET_PATTERNS = [\n    re.compile(r\"ghp_[A-Za-z0-9]{36,}\"),                          # GitHub PAT\n    re.compile(r\"AKIA[0-9A-Z]{16}\"),                               # AWS Access Key\n    re.compile(r\"sk-[A-Za-z0-9]{32,}\"),                           # OpenAI / Stripe\n    re.compile(r\"xox[baprs]-[A-Za-z0-9\\-]{10,}\"),                 # Slack tokens\n    re.compile(r\"eyJ[A-Za-z0-9\\-_]{10,}\\.[A-Za-z0-9\\-_]{10,}\"),  # JWT\n    re.compile(r\"-----BEGIN (?:RSA |EC |OPENSSH )?PRIVATE KEY-----\"),\n]\n\n# (Note: _RE_PATH, _CODE_KEYWORDS, and _CODE_BRACKETS are omitted here for brevity)\n\n# ── Pool-based keyspace entropy estimate ──────────────────────────────────────\n\ndef _entropy_bits(text: str) -> float:\n    \"\"\"Estimate theoretical keyspace entropy: H = len * log₂(pool_size).\"\"\"\n    pool = 0\n    if any(c in string.ascii_lowercase for c in text): pool += 26\n    if any(c in string.ascii_uppercase for c in text): pool += 26\n    if any(c in string.digits           for c in text): pool += 10\n    # Any non-alphanumeric character (symbols, punctuation, or non-ASCII) adds 32\n    if any(c not in string.ascii_letters + string.digits for c in text): pool += 32\n    return math.log2(pool or 26) * len(text)\n\n# ── Main classification function ──────────────────────────────────────────────\n\ndef detect_tags(content: str) -> list[str]:\n    if not content or not isinstance(content, str):\n        return []\n    tags: list[str] = []\n    text = content.strip()\n\n    if _RE_URL.search(text):    tags.append(\"#link\")\n    if _RE_EMAIL.search(text):  tags.append(\"#email\")\n    if _RE_IPV4.search(text) or _RE_IPV6.search(text): tags.append(\"#ip\")\n    if _RE_PATH.search(text) and \"#link\" not in tags:   tags.append(\"#path\")\n    if _RE_HEX_HASH.search(text):                       tags.append(\"#hash\")\n\n    # Secret detection: known prefix first, entropy gate as fallback\n    is_secret = any(p.search(text) for p in _RE_SECRET_PATTERNS)\n    if not is_secret and len(text) >= 16 and \" \" not in text and \"\\n\" not in text:\n        if _entropy_bits(text) >= 90:\n            is_secret = True\n    if is_secret:\n        tags.append(\"#secret\")\n\n    # JSON: only attempt parse if it looks like an object or array\n    if text.lstrip().startswith((\"{\", \"[\")):\n        try:\n            json.loads(text)\n            tags.append(\"#json\")\n        except (json.JSONDecodeError, ValueError):\n            pass\n\n    # Code: multi-line with language keywords or dense bracket patterns\n    if \"\\n\" in text and len(text) > 40:\n        if _CODE_KEYWORDS.search(text) or _CODE_BRACKETS.search(text):\n            tags.append(\"#code\")\n\n    return list(dict.fromkeys(tags))  # deduplicate, preserve order\n```\n\nThe `#secret` classifier uses a two-step approach:\n\n**Layer 1 — Known prefixes.** Patterns like `ghp_`, `AKIA`, `eyJ` (JWT header) are unambiguous. Match any of them and the tag is assigned immediately, no entropy calculation needed.\n\n**Layer 2 — Entropy gate.** For everything else: single-line, no spaces, at least 16 characters, and a pool-based keyspace entropy ≥ 90 bits. The 90-bit threshold was tuned empirically to catch random high-entropy tokens, API keys, and strong bearer credentials while skipping short identifiers, dictionary words, and common camelCase variable names.\n\nBoth conditions guard against a common pitfall: pure Shannon entropy on raw text produces too many false positives on long sentences or dense code snippets. The entropy function here is `H = len × log₂(pool_size)` — a measure of the *theoretical keyspace*, not character frequency distribution.\n\n**Note on non-ASCII characters:** Because `any(c not in string.ascii_letters + string.digits ...)` catches non-ASCII characters (e.g. Arabic script, accented letters, emoji), unicode text will trigger the `+32` symbols branch. This heuristic is tuned primarily for ASCII credentials; multilingual script entropy is handled gracefully without throwing exceptions.\n\nHonest engineering requires talking about edge cases:\n\n`#secret`:`/home/kareem/StudioProjects/DotGhostBoard/main.py` currently receives both `#path` and `/`. We've logged this as a known limitation; v2.2 will add `pool = 36`), yielding `32 × log₂(36) ≈ 165 bits`, which clears the ≥ 90-bit gate. Consequently, MD5 and SHA digests receive both `#hash` and Tags are displayed as colored pill chips on each card. More importantly, they drive **Contextual Smart Actions** — buttons that appear dynamically in the card toolbar based on what was detected.\n\n| Tag | Action | What it does | \n|---|---|---|\n| `#link` | 🔗 Open Link | `QDesktopServices.openUrl()` — no copy-paste to browser | \n| `#json` | `{ }` Format JSON | Parses, re-serializes with 2-space indent, writes back to clipboard | \n| `#email` | ✉ Compose | Constructs a `mailto:` URI, hands off to default mail client | \n| `#ip` | 📡 Copy IP | Extracts just the IP from surrounding log noise | \n| `#secret` | 🛡 → Vault | Opens Vault drawer, prefills secret dialog, purges plaintext from history | \n\nHere's the JSON formatter — the action developers use most:\n\n``` php\n# ui/widgets/item_card.py (simplified)\nimport json\n\ndef _on_format_json(self) -> None:\n    try:\n        parsed = json.loads(self._raw_text)\n        formatted = json.dumps(parsed, indent=2, ensure_ascii=False)\n        QApplication.clipboard().setText(formatted)\n        self._flash_badge(\"✓ Formatted\")\n    except json.JSONDecodeError:\n        self._flash_badge(\"✗ Invalid JSON\")\n```\n\nOne button. Minified JSON in, readable JSON back in your clipboard. No editor, no extra window.\n\n*Figure 1: Auto-detected tags (`#link`, `#ip`, `#code`) and dynamic contextual smart actions (`🔗 Open Link`, `📡 Copy IP`) on clipboard cards.*\n\nThe classifier only stays fast because it runs inside a well-isolated pipeline:\n\n*Figure 2: The 4-phase sequential classification and dispatch pipeline.*\n\n```\n[System Clipboard] ──► ClipboardWatcher (background polling / events)\n                             │\n                             ▼\n                    ClipboardPipeline (pure policy layer)\n                      ├── App filter      (skip blacklisted sources)\n                      ├── Self-paste guard (skip our own copy events)\n                      ├── AutoTagger      (sync, pure, no I/O)\n                      └── Duplicate & pin policy\n                             │\n                             ▼\n                    HistoryService ──► SQLite (PRAGMA secure_delete)\n                             │\n                             ▼\n                    HistoryController ──► CardsView (lazy render)\n```\n\n`AutoTagger` has zero dependencies on Qt, SQLite, or any network code. That isolation means:\n\nWe wanted hard performance guarantees rather than assumptions. Here are `timeit` measurements on typical clipboard payloads running on an **Intel Core i5-8350U @ 1.70GHz (Python 3.14.7, Linux x86_64)**, 1,000 iterations each:\n\n*Figure 3: Microsecond-level classification benchmark measurements across typical clipboard payloads.*\n\n```\nPayload                      min µs    median µs    max µs\n──────────────────────────────────────────────────────────\nURL                             8.2          8.8      10.7\nGitHub PAT                      6.5          6.8      10.1\nIPv4 address                    7.1          7.5      12.7\nMD5 hash                        7.1          7.2       7.6\nJSON (small, ~80 chars)        14.1         14.4      16.1\nCode snippet                   20.3         20.4      21.4\nPlain English text             14.8         15.2      15.3\nJSON (1KB payload)             80.3         81.8      82.6\nWorst-case dense (7.8KB)     1380.0       1410.6    1460.0\n```\n\nThe \"< 1ms\" guarantee holds for typical clipboard payloads (< 5KB), with small items (URLs, hashes, JSON) processing in 8–82 µs. At the 8KB cap, worst-case dense tokens (7.8KB of unspaced text forcing character-by-character entropy evaluation) measure ~1.4ms. Payloads exceeding 8KB automatically bypass the entropy gate to ensure the UI loop never hitches on huge paste events.\n\nThe productivity layer is half the story. The other half is what happens to sensitive data after it's detected.\n\nIn **[Part 2]**, we cover the security architecture:\n\n`⛔ EXPIRED` and `⚠️ 3d left` badges.\n100% free and open-source — **Apache-2.0 License.**\n\n```\ngit clone https://github.com/kareem2099/DotGhostBoard.git\ncd DotGhostBoard\n\npython3 -m venv venv && source venv/bin/activate\npip install -r requirements.txt\n\npython3 main.py\n```\n\n| Shortcut | Action | \n|---|---|\n| `Ctrl+Alt+V` | Summon / hide Dashboard (migrates to active workspace) | \n| `Ctrl+Alt+Space` | Floating Spotlight Quick Search | \n| `Ctrl+Shift+V` | Open / close The Vault encrypted drawer | \n\nQuestions about the classification heuristics, the entropy threshold choice, or the action pipeline? Drop them in the comments. 👻", "url": "https://wpnews.pro/news/dotghostboard-2-1-leviathan-smart-clipboard-tagging-contextual-actions-on-linux", "canonical_source": "https://dev.to/freerave/dotghostboard-21-leviathan-smart-clipboard-tagging-contextual-actions-on-linux-part-1-5gb8", "published_at": "2026-10-02 20:59:46+00:00", "updated_at": "2026-10-02 21:07:20.009847+00:00", "lang": "en", "topics": ["developer-tools", "ai-tools"], "entities": ["DotGhostBoard", "DotGhostBoard 2.1 \"Leviathan\"", "GitHub", "AWS", "OpenAI", "Stripe", "Slack"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/dotghostboard-2-1-leviathan-smart-clipboard-tagging-contextual-actions-on-linux", "markdown": "https://wpnews.pro/news/dotghostboard-2-1-leviathan-smart-clipboard-tagging-contextual-actions-on-linux.md", "text": "https://wpnews.pro/news/dotghostboard-2-1-leviathan-smart-clipboard-tagging-contextual-actions-on-linux.txt", "jsonld": "https://wpnews.pro/news/dotghostboard-2-1-leviathan-smart-clipboard-tagging-contextual-actions-on-linux.jsonld"}}