DotGhostBoard 2.1 'Leviathan' — Smart Clipboard Tagging & Contextual Actions on Linux (Part 1) A developer built DotGhostBoard 2.1 "Leviathan," a Linux clipboard manager that classifies copied content locally using pre-compiled regular expressions, structural heuristics, and a pool-based keyspace entropy estimate rather than an LLM or cloud API. The engine, in core/security/auto_tagger.py, runs synchronously on every new clipboard item before storage and UI rendering, tagging items as links, JSON, code, secrets, IPs, emails, paths, or hashes. The developer reports classification completes in under 1ms for typical payloads under 5KB, scaling to about 1.4ms worst-case at the 8KB cap. You copy a minified JSON blob from a terminal log. To actually read it, you paste it into an editor, run a formatter, then copy the result back. You copy an API key from a config file and your clipboard manager stores it as plain text, alongside screenshots and grocery lists, until you manually hunt it down and secure it. Most clipboard tools treat every copied item the same: inert text in a database. For anyone who spends time in terminals, editors, and browser devtools, that uniformity creates constant small frictions that accumulate fast. DotGhostBoard 2.1 "Leviathan" takes a different approach. The clipboard understands what it's holding and surfaces the right action immediately — format this JSON, open that URL, secure this credential. This is a two-part series on how we built it: .vault Backups, CSPRNG Password Generator, Secret Expiry, and Encrypted Version History. The obvious shortcut for text classification is an LLM or a cloud API. For a clipboard manager, that's a non-starter on three counts. Your clipboard routinely holds API keys, internal paths, private tokens, and proprietary code snippets. Sending that to any remote service — even a local on-device model — introduces a trust boundary that shouldn't exist in a tool people actually rely on for security. Beyond that, loading inference libraries adds hundreds of MB of idle RAM, and any latency between Ctrl+C and the item appearing in history is enough to make the tool feel broken. Our constraint was concrete: every classification must run locally, produce deterministic results, and complete in under 1ms for typical clipboard payloads < 5KB , scaling to ~1.4ms worst-case at the 8KB classification cap. The approach: pre-compiled regular expressions, structural heuristics, and a pool-based keyspace entropy estimate — no opaque models. The engine lives in core/security/auto tagger.py . It runs synchronously on every new clipboard item before storage and UI rendering. Captured Text ──► Pre-compiled Patterns & Heuristics ──► Tags ──► Storage & UI ├── URL regex ──► link ├── JSON fast-path + json.loads ──► json ├── Code keywords + bracket density ──► code ├── Known prefixes + entropy gate ──► secret ├── IPv4 regex + IPv6 regex ──► ip ├── RFC-like email regex ──► email ├── Unix / Windows path regex ──► path └── Hex length check {32,64} ──► hash All patterns compile once at module load time. The function is pure — no I/O, no Qt, no database access — which makes it fast to call and trivial to test. Here's the real implementation, abbreviated: python core/security/auto tagger.py import json, re, math, string ── Compiled patterns module-level — compiled once, reused forever ────────── RE URL = re.compile r"https?:// ^\s\"'< {3,}|ftp:// ^\s\"'< {3,}", re.IGNORECASE RE EMAIL = re.compile r"\b A-Za-z0-9. %+\- +@ A-Za-z0-9.\- +\. A-Za-z {2,}\b" RE IPV4 = re.compile validates 0-255 per octet r"\b ?: ?:25 0-5 |2 0-4 \d| 01 ?\d\d? \. {3}" r" ?:25 0-5 |2 0-4 \d| 01 ?\d\d? \b" RE IPV6 = re.compile r"\b ?: 0-9a-fA-F {1,4}: {2,7} 0-9a-fA-F {1,4}\b" Pragmatic range matching MD5 32 , SHA-1 40 , SHA-256 64 ; strict lengths planned for v2.2 RE HEX HASH = re.compile r"\b 0-9a-fA-F {32,64}\b" Known credential prefixes — the first line of secret detection RE SECRET PATTERNS = re.compile r"ghp A-Za-z0-9 {36,}" , GitHub PAT re.compile r"AKIA 0-9A-Z {16}" , AWS Access Key re.compile r"sk- A-Za-z0-9 {32,}" , OpenAI / Stripe re.compile r"xox baprs - A-Za-z0-9\- {10,}" , Slack tokens re.compile r"eyJ A-Za-z0-9\- {10,}\. A-Za-z0-9\- {10,}" , JWT re.compile r"-----BEGIN ?:RSA |EC |OPENSSH ?PRIVATE KEY-----" , Note: RE PATH, CODE KEYWORDS, and CODE BRACKETS are omitted here for brevity ── Pool-based keyspace entropy estimate ────────────────────────────────────── def entropy bits text: str - float: """Estimate theoretical keyspace entropy: H = len log₂ pool size .""" pool = 0 if any c in string.ascii lowercase for c in text : pool += 26 if any c in string.ascii uppercase for c in text : pool += 26 if any c in string.digits for c in text : pool += 10 Any non-alphanumeric character symbols, punctuation, or non-ASCII adds 32 if any c not in string.ascii letters + string.digits for c in text : pool += 32 return math.log2 pool or 26 len text ── Main classification function ────────────────────────────────────────────── def detect tags content: str - list str : if not content or not isinstance content, str : return tags: list str = text = content.strip if RE URL.search text : tags.append " link" if RE EMAIL.search text : tags.append " email" if RE IPV4.search text or RE IPV6.search text : tags.append " ip" if RE PATH.search text and " link" not in tags: tags.append " path" if RE HEX HASH.search text : tags.append " hash" Secret detection: known prefix first, entropy gate as fallback is secret = any p.search text for p in RE SECRET PATTERNS if not is secret and len text = 16 and " " not in text and "\n" not in text: if entropy bits text = 90: is secret = True if is secret: tags.append " secret" JSON: only attempt parse if it looks like an object or array if text.lstrip .startswith "{", " " : try: json.loads text tags.append " json" except json.JSONDecodeError, ValueError : pass Code: multi-line with language keywords or dense bracket patterns if "\n" in text and len text 40: if CODE KEYWORDS.search text or CODE BRACKETS.search text : tags.append " code" return list dict.fromkeys tags deduplicate, preserve order The secret classifier uses a two-step approach: Layer 1 — Known prefixes. Patterns like ghp , AKIA , eyJ JWT header are unambiguous. Match any of them and the tag is assigned immediately, no entropy calculation needed. Layer 2 — Entropy gate. For everything else: single-line, no spaces, at least 16 characters, and a pool-based keyspace entropy ≥ 90 bits. The 90-bit threshold was tuned empirically to catch random high-entropy tokens, API keys, and strong bearer credentials while skipping short identifiers, dictionary words, and common camelCase variable names. Both conditions guard against a common pitfall: pure Shannon entropy on raw text produces too many false positives on long sentences or dense code snippets. The entropy function here is H = len × log₂ pool size — a measure of the theoretical keyspace , not character frequency distribution. Note on non-ASCII characters: Because any c not in string.ascii letters + string.digits ... catches non-ASCII characters e.g. Arabic script, accented letters, emoji , unicode text will trigger the +32 symbols branch. This heuristic is tuned primarily for ASCII credentials; multilingual script entropy is handled gracefully without throwing exceptions. Honest engineering requires talking about edge cases: secret : /home/kareem/StudioProjects/DotGhostBoard/main.py currently receives both path and / . We've logged this as a known limitation; v2.2 will add pool = 36 , yielding 32 × log₂ 36 ≈ 165 bits , which clears the ≥ 90-bit gate. Consequently, MD5 and SHA digests receive both hash and Tags are displayed as colored pill chips on each card. More importantly, they drive Contextual Smart Actions — buttons that appear dynamically in the card toolbar based on what was detected. | Tag | Action | What it does | |---|---|---| | link | 🔗 Open Link | QDesktopServices.openUrl — no copy-paste to browser | | json | { } Format JSON | Parses, re-serializes with 2-space indent, writes back to clipboard | | email | ✉ Compose | Constructs a mailto: URI, hands off to default mail client | | ip | 📡 Copy IP | Extracts just the IP from surrounding log noise | | secret | 🛡 → Vault | Opens Vault drawer, prefills secret dialog, purges plaintext from history | Here's the JSON formatter — the action developers use most: php ui/widgets/item card.py simplified import json def on format json self - None: try: parsed = json.loads self. raw text formatted = json.dumps parsed, indent=2, ensure ascii=False QApplication.clipboard .setText formatted self. flash badge "✓ Formatted" except json.JSONDecodeError: self. flash badge "✗ Invalid JSON" One button. Minified JSON in, readable JSON back in your clipboard. No editor, no extra window. Figure 1: Auto-detected tags link , ip , code and dynamic contextual smart actions 🔗 Open Link , 📡 Copy IP on clipboard cards. The classifier only stays fast because it runs inside a well-isolated pipeline: Figure 2: The 4-phase sequential classification and dispatch pipeline. System Clipboard ──► ClipboardWatcher background polling / events │ ▼ ClipboardPipeline pure policy layer ├── App filter skip blacklisted sources ├── Self-paste guard skip our own copy events ├── AutoTagger sync, pure, no I/O └── Duplicate & pin policy │ ▼ HistoryService ──► SQLite PRAGMA secure delete │ ▼ HistoryController ──► CardsView lazy render AutoTagger has zero dependencies on Qt, SQLite, or any network code. That isolation means: We wanted hard performance guarantees rather than assumptions. Here are timeit measurements on typical clipboard payloads running on an Intel Core i5-8350U @ 1.70GHz Python 3.14.7, Linux x86 64 , 1,000 iterations each: Figure 3: Microsecond-level classification benchmark measurements across typical clipboard payloads. Payload min µs median µs max µs ────────────────────────────────────────────────────────── URL 8.2 8.8 10.7 GitHub PAT 6.5 6.8 10.1 IPv4 address 7.1 7.5 12.7 MD5 hash 7.1 7.2 7.6 JSON small, ~80 chars 14.1 14.4 16.1 Code snippet 20.3 20.4 21.4 Plain English text 14.8 15.2 15.3 JSON 1KB payload 80.3 81.8 82.6 Worst-case dense 7.8KB 1380.0 1410.6 1460.0 The "< 1ms" guarantee holds for typical clipboard payloads < 5KB , with small items URLs, hashes, JSON processing in 8–82 µs. At the 8KB cap, worst-case dense tokens 7.8KB of unspaced text forcing character-by-character entropy evaluation measure ~1.4ms. Payloads exceeding 8KB automatically bypass the entropy gate to ensure the UI loop never hitches on huge paste events. The productivity layer is half the story. The other half is what happens to sensitive data after it's detected. In Part 2 , we cover the security architecture: ⛔ EXPIRED and ⚠️ 3d left badges. 100% free and open-source — Apache-2.0 License. git clone https://github.com/kareem2099/DotGhostBoard.git cd DotGhostBoard python3 -m venv venv && source venv/bin/activate pip install -r requirements.txt python3 main.py | Shortcut | Action | |---|---| | Ctrl+Alt+V | Summon / hide Dashboard migrates to active workspace | | Ctrl+Alt+Space | Floating Spotlight Quick Search | | Ctrl+Shift+V | Open / close The Vault encrypted drawer | Questions about the classification heuristics, the entropy threshold choice, or the action pipeline? Drop them in the comments. 👻