# LLM-Shield-Proxy - Zero-Egress PII Proxy for LLM SoC 2 (24MB RAM)

> Source: <https://github.com/ninadphalak/LLM-Shield-Proxy>
> Published: 2026-08-04 12:11:37+00:00

SOC 2 and HIPAA compliance for LLM streams without breaking real-time latency.

**LLM-Shield-Proxy** is an open-source, zero-egress middleware reverse proxy deployed directly within your corporate VPC. It intercepts OpenAI-compatible LLM API requests, redacts Personally Identifiable Information (PII) before it leaves your infrastructure, and deterministically re-hydrates real-time Server-Sent Events (SSE) chat responses with ultra-low stream latency.

Designed to unblock enterprise privacy compliance (**SOC 2 / HIPAA**).

Author & Core Maintainer: **Ninad Phalak** (`ninadphalak@gmail.com`

)

```
pip install llm-shield-proxy "uvicorn[standard]"
docker run -d -p 8000:8000 \
  -e OPENAI_API_KEY="sk-your-openai-api-key" \
  --name llm-shield-proxy \
  ghcr.io/ninadphalak/llm-shield-proxy:latest
version: "3.8"

services:
  llm-shield-proxy:
    image: ghcr.io/ninadphalak/llm-shield-proxy:latest
    ports:
      - "8000:8000"
    environment:
      - OPENAI_API_KEY=sk-your-openai-key-here
      - REDIS_URL=redis://redis:6379/0
    depends_on:
      - redis

  redis:
    image: redis:7-alpine
    ports:
      - "6379:6379"
```

Point your existing OpenAI SDK `base_url`

to your local LLM-Shield-Proxy instance:

``` python
from openai import OpenAI

client = OpenAI(
    api_key="your-openai-api-key",
    base_url="http://localhost:8000/v1"  # Point to LLM-Shield-Proxy
)

response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[
        {"role": "user", "content": "Contact Sarah Connor at sarah@example.com or 555-0199."}
    ],
    stream=True
)

for chunk in response:
    print(chunk.choices[0].delta.content or "", end="")
```

| Existing Legacy Proxies | LLM-Shield-Proxy |
|---|---|
Destroys Real-Time SSE Streaming: Buffers entire responses before scanning, causing multi-second UI latency stalls. |
Ultra-Low Latency Streaming: Redacts and re-hydrates delta-by-delta as SSE packets stream. |
Heavy Memory Footprint: Requires 1GB–2GB RAM for heavy spaCy or PyTorch NLP libraries. |
Ultra-Lightweight <24MB RAM: Runs on a microsecond compiled regex + quantized ONNX NER engine. |
Data Liability: Stores user PII in long-term databases. |
Zero Long-Term Storage: Self-destructing TTL session vault built for zero data liability. |
Complex Cloud Egress: Routes data to 3rd-party SaaS inspection APIs. |
100% Zero-Egress VPC: All scanning happens locally inside your secure corporate boundary. |

LLM-Shield-Proxy delivers enterprise security through two core architectural breakthroughs:

When streaming LLM responses, Server-Sent Events (SSE) send text in arbitrary token chunks. An SSE delta chunk might split a redacted placeholder tag directly across two network packets:

**Chunk N:**`Hello [PER`

**Chunk N+1:**`SON_1]! How can I help you today?`

If unbuffered, `[PER`

leaks to the user's screen as raw un-hydrated text.

**The Engineering Solution:** An asynchronous `SSERehydrationBuffer`

tracks bracket boundaries (`[`

and `]`

). When an open bracket is detected near the tail of an incoming delta without a matching closing bracket, the buffer holds back the tail bytes until the completing chunk arrives. Once complete, the deterministic token is re-hydrated to its original value with zero UI jitter or streaming stalls.

To achieve sub-millisecond execution without blowing up infrastructure costs:

**Tier 1 (Sub-millisecond Compiled Regex):** Scans structured secrets (SSNs, Credit Cards, Emails, Phone Numbers, IPv4/IPv6, API Keys) in**<0.03ms**.** Tier 2 (Quantized Local ONNX NER):**Uses a tiny, quantized ONNX Named Entity Recognition (NER) model to catch unstructured person names in**~5–12ms**.

By avoiding heavy NLP libraries like spaCy or HuggingFace transformers, LLM-Shield-Proxy runs inside a **24MB RAM process footprint** — making it fast, deterministic, and ideal for microservice sidecars. This means you can run dozens of proxy containers side-by-side on cheap micro-instances (like AWS `t4g.nano`

or Docker Swarm/Kubernetes pods) for virtually zero RAM cost.

**Zero-Egress Security:** 100% of PII scanning and re-hydration happens locally within your VPC. No prompt data or telemetry ever leaves your server.**Stateless Privacy (Self-Destructing Redis TTL):** Real PII is mapped to session-bound tokens (e.g.`Sarah`

->`[PERSON_1]`

) stored in an in-memory vault backed by strict Time-To-Live (TTL) expiration rules. When configured with Redis (`REDIS_URL`

), vaults are shared across multi-replica clusters without building a permanent database of user PII.

Emits structured JSON audit events (`app/audit.py`

) directly to `stdout`

compatible with Datadog, Splunk, Elastic, Vanta, and Drata to prove compliance for **SOC 2 Type II** and **HIPAA** audits:

```
{
  "timestamp": "2026-08-04T01:48:00Z",
  "event": "pii_redaction",
  "session_id": "sess_8f179f3",
  "path": "/v1/chat/completions",
  "redactions_summary": {
    "SSN": 1,
    "EMAIL": 2,
    "PERSON": 1
  },
  "compliance_status": "zero_egress_passed"
}
flowchart TD
    classDef client fill:#e0f2fe,stroke:#0284c7,stroke-width:2px,color:#0369a1,font-weight:bold;
    classDef proxyEngine fill:#f8fafc,stroke:#475569,stroke-width:2px,color:#0f172a,font-weight:bold;
    classDef piiSecurity fill:#fef2f2,stroke:#ef4444,stroke-width:2px,color:#991b1b,font-weight:bold;
    classDef vault fill:#fffbebe,stroke:#f59e0b,stroke-width:2px,color:#92400e,font-weight:bold;
    classDef upstream fill:#f3e8ff,stroke:#9333ea,stroke-width:2px,color:#6b21a8,font-weight:bold;

    UserApp["👤 User Application\n(OpenAI / LangChain SDK)"]:::client

    subgraph SecurityMoat ["🛡️ Zero-Egress Local Environment (Apache 2.0 Licensed)"]
        direction TD
        FastAPIProxy["⚡ FastAPI Catch-All Proxy\n(/{path:path})"]:::proxyEngine

        subgraph CascadeEngine ["🔒 Two-Tier PII Cascade Engine"]
            Tier1["Tier 1: Compiled Regex"]:::piiSecurity
            Tier2["Tier 2: Quantized ONNX NER"]:::piiSecurity
            Tier1 --> Tier2
        end

        VaultStore[("🔑 Session Vault Store\n(Deterministic Tokens)")]:::vault
        LookaheadBuffer["⏱️ Sliding-Window Lookahead Buffer\n(Prevent SSE Tag Leaks)"]:::proxyEngine
        Rehydrator["🔄 Stream Re-hydrator\n(Token -> Original Value)"]:::proxyEngine
    end

    UpstreamLLM["☁️ Upstream LLM Provider\n(OpenAI / Anthropic / vLLM)"]:::upstream

    %% Inbound Flow (Prompt Sanitization)
    UserApp -- "1. Inbound Raw Prompt Payload" --> FastAPIProxy
    FastAPIProxy -- "2. Scan Payload" --> Tier1
    Tier2 -- "3. Store Vault Keys" --> VaultStore
    Tier2 -- "4. Redacted JSON Payload" --> UpstreamLLM

    %% Outbound Flow (Streaming De-redaction)
    UpstreamLLM -. "5. Raw SSE Stream Deltas" .-> LookaheadBuffer
    LookaheadBuffer -- "6. Tag-Safe Assembly" --> Rehydrator
    Rehydrator <--> VaultStore
    Rehydrator -. "7. Sanitized Real-Time Stream" .-> UserApp

    style SecurityMoat fill:#f8fafc,stroke:#0284c7,stroke-width:2px,stroke-dasharray: 5 5,color:#0f172a
    style CascadeEngine fill:#ffffff,stroke:#cbd5e1,stroke-width:1px
```

**Intercept:** Your application sends a standard OpenAI / LangChain payload to`localhost:8000`

.**Cascade Redaction:** The proxy intercepts the JSON and routes text through a high-speed compiled Regex engine (SSNs, emails, credit cards), falling back to a local ONNX model for unstructured names.**Vault Storage:** The original PII is mapped to a deterministic tag (e.g.,`[PERSON_1]`

) and stored locally in a TTL-backed session vault.**Clean Egress:** A 100% sanitized payload is forwarded to OpenAI. OpenAI never sees your raw sensitive data.

**SSE Stream Intercept:** OpenAI streams the response back chunk-by-chunk via Server-Sent Events (SSE).**Lookahead Buffer:** Because tags can be split across SSE chunks (e.g.,`[PER`

in chunk N and`SON_1]`

in chunk N+1), the proxy's sliding-window buffer holds back unclosed brackets to prevent tag leaks.**Re-hydration:** Once a tag is fully assembled, the proxy swaps the real data back from the local vault and streams the final, un-redacted text to the user's application in real-time.

LLM-Shield-Proxy is engineered for sub-millisecond overhead and ultra-lightweight resource usage. Measured over 1,000 production streaming iterations:

| Metric | Average Latency | Median Latency | Footprint / Notes |
|---|---|---|---|
Tier 1 Regex Overhead |
`0.0294 ms` |
`0.0291 ms` (`29.10 µs` ) |
Microsecond pattern scan |
Tier 2 NER Overhead |
`0.0033 ms` |
`0.0032 ms` (`3.20 µs` ) |
Quantized local NER scan |
Total SSE Stream Overhead |
`0.0010 ms` |
`0.0010 ms` (`0.97 µs` ) |
Added latency per SSE delta chunk |
Process RAM Footprint |
- | - | `24.55 MB` Resident Set Size |

To run the automated benchmark suite locally:

```
py tests/benchmark.py
```

Transparency is critical for security tooling. Please be aware of the following current limitations:

**Text Only:** The proxy does not currently scan or redact text embedded inside base64 image payloads (e.g., OpenAI Vision models).**Supported Languages:** The Tier-2 ONNX NER model is currently optimized for English-language entities.**Non-Standard Streaming:** Designed for standard Server-Sent Events (SSE). Custom or proprietary streaming protocols may bypass the sliding-window buffer.

Run the full automated test suite:

```
py -m pytest tests/
```

Designed for zero-friction adoption by DevOps, Site Reliability Engineers (SREs), and Network Administrators:

Built-in liveness and readiness endpoints return `HTTP 200 OK`

for Kubernetes, Docker Swarm, or AWS ECS health monitors:

```
curl http://localhost:8000/health
# Output: {"status":"ok","service":"llm-shield-proxy","version":"1.0.4"}

curl http://localhost:8000/livez
# Output: {"status":"ok","service":"llm-shield-proxy","version":"1.0.4"}
```

100% compliant with 12-factor app standards. All upstream target routing and API keys are injected via environment variables or a `.env`

file without code modifications:

`UPSTREAM_BASE_URL`

: Base target URL (e.g.`https://api.openai.com`

or internal`vLLM`

server).`OPENAI_API_KEY`

: Upstream API key passed to target providers.`REDIS_URL`

: Optional Redis connection string for distributed multi-instance session caching.

LLM-Shield-Proxy runs completely stateless by default. For high-volume enterprise deployments, instances scale horizontally behind edge proxies (NGINX, Traefik, AWS ALB):

```
docker-compose up -d --scale proxy=5
```

When configured with `REDIS_URL`

, session vaults are shared across all proxy replicas, ensuring seamless session isolation across multi-instance clusters.

Every published release includes automated SHA-256 checksums (`checksums.txt`

) and GPG detached signatures (`checksums.txt.asc`

) signed by maintainer **Ninad Phalak**. You can verify checksums and cryptographic authenticity before deployment using:

```
# 1. Verify SHA-256 Checksums (Linux / macOS):
sha256sum -c checksums.txt

# On Windows (PowerShell):
Get-FileHash llm-shield-proxy-source-v1.0.4.zip -Algorithm SHA256

# 2. Verify Cryptographic GPG Signature:
gpg --verify checksums.txt.asc checksums.txt
```

Currently, LLM-Shield-Proxy's Tier 1 Regex engine is optimized for North American PII (US SSNs, Phone Formats). To support global GDPR compliance, I am actively looking for contributors to help expand regex payloads and Tier 2 ONNX models for:

**European Formats:** UK NIN, EU Phone Numbers, IBANs.**APAC Data Structures:** India Aadhaar, APAC localized identifiers.**Multilingual NER ONNX Models:** Multilingual entity recognition models.

If you want to contribute to enterprise AI security, check out [CONTRIBUTING.md](/ninadphalak/LLM-Shield-Proxy/blob/main/CONTRIBUTING.md) and claim a locale!

I am committed to maintaining LLM-Shield-Proxy as the fastest ultra-low latency redaction engine for LLMs. Here are the core architectural optimizations planned for upcoming releases — contributions and PRs are warmly welcomed:

-
**ONNX Thread Tuning (Preventing CPU Contention)*** Problem:*By default, ONNX Runtime attempts to use every available CPU core. In FastAPI, this competes with the event loop handling thousands of concurrent connections.*The Fix:*Restrict ONNX by setting`sess_options.intra_op_num_threads = 1`

. This forces ONNX execution onto a single thread, keeping CPU cores free for FastAPI's event loop to stream packets instantly.

-
**Persistent Connection Pooling (The TLS Trick)*** Problem:*Opening a new TLS/SSL connection to OpenAI per request adds 50–100ms latency.*The Fix:*Maintain a persistent`httpx.AsyncClient`

HTTP/2 connection pool on server startup. The proxy opens pre-warmed secure tunnels, routing requests instantly with zero TLS setup overhead.

-
**Swap to**`orjson`

for Chunk Parsing*Problem:*In an SSE stream, standard Python`json.loads`

parses hundreds of delta chunks per second.*The Fix:*Swap built-in`json`

for`orjson`

(written in Rust). It parses streaming LLM chunks up to 10x faster, dropping proxy overhead to near zero.

-
**Cythonize the Sliding-Window Buffer*** Problem:*The sliding-window buffer performs frequent string slicing and bracket matching.*The Fix:*Use Cython or`mypyc`

to compile`streaming.py`

directly into a C-extension binary module. Retains Python readability while executing string operations at native C speed.

I am actively working with enterprise security teams to map out advanced compliance features. If your startup or organization is using LLM-Shield-Proxy to unblock LLM streaming or pass SOC 2/HIPAA audits, I would love to hear from you.

Email the core maintainer at [ninadphalak@gmail.com](mailto:ninadphalak@gmail.com) to share your feedback, request a feature, or feature your team as a case study.
