SOC 2 and HIPAA compliance for LLM streams without breaking real-time latency.
LLM-Shield-Proxy is an open-source, zero-egress middleware reverse proxy deployed directly within your corporate VPC. It intercepts OpenAI-compatible LLM API requests, redacts Personally Identifiable Information (PII) before it leaves your infrastructure, and deterministically re-hydrates real-time Server-Sent Events (SSE) chat responses with ultra-low stream latency.
Designed to unblock enterprise privacy compliance (SOC 2 / HIPAA).
Author & Core Maintainer: Ninad Phalak (ninadphalak@gmail.com
)
pip install llm-shield-proxy "uvicorn[standard]"
docker run -d -p 8000:8000 \
-e OPENAI_API_KEY="sk-your-openai-api-key" \
--name llm-shield-proxy \
ghcr.io/ninadphalak/llm-shield-proxy:latest
version: "3.8"
services:
llm-shield-proxy:
image: ghcr.io/ninadphalak/llm-shield-proxy:latest
ports:
- "8000:8000"
environment:
- OPENAI_API_KEY=sk-your-openai-key-here
- REDIS_URL=redis://redis:6379/0
depends_on:
- redis
redis:
image: redis:7-alpine
ports:
- "6379:6379"
Point your existing OpenAI SDK base_url
to your local LLM-Shield-Proxy instance:
from openai import OpenAI
client = OpenAI(
api_key="your-openai-api-key",
base_url="http://localhost:8000/v1" # Point to LLM-Shield-Proxy
)
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{"role": "user", "content": "Contact Sarah Connor at sarah@example.com or 555-0199."}
],
stream=True
)
for chunk in response:
print(chunk.choices[0].delta.content or "", end="")
| Existing Legacy Proxies | LLM-Shield-Proxy |
|---|---|
| Destroys Real-Time SSE Streaming: Buffers entire responses before scanning, causing multi-second UI latency stalls. | |
| Ultra-Low Latency Streaming: Redacts and re-hydrates delta-by-delta as SSE packets stream. | |
| Heavy Memory Footprint: Requires 1GB–2GB RAM for heavy spaCy or PyTorch NLP libraries. | |
| Ultra-Lightweight <24MB RAM: Runs on a microsecond compiled regex + quantized ONNX NER engine. | |
| Data Liability: Stores user PII in long-term databases. | |
| Zero Long-Term Storage: Self-destructing TTL session vault built for zero data liability. | |
| Complex Cloud Egress: Routes data to 3rd-party SaaS inspection APIs. | |
| 100% Zero-Egress VPC: All scanning happens locally inside your secure corporate boundary. |
LLM-Shield-Proxy delivers enterprise security through two core architectural breakthroughs:
When streaming LLM responses, Server-Sent Events (SSE) send text in arbitrary token chunks. An SSE delta chunk might split a redacted placeholder tag directly across two network packets:
Chunk N:Hello [PER
Chunk N+1:SON_1]! How can I help you today?
If unbuffered, [PER
leaks to the user's screen as raw un-hydrated text.
The Engineering Solution: An asynchronous SSERehydrationBuffer
tracks bracket boundaries ([
and ]
). When an open bracket is detected near the tail of an incoming delta without a matching closing bracket, the buffer holds back the tail bytes until the completing chunk arrives. Once complete, the deterministic token is re-hydrated to its original value with zero UI jitter or streaming stalls.
To achieve sub-millisecond execution without blowing up infrastructure costs:
Tier 1 (Sub-millisecond Compiled Regex): Scans structured secrets (SSNs, Credit Cards, Emails, Phone Numbers, IPv4/IPv6, API Keys) in**<0.03ms**.** Tier 2 (Quantized Local ONNX NER):Uses a tiny, quantized ONNX Named Entity Recognition (NER) model to catch unstructured person names in~5–12ms**.
By avoiding heavy NLP libraries like spaCy or HuggingFace transformers, LLM-Shield-Proxy runs inside a 24MB RAM process footprint — making it fast, deterministic, and ideal for microservice sidecars. This means you can run dozens of proxy containers side-by-side on cheap micro-instances (like AWS t4g.nano
or Docker Swarm/Kubernetes pods) for virtually zero RAM cost.
Zero-Egress Security: 100% of PII scanning and re-hydration happens locally within your VPC. No prompt data or telemetry ever leaves your server.Stateless Privacy (Self-Destructing Redis TTL): Real PII is mapped to session-bound tokens (e.g.Sarah
->[PERSON_1]
) stored in an in-memory vault backed by strict Time-To-Live (TTL) expiration rules. When configured with Redis (REDIS_URL
), vaults are shared across multi-replica clusters without building a permanent database of user PII.
Emits structured JSON audit events (app/audit.py
) directly to stdout
compatible with Datadog, Splunk, Elastic, Vanta, and Drata to prove compliance for SOC 2 Type II and HIPAA audits:
{
"timestamp": "2026-08-04T01:48:00Z",
"event": "pii_redaction",
"session_id": "sess_8f179f3",
"path": "/v1/chat/completions",
"redactions_summary": {
"SSN": 1,
"EMAIL": 2,
"PERSON": 1
},
"compliance_status": "zero_egress_passed"
}
flowchart TD
classDef client fill:#e0f2fe,stroke:#0284c7,stroke-width:2px,color:#0369a1,font-weight:bold;
classDef proxyEngine fill:#f8fafc,stroke:#475569,stroke-width:2px,color:#0f172a,font-weight:bold;
classDef piiSecurity fill:#fef2f2,stroke:#ef4444,stroke-width:2px,color:#991b1b,font-weight:bold;
classDef vault fill:#fffbebe,stroke:#f59e0b,stroke-width:2px,color:#92400e,font-weight:bold;
classDef upstream fill:#f3e8ff,stroke:#9333ea,stroke-width:2px,color:#6b21a8,font-weight:bold;
UserApp["👤 User Application\n(OpenAI / LangChain SDK)"]:::client
subgraph SecurityMoat ["🛡️ Zero-Egress Local Environment (Apache 2.0 Licensed)"]
direction TD
FastAPIProxy["⚡ FastAPI Catch-All Proxy\n(/{path:path})"]:::proxyEngine
subgraph CascadeEngine ["🔒 Two-Tier PII Cascade Engine"]
Tier1["Tier 1: Compiled Regex"]:::piiSecurity
Tier2["Tier 2: Quantized ONNX NER"]:::piiSecurity
Tier1 --> Tier2
end
VaultStore[("🔑 Session Vault Store\n(Deterministic Tokens)")]:::vault
LookaheadBuffer["⏱️ Sliding-Window Lookahead Buffer\n(Prevent SSE Tag Leaks)"]:::proxyEngine
Rehydrator["🔄 Stream Re-hydrator\n(Token -> Original Value)"]:::proxyEngine
end
UpstreamLLM["☁️ Upstream LLM Provider\n(OpenAI / Anthropic / vLLM)"]:::upstream
%% Inbound Flow (Prompt Sanitization)
UserApp -- "1. Inbound Raw Prompt Payload" --> FastAPIProxy
FastAPIProxy -- "2. Scan Payload" --> Tier1
Tier2 -- "3. Store Vault Keys" --> VaultStore
Tier2 -- "4. Redacted JSON Payload" --> UpstreamLLM
%% Outbound Flow (Streaming De-redaction)
UpstreamLLM -. "5. Raw SSE Stream Deltas" .-> LookaheadBuffer
LookaheadBuffer -- "6. Tag-Safe Assembly" --> Rehydrator
Rehydrator <--> VaultStore
Rehydrator -. "7. Sanitized Real-Time Stream" .-> UserApp
style SecurityMoat fill:#f8fafc,stroke:#0284c7,stroke-width:2px,stroke-dasharray: 5 5,color:#0f172a
style CascadeEngine fill:#ffffff,stroke:#cbd5e1,stroke-width:1px
Intercept: Your application sends a standard OpenAI / LangChain payload tolocalhost:8000
.Cascade Redaction: The proxy intercepts the JSON and routes text through a high-speed compiled Regex engine (SSNs, emails, credit cards), falling back to a local ONNX model for unstructured names.Vault Storage: The original PII is mapped to a deterministic tag (e.g.,[PERSON_1]
) and stored locally in a TTL-backed session vault.Clean Egress: A 100% sanitized payload is forwarded to OpenAI. OpenAI never sees your raw sensitive data.
SSE Stream Intercept: OpenAI streams the response back chunk-by-chunk via Server-Sent Events (SSE).Lookahead Buffer: Because tags can be split across SSE chunks (e.g.,[PER
in chunk N andSON_1]
in chunk N+1), the proxy's sliding-window buffer holds back unclosed brackets to prevent tag leaks.Re-hydration: Once a tag is fully assembled, the proxy swaps the real data back from the local vault and streams the final, un-redacted text to the user's application in real-time.
LLM-Shield-Proxy is engineered for sub-millisecond overhead and ultra-lightweight resource usage. Measured over 1,000 production streaming iterations:
| Metric | Average Latency | Median Latency | Footprint / Notes |
|---|---|---|---|
| Tier 1 Regex Overhead | |||
0.0294 ms |
|||
0.0291 ms (29.10 µs ) |
|||
| Microsecond pattern scan | |||
| Tier 2 NER Overhead | |||
0.0033 ms |
|||
0.0032 ms (3.20 µs ) |
|||
| Quantized local NER scan | |||
| Total SSE Stream Overhead | |||
0.0010 ms |
|||
0.0010 ms (0.97 µs ) |
|||
| Added latency per SSE delta chunk | |||
| Process RAM Footprint |
- | - |
24.55 MBResident Set Size |
To run the automated benchmark suite locally:
py tests/benchmark.py
Transparency is critical for security tooling. Please be aware of the following current limitations:
Text Only: The proxy does not currently scan or redact text embedded inside base64 image payloads (e.g., OpenAI Vision models).Supported Languages: The Tier-2 ONNX NER model is currently optimized for English-language entities.Non-Standard Streaming: Designed for standard Server-Sent Events (SSE). Custom or proprietary streaming protocols may bypass the sliding-window buffer.
Run the full automated test suite:
py -m pytest tests/
Designed for zero-friction adoption by DevOps, Site Reliability Engineers (SREs), and Network Administrators:
Built-in liveness and readiness endpoints return HTTP 200 OK
for Kubernetes, Docker Swarm, or AWS ECS health monitors:
curl http://localhost:8000/health
curl http://localhost:8000/livez
100% compliant with 12-factor app standards. All upstream target routing and API keys are injected via environment variables or a .env
file without code modifications:
UPSTREAM_BASE_URL
: Base target URL (e.g.https://api.openai.com
or internalvLLM
server).OPENAI_API_KEY
: Upstream API key passed to target providers.REDIS_URL
: Optional Redis connection string for distributed multi-instance session caching.
LLM-Shield-Proxy runs completely stateless by default. For high-volume enterprise deployments, instances scale horizontally behind edge proxies (NGINX, Traefik, AWS ALB):
docker-compose up -d --scale proxy=5
When configured with REDIS_URL
, session vaults are shared across all proxy replicas, ensuring seamless session isolation across multi-instance clusters.
Every published release includes automated SHA-256 checksums (checksums.txt
) and GPG detached signatures (checksums.txt.asc
) signed by maintainer Ninad Phalak. You can verify checksums and cryptographic authenticity before deployment using:
sha256sum -c checksums.txt
Get-FileHash llm-shield-proxy-source-v1.0.4.zip -Algorithm SHA256
gpg --verify checksums.txt.asc checksums.txt
Currently, LLM-Shield-Proxy's Tier 1 Regex engine is optimized for North American PII (US SSNs, Phone Formats). To support global GDPR compliance, I am actively looking for contributors to help expand regex payloads and Tier 2 ONNX models for:
European Formats: UK NIN, EU Phone Numbers, IBANs.APAC Data Structures: India Aadhaar, APAC localized identifiers.Multilingual NER ONNX Models: Multilingual entity recognition models.
If you want to contribute to enterprise AI security, check out CONTRIBUTING.md and claim a locale!
I am committed to maintaining LLM-Shield-Proxy as the fastest ultra-low latency redaction engine for LLMs. Here are the core architectural optimizations planned for upcoming releases — contributions and PRs are warmly welcomed:
ONNX Thread Tuning (Preventing CPU Contention)* Problem:*By default, ONNX Runtime attempts to use every available CPU core. In FastAPI, this competes with the event loop handling thousands of concurrent connections.*The Fix:*Restrict ONNX by settingsess_options.intra_op_num_threads = 1
. This forces ONNX execution onto a single thread, keeping CPU cores free for FastAPI's event loop to stream packets instantly.
Persistent Connection Pooling (The TLS Trick)* Problem:*Opening a new TLS/SSL connection to OpenAI per request adds 50–100ms latency.*The Fix:*Maintain a persistenthttpx.AsyncClient
HTTP/2 connection pool on server startup. The proxy opens pre-warmed secure tunnels, routing requests instantly with zero TLS setup overhead.
Swap toorjson
for Chunk Parsing*Problem:*In an SSE stream, standard Pythonjson.loads
parses hundreds of delta chunks per second.*The Fix:*Swap built-injson
fororjson
(written in Rust). It parses streaming LLM chunks up to 10x faster, dropping proxy overhead to near zero.
Cythonize the Sliding-Window Buffer* Problem:*The sliding-window buffer performs frequent string slicing and bracket matching.*The Fix:*Use Cython ormypyc
to compilestreaming.py
directly into a C-extension binary module. Retains Python readability while executing string operations at native C speed.
I am actively working with enterprise security teams to map out advanced compliance features. If your startup or organization is using LLM-Shield-Proxy to unblock LLM streaming or pass SOC 2/HIPAA audits, I would love to hear from you.
Email the core maintainer at ninadphalak@gmail.com to share your feedback, request a feature, or feature your team as a case study.