Walled AI Guardrails in Practice Yaala Labs detailed its integration of Walled AI's Protect and Redact guardrail APIs into its Agent Kernel agent framework, reporting that the provider delivers fast baseline safety checks and PII masking but has call-scoped placeholder IDs, text-only validation, and an INPUT_SHORT failure mode on very short messages. The team hardened the integration by bypassing redaction on INPUT_SHORT, logging and re-raising other redaction failures, and preserving prompt and trace metadata through unmasking, and it recommends treating provider behavior as distinct from application behavior. By Yaala Labs https://github.com/yaalalabs Walled AI gives us a clean, practical safety + PII masking pipeline for agent input/output flows. It is easy to integrate, fast to test, and useful for production guardrail baselines. But like every real provider integration, there are edge cases. This post covers the capabilities that made Walled AI a good fit for Agent Kernel, the real-world edge cases we encountered, and the implementation choices we made to keep behavior safe and predictable. For Agent Kernel, we needed guardrails that could do two things quickly: Walled AI gives both via: Protect for safety checks. Redact for masking sensitive values. That made it a strong fit for a provider-level guardrail option in our framework. Before discussing edge cases, it is important to highlight the strengths that made this integration practical: For many teams, this baseline is enough to ship a strong first guardrail layer quickly. At a high level: session.get non volatile cache . In Agent Kernel, we intentionally process requests individually and preserve non-text request objects files/images/other without suppressing them by default. Since Walled AI is text-focused, non-text validation should be implemented via separate hooks/policies when needed. Walled AI also supports a local moderation path through walledai/walledguard-edge Hugging Face , which is useful for: In Agent Kernel, this local path is treated as an optional complement to the default API-driven Walled AI guardrail integration. This gives teams flexibility: Reference links: Walled AI is strong for baseline safety and PII redaction, but there are provider-level limitations teams should understand before production rollout. Walled AI does not keep a session memory of placeholders across redaction calls, and it does not maintain thread history between requests. Example: my name is Person 1 my brother is Person 1 The same label can appear again for a different value in later calls. Teams should treat placeholder IDs as call-scoped unless they add an application-side session strategy. Fine-grained field-level controls can be limited for some domain requirements. Example requirement: If you need this type of selective policy, you may need an extra policy layer before or after provider redaction. The primary safety/redaction operations are text-focused. Mixed-content pipelines text + image/file/other still need runtime logic to preserve and route non-text objects correctly. These are practical issues observed while integrating current features in Agent Kernel. INPUT SHORT Very short messages like "hi", "ok", or "23" can return INPUT SHORT during redaction. Agent Kernel handles this by bypassing redaction for that request and continuing. If redaction fails for reasons other than INPUT SHORT , allowing the exception to bubble can fail the full request path. In the current implementation, Agent Kernel logs the exception and re-raises it rather than returning a fallback response. If unmasking creates a new reply object without preserving metadata, tracing tools can lose input-output linkage. In the current implementation, unmasking returns a new AgentReplyText with rewritten text and does not preserve prompt metadata by default. Based on integration feedback and code reviews, we implemented these hardening changes: If you are integrating Walled AI into your own framework/runtime: prompt , session IDs, trace metadata intact. Walled AI is a strong guardrail building block, especially for fast safety + PII integration. Its biggest advantage is speed to value: you get practical safety and masking quickly, with a clean integration model. The key is not assuming provider behavior equals application behavior. Production-safe integrations come from the combination of provider checks, session-aware runtime logic, and explicit handling of edge cases like short inputs, placeholder collisions, and tracing continuity. That combination is where reliability comes from. Originally published at kernel.yaala.ai https://kernel.yaala.ai/blog/walledai-guardrails on March 10, 2026.