RAI Guardrails
Overview
IAF integrates Responsible AI (RAI) Guardrails to ensure all interactions between users and LLM-powered agents are safe, compliant, and privacy-preserving. The guardrails automatically screen every user input and (optionally) every LLM response for harmful content and sensitive personal information before anything reaches the model or is returned to the user.
Two categories of guardrails are enforced:
| Category | What It Does |
|---|---|
| Content Moderation | Detects and blocks jailbreak attempts, toxic language, profanity, and restricted topics |
| PII Privacy Protection | Detects personally identifiable information (PII), blocks critical entities, and anonymises the rest before sending to the LLM |
These guardrails are powered by the Infosys Responsible AI Toolkit and are transparently integrated — no changes are required to agent configurations or user workflows once enabled.
LiteLLM Proxy Required
RAI Guardrails are implemented within the LiteLLM Proxy layer. To enable or use RAI Guardrails, you must set USE_LITELLM_PROXY_FLAG=true in your IAF environment configuration. Without this flag, LLM requests bypass the proxy and no guardrail checks will be applied.
Guardrail Checks
1. Jailbreak Detection
Identifies attempts to manipulate the LLM into bypassing its safety instructions through prompt injection or jailbreak techniques.
- What it detects: Prompt injection patterns, role-play instructions designed to override system prompts, instructions asking the model to ignore previous guidelines
- Threshold: Configurable (default: 0.7) — scores at or above this value trigger a block
- Action on violation: Request is blocked and the user receives a Content Policy Alert
2. Toxicity Analysis
Evaluates user input for harmful, hateful, or dangerous language across multiple dimensions.
Sub-checks include:
| Sub-check | What It Detects |
|---|---|
| General Toxicity | Broadly harmful or negative content |
| Severe Toxicity | Extremely harmful content |
| Obscenity | Obscene or vulgar content |
| Threats | Language indicating intent to harm |
| Insults | Demeaning or degrading language |
| Identity Attacks | Content targeting specific identity groups |
| Sexually Explicit | Inappropriate sexual content |
- Threshold: Configurable per sub-check (default: 0.6 for all)
- Action on violation: Request is blocked if any sub-check exceeds its threshold
3. Profanity Filtering
Scans input text for profane or offensive words.
- What it detects: Known profane words and phrases in the user's message
- Threshold: Number of profane words allowed before triggering (default: 1 — meaning even a single profane word triggers a block)
- Action on violation: Request is blocked and flagged words are reported
4. Restricted Topic Detection
Prevents the LLM from being used to discuss or generate content about topics that your organisation has deemed off-limits.
- Default restricted topics: Terrorism, Explosives, Nudity, Cruelty, Cheating, Fraud, Crime, Hacking, Immoral, Unethical, Illegal, Robbery, Forgery, Misinformation
- Threshold: Topic relevance score (default: 0.7)
- Customisable: Administrators can add or remove topics based on organisational policies
- Action on violation: Request is blocked and the detected topic is reported
5. PII (Personally Identifiable Information) Protection
Automatically detects sensitive personal data in user messages using a RoBERTa NLP model and applies one of three actions per entity type:
| Action | Behaviour | Use Case |
|---|---|---|
| Block | Entire request is rejected | For highly sensitive entities (e.g., email addresses) that should never be sent to an LLM |
| Anonymise | Entity is replaced with a placeholder (e.g., [PHONE_NUMBER_1]) before sending to the LLM, then restored in the response |
For entities that need privacy protection but the conversation should continue |
| Skip | Entity is left unchanged | For entities that are acceptable to send to the LLM (e.g., location names, organisation names) |
Detectable PII entity types include:
- Email addresses
- Phone numbers
- Credit card numbers
- Social security numbers
- Person names
- Locations
- Dates
- URLs
- Organisation names
- Nationalities / Religious / Political groups (NRP)
The classification of which entities to block, anonymise, or skip is fully configurable via environment variables.
How It Works
Architecture
The guardrails operate through a proxy layer (LiteLLM Proxy) that sits between IAF and the LLM provider. This architecture ensures guardrails are applied consistently to all agent types without modifying individual agent logic.
┌────────────────────┐
│ User sends │
│ message to agent │
└────────┬───────────┘
▼
┌────────────────────┐
│ IAF Agent │
│ (any type) │
└────────┬───────────┘
▼
┌────────────────────────────────────────────────────────────┐
│ LiteLLM Proxy — Pre-call Guardrails │
│ │
│ Step 1: PII Protection │
│ • Scan for personal information │
│ • Block if critical PII found │
│ • Anonymise remaining PII with placeholders │
│ │
│ Step 2: Content Moderation │
│ • Check for jailbreak attempts │
│ • Check for toxicity │
│ • Check for profanity │
│ • Check for restricted topics │
│ • Block if any check fails │
└────────┬────────────────────────────────────────────────────┘
▼ (safe, anonymised request)
┌────────────────────┐
│ LLM Provider │
│ (Azure/OpenAI) │
└────────┬───────────┘
▼ (LLM response)
┌────────────────────────────────────────────────────────────┐
│ LiteLLM Proxy — Post-call Guardrails │
│ │
│ Step 3: Response Moderation (optional) │
│ • Check LLM output for harmful content │
│ • Replace with safe message if violated │
│ │
│ Step 4: PII De-anonymisation │
│ • Restore original values in place of placeholders │
└────────┬────────────────────────────────────────────────────┘
▼
┌────────────────────┐
│ Response returned │
│ to user │
└─────────────────────┘
What Gets Checked
Only user-provided content is moderated. System instructions and framework-generated prompts are excluded from guardrail checks to avoid false positives. The system intelligently identifies user content by looking for specific markers in the message (e.g., "User Query:" or "Input Query:").
Response Moderation (Optional)
By default, only user inputs are moderated. Optionally, you can also enable output moderation to check LLM responses for harmful content. When enabled:
- If the LLM response violates content policy, it is silently replaced with a safe message
- The request is not blocked (since the user's input was fine) — only the response is sanitised
This is recommended for customer-facing applications where LLM output safety is critical.
User Experience
When a guardrail is triggered, the user receives a clear, structured alert explaining what happened:
Example — Content Policy Violation:
**Content Policy Alert**
Status: FAILED
Check Results:
✅ JailBreak — PASSED (score: 0.12, threshold: 0.70)
❌ Toxicity — FAILED (score: 0.85, threshold: 0.60)
• ToxicityScore: 0.85
• InsultScore: 0.72
✅ Profanity — PASSED
✅ RestrictTopic — PASSED
Example — PII Privacy Violation:
**Privacy Protection Alert**
Your request was blocked because it contains sensitive personal information:
• EMAIL_ADDRESS detected (confidence: 0.99, position: 10-26)
Entities configured to block: EMAIL_ADDRESS
Please remove the identified information and try again.
The guardrail message is returned directly to the user — no further processing or tool calls are attempted once a violation is detected.
Integration with Agent Types
RAI Guardrails are integrated across all agent types in IAF. Once enabled, every agent automatically benefits from guardrail protection:
| Agent Type | Guardrail Behaviour |
|---|---|
| React Agent | Catches violations during LLM call, skips tool execution, returns guardrail alert directly |
| React Critic Agent | Same as React — violation detected before critic evaluation |
| Planner Executor Agent | Scans planner output for violations; catches executor-level violations; early exit on detection |
| Meta Agent | Catches violations at sub-agent delegation; skips remaining sub-agents on block |
| Planner Meta Agent | Combines planner scanning with meta-agent delegation protection |
| Hybrid Agent | Detects violations in agent response and returns formatted alert to user |
| Workflows | Each sub-agent within the workflow enforces guardrails independently at its own step |
Tool Onboarding
Guardrails also protect the tool onboarding process. When new tools are registered, their code is sent through the LLM for validation and documentation generation. If the tool code contains content that violates policy (e.g., references to restricted topics), the onboarding process returns a clear policy-violation message rather than failing silently.
Audit Logging
Every guardrail check (both passes and failures) is logged to a PostgreSQL database for compliance and observability purposes. The audit log captures:
- Timestamp of the check
- Whether it was a pre-call (input) or post-call (output) check
- The content that was moderated
- Overall pass/fail status
- Individual check results and scores (jailbreak, toxicity, profanity, restricted topics)
- PII entities detected and actions taken
- Full API response for debugging
This enables security teams to review moderation activity, identify patterns, and tune thresholds over time.
Failure Handling & Resilience
The guardrails follow a fail-open strategy — if the RAI Toolkit APIs are unreachable, requests pass through to the LLM without guardrail protection rather than blocking all traffic. This ensures service availability is not impacted by guardrail infrastructure issues.
| Scenario | Behaviour |
|---|---|
| RAI Moderation API unreachable | Warning logged; request passes to LLM unmoderated |
| PII Analyze API unreachable | Warning logged after retry; request passes without PII protection |
RAI Moderation returns PASSED |
Request passes unchanged |
| No PII entities detected | Request passes unchanged |
| Guardrail APIs not configured | Guardrails are skipped entirely |
Configuration
Enabling Guardrails in IAF
Set the following environment variables in the IAF .env file:
# Route LLM traffic through the proxy (required for guardrails)
USE_LITELLM_PROXY_FLAG=true
LITELLM_ENDPOINT=http://<proxy-host>:<proxy-port>
LITELLM_API_KEY=<api-key>
LITELLM_API_VERSION=<azure-api-version>
LITELLM_MODELS=<comma-separated-model-names>
Warning
USE_LITELLM_PROXY_FLAG=true must be set for end-to-end guardrail protection. The guardrail type is configured per agent via the Guardrail Type dropdown during agent creation.
Selecting Guardrail Type per Agent
When creating or onboarding an agent, use the Guardrail Type dropdown to select the guardrail checks to apply for that specific agent. The selected guardrail type is sent to the LiteLLM Proxy, which uses a guardrail router to apply the appropriate checks.
Info
To remove guardrails from an agent, update the agent and select None from the Guardrail Type dropdown.
Guardrail Checks for Non-Inference LLM Calls
To enable guardrail checks for non-inference LLM calls within the framework, navigate to the Admin Screen, open the Config tab, and select the required guardrail from the Guardrail dropdown.
Configuring the LiteLLM Proxy
The proxy requires the following environment variables:
| Variable | Description |
|---|---|
RAI_API_URL |
URL of the RAI Moderation API |
PII_ANALYZE_URL |
URL of the PII Privacy Analyze API |
PII_ENTITIES_TO_BLOCK |
Comma-separated PII entity types to hard-block (e.g., EMAIL_ADDRESS) |
PII_ENTITIES_TO_SKIP |
Comma-separated PII entity types to leave unchanged (e.g., LOCATION,PERSON,DATE_TIME,NRP,ORGANIZATION,URL) |
DATABASE_URL |
PostgreSQL connection string for audit logging |
Guardrail Router (LiteLLM Proxy — server.py)
The LiteLLM Proxy implements a GuardrailRouter in server.py that routes each incoming LLM request to the correct guardrail provider based on the x-guardrail-provider HTTP header sent by the IAF server. The header value corresponds to the guardrail type selected for the agent.
- If the header is absent or empty, the request passes through without moderation (equivalent to
guardrail_type = "none"). - Currently, the Infosys RAI provider is registered by default.
- To integrate a custom guardrail, implement the guardrail service and register it with the router using
guardrail_router.register(...).
Registering a Custom Guardrail
Call guardrail_router.register() with a unique key, the service instance, a display label, and an optional description. Once registered, the new provider appears automatically in the Guardrail Type dropdown in the IAF UI.
Guardrail Registry (IAF Server — guardrail_helper.py)
On startup, the IAF server calls GuardrailRegistry.sync_from_proxy() to fetch all registered guardrail providers from the LiteLLM Proxy (GET /guardrail-providers). This drives the Guardrail Type dropdown in the agent creation/update UI — no hardcoded guardrail types exist in IAF.
- If the proxy is unreachable at startup, the registry stays empty and only
"none"is valid. - Any new provider registered in
server.pyis automatically surfaced to the IAF dropdown after a restart.
Customising Thresholds
All moderation thresholds are configurable in the proxy's constants.py file:
| Check | Default Threshold | Description |
|---|---|---|
| Jailbreak | 0.7 | Score at or above triggers block |
| Prompt Injection | 0.7 | Score at or above triggers block |
| Toxicity (all sub-checks) | 0.6 | Score at or above triggers block |
| Profanity | 1 word | Number of profane words to trigger |
| Restricted Topics | 0.7 | Topic relevance score to trigger |
| PII Detection Confidence | 0.7 | Minimum confidence to consider a PII entity valid |
Customising Restricted Topics
Add or remove topics in the constants.py configuration:
"Restrictedtopics": [
"Terrorism", "Explosives", "Nudity", "Cruelty", "Cheating",
"Fraud", "Crime", "Hacking", "Immoral", "Unethical",
"Illegal", "Robbery", "Forgery", "Misinformation",
# Add your custom topics here
]
Prerequisites
- Infosys RAI Toolkit — The Moderation and Privacy modules must be deployed on VMs and accessible via HTTP
- LiteLLM Proxy — Must be set up and running with guardrails registered in
config.yaml - PostgreSQL — Required for audit logging (shared with IAF's existing database or separate)
- Kafka Workers — Must be running if using guardrails with scheduled (cron) agent executions
Hosting RAI Guardrails on a VM
The RAI Guardrails rely on two open-source modules from the Infosys Responsible AI Toolkit — the Moderation Layer and the Privacy module. Both must be deployed and running as HTTP services before IAF can use guardrail protection.
Modules to Deploy:
| Module | Purpose | Repository |
|---|---|---|
| Moderation Layer | Provides jailbreak detection, toxicity analysis, profanity filtering, and restricted topic checks | responsible-ai-moderationlayer |
| Privacy Module | Provides PII detection, entity classification, anonymisation, and de-anonymisation | responsible-ai-privacy |
Deployment Steps (Summary):
-
Clone the Infosys Responsible AI Toolkit repository on your VM:
git clone https://github.com/Infosys/Infosys-Responsible-AI-Toolkit.git -
Follow the setup instructions in each module's README:
- Moderation Layer:
responsible-ai-moderationlayer/— install dependencies, configure models, and start the API server - Privacy Module:
responsible-ai-privacy/— install dependencies, configure PII detection models, and start the API server
- Moderation Layer:
-
Once both services are running, note their URLs (e.g.,
http://<vm-ip>:8081for moderation,http://<vm-ip>:8082for privacy) -
Configure the LiteLLM Proxy environment variables to point to these services:
RAI_API_URL=http://<vm-ip>:<moderation-port> PII_ANALYZE_URL=http://<vm-ip>:<privacy-port>/analyze
Deployment Recommendations
- Deploy both modules on the same VM or within the same network as the LiteLLM Proxy for low-latency checks
- Ensure the VM has a GPU if you need faster inference for the ML models used in moderation and PII detection
- Monitor service health and set up restart policies (e.g., via systemd or Docker) to ensure availability
Quick-Start Checklist
- Deploy the Infosys RAI Toolkit APIs (Moderation + Privacy modules) on VMs
- Set up the LiteLLM proxy with
.env,config.yaml, andconstants.py - Install proxy dependencies and start the proxy server
- Configure IAF
.envwithUSE_LITELLM_PROXY_FLAG=true - Restart IAF — all agents will automatically use guardrail protection
- Verify by sending a test message with known policy-violating content and confirming the guardrail alert is returned