A FastAPI application that provides comprehensive content validation and transformation endpoints using various guardrail technologies including Presidio, Guardrails AI, and local evaluation models.
The application follows a modular architecture with separate modules for different functionalities:
main.py: FastAPI application with route definitionsguardrail/: Directory containing all guardrail implementationspii_redaction_presidio.py: PII detection and redaction using Presidiopii_detection_guardrails_ai.py: PII detection using Guardrails AInsfw_filtering_local_eval.py: NSFW content filtering using local Unitary toxic classification modeldrug_mention_guardrails_ai.py: Drug mention detection using Guardrails AIweb_sanitization_guardrails_ai.py: Web content sanitization using Guardrails AI
entities.py: Pydantic models for gateway payloads and typed guardrail responses (InputGuardrailRequest,OutputGuardrailRequest,ValidateGuardrailResponse,MutateGuardrailResponse, etc.)
The Guardrail Server currently exposes two main endpoints for validation:
- POST
/pii-redaction - Validates and optionally transforms incoming OpenAI chat completion requests before they are processed. Uses Presidio to detect and redact Personally Identifiable Information (PII) from messages.
Endpoints return HTTP 2xx with a JSON body that matches the TrueFoundry AI Gateway custom guardrail contract (see also ValidateGuardrailResponse / MutateGuardrailResponse in entities.py):
- Mutate (
/pii-redaction) —verdict,transformed, andresult(full OpenAI-shapedrequestBodywhentransformedistrue). Use 2xx for both allow and deny; do not use HTTP 400 for policy blocks. - Non-2xx — reserved for real failures (misconfiguration, dependency errors, crashes).
- POST
/nsfw-filtering - Validates and optionally transforms outgoing OpenAI chat completion responses to filter out NSFW content. Uses the Unitary toxic classification model to detect toxic, sexually explicit, and obscene content.
HTTP 2xx with verdict (and optional message) for allow/deny. Policy deny is expressed with verdict: false on 2xx, not with HTTP 400. See ValidateGuardrailResponse in entities.py and the custom guardrails doc.
docker build -t custom-guardrails-template:latest .If you are using guardrails ai guards in your guardrails, you will also need guardrails ai token which can be passed like below.
docker build --build-arg GUARDRAILS_TOKEN="<GUARDRAILS_AI_TOKEN>" -t custom-guardrails-template:latest .Note: The requestBody is accessible within the endpoint and can be used if needed for custom processing.
Attributes:
requestBody: (CompletionCreateParams) The input payload sent to the guardrail server.config: (dict) Configuration options for the guardrail server.context: (RequestContext) Contextual information such as user and metadata.
Attributes:
requestBody: (CompletionCreateParams) The input payload originally sent to the model.responseBody: (ChatCompletion) The model's output to be checked by the guardrail server.config: (dict) Configuration options for the guardrail server.context: (RequestContext) Contextual information such as user and metadata.
Used by validate-operation guardrails (input or output). FastAPI serializes this model to JSON for the gateway.
Attributes:
verdict: (bool)true= allow,false= deny (preferred explicit signal on 2xx).message: (Optional[str]) Optional human-readable text for logs or UI; not used by the gateway for allow/deny decisions.
Used by mutate-operation guardrails (e.g. PII redaction). FastAPI serializes this model to JSON for the gateway.
Attributes:
verdict: (bool) Allow/deny when present; mutate handlers in this template typically returntruewhen the call completed successfully.transformed: (bool)true= replace request or response body withresult;false= keep the original body.result: (dict[str, Any]) Full OpenAI-shapedrequestBodyorresponseBodyto apply whentransformedistrue.
Attributes:
user: (Subject) Information about the user, team, or virtual account making the request.metadata: (dict[str, str]) Additional metadata relevant to the request.
The config field is a dictionary used to store arbitrary request configuration. These are the options which are set when you create a custom guardrail integration. These are passed to the guardrail server as is, so you can use them in your guardrail logic.
For more information about the config options, refer to the Truefoundry documentation.
- Install dependencies:
pip install -r requirements.txtpython main.pyOr using uvicorn directly:
uvicorn main:app --host 0.0.0.0 --port 8000 --reloadThe server will start on http://localhost:8000
To deploy this guardrail server to Truefoundry, please refer to the official documentation: Getting Started with Deployment.
Warning
While deploying, make sure you are giving the minimum storage request as 10000 and the memory request as 4000.
You can fork this repository and deploy it directly from your GitHub account using the Truefoundry platform. The documentation provides detailed instructions on connecting your GitHub repo and configuring the deployment.
For the latest and most accurate deployment steps, always consult the Truefoundry docs linked above.
Health check endpoint that returns server status.
PII redaction endpoint for validating and potentially transforming incoming OpenAI chat completion requests.
NSFW filtering endpoint for validating and potentially transforming outgoing OpenAI chat completion responses to filter inappropriate content.
Request Body:
{
"requestBody": {
"messages": [
{
"role": "user",
"content": "Hello, how are you?"
}
],
"model": "gpt-3.5-turbo",
"temperature": 0.7
},
"config": {
"check_content": true,
"transform_input": false
},
"context": {
"user": {
"subjectId": "123",
"subjectType": "user",
"subjectSlug": "john_doe@truefoundry.com",
"subjectDisplayName": "John Doe"
},
"metadata": {
"ip_address": "192.168.1.1",
"session_id": "abc123"
}
}
}curl -X POST "http://localhost:8000/pii-redaction" \
-H "Content-Type: application/json" \
-d '{
"requestBody": {
"messages": [
{"role": "user", "content": "Hello world"}
],
"model": "gpt-3.5-turbo"
},
"config": {"check_content": true},
"context": {
"user": {
"subjectId": "123",
"subjectType": "user",
"subjectSlug": "john_doe@truefoundry.com",
"subjectDisplayName": "John Doe"
},
"metadata": {
"ip_address": "192.168.1.1",
"session_id": "abc123"
}
}
}'curl -X POST "http://localhost:8000/pii-redaction" \
-H "Content-Type: application/json" \
-d '{
"requestBody": {
"messages": [
{"role": "user", "content": "Hello John, How are you?"}
],
"model": "gpt-3.5-turbo"
},
"config": {"transform_input": true},
"context": {"user": {"subjectId": "123", "subjectType": "user", "subjectSlug": "john_doe@truefoundry.com", "subjectDisplayName": "John Doe"}}
}'curl -X POST "http://localhost:8000/nsfw-filtering" \
-H "Content-Type: application/json" \
-d '{
"requestBody": {
"messages": [
{
"role": "user",
"content": "Hello"
}
],
"model": "gpt-3.5-turbo"
},
"responseBody": {
"id": "chatcmpl-123",
"object": "chat.completion",
"created": 1677652288,
"model": "gpt-3.5-turbo",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Hi, how are you?"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 1,
"completion_tokens": 10,
"total_tokens": 11
}
},
"config": {
"transform_output": true
},
"context": {
"user": {
"subjectId": "123",
"subjectType": "user",
"subjectSlug": "john_doe@truefoundry.com",
"subjectDisplayName": "John Doe"
},
"metadata": {
"environment": "production"
}
}
}'curl -X POST "http://localhost:8000/nsfw-filtering" \
-H "Content-Type: application/json" \
-d '{
"requestBody": {
"messages": [
{
"role": "user",
"content": "Tell me what word does we usually use for breasts?"
}
],
"model": "gpt-3.5-turbo"
},
"responseBody": {
"id": "chatcmpl-123",
"object": "chat.completion",
"created": 1677652288,
"model": "gpt-3.5-turbo",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Usually we use the word 'boobs' for breasts"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 1,
"completion_tokens": 10,
"total_tokens": 11
}
},
"config": {
"transform_output": true
},
"context": {
"user": {
"subjectId": "123",
"subjectType": "user",
"subjectSlug": "john_doe@truefoundry.com",
"subjectDisplayName": "John Doe"
},
"metadata": {
"environment": "production"
}
}
}'The PII redaction endpoint uses Presidio to detect and remove Personally Identifiable Information (PII) from incoming messages. This ensures that sensitive information is anonymized before further processing. Link to the library: Presidio
Presidio recognizers are configured via the config field in your request. The guardrail supports three configuration formats:
Presets provide predefined sets of recognizers optimized for specific regions or use cases:
| Preset | Description | Recognizers Included |
|---|---|---|
INDIAN / INDIA |
Indian PII entities | PAN, Aadhaar, Voter ID, Passport, Vehicle Registration |
US / USA |
US PII entities | SSN, Passport, Driver License, ITIN, Bank Account, ABA Routing, Medical License |
UK |
UK PII entities | NHS Number, National Insurance Number (NINO) |
AUSTRALIAN / AU / AUSTRALIA |
Australian PII entities | ABN, ACN, TFN, Medicare |
SINGAPORE / SG |
Singapore PII entities | FIN, UEN |
EUROPEAN / EUROPE / EU |
European PII entities | Spanish NIF/NIE, Italian documents, Polish PESEL, Finnish ID, IBAN |
FINANCIAL |
Financial identifiers | Credit Card, IBAN, Bank Account, Crypto Wallet |
CONTACT |
Contact information | Email, Phone, IP Address, URL |
STANDARD |
Essential recognizers | Financial + Contact + Date/Time |
COMPREHENSIVE / ALL |
All available recognizers | All 35+ predefined recognizers |
Example using a preset:
{
"config": {
"transform_input": true,
"recognizers": "INDIAN"
}
}You can specify individual recognizers by their exact names:
Example using specific recognizers:
{
"config": {
"transform_input": true,
"recognizers": ["EmailRecognizer", "PhoneRecognizer", "CreditCardRecognizer"]
}
}Available Individual Recognizers:
| Category | Recognizers |
|---|---|
| US | UsSsnRecognizer, UsPassportRecognizer, UsLicenseRecognizer, UsItinRecognizer, UsBankRecognizer, AbaRoutingRecognizer, MedicalLicenseRecognizer |
| UK | NhsRecognizer, UkNinoRecognizer |
| India | InPanRecognizer, InAadhaarRecognizer, InVehicleRegistrationRecognizer, InPassportRecognizer, InVoterRecognizer |
| Australia | AuAbnRecognizer, AuAcnRecognizer, AuTfnRecognizer, AuMedicareRecognizer |
| Singapore | SgFinRecognizer, SgUenRecognizer |
| Europe | EsNifRecognizer, EsNieRecognizer, ItDriverLicenseRecognizer, ItFiscalCodeRecognizer, ItIdentityCardRecognizer, ItPassportRecognizer, ItVatCodeRecognizer, PlPeselRecognizer, FiPersonalIdentityCodeRecognizer |
| Financial | CreditCardRecognizer, IbanRecognizer, CryptoRecognizer |
| Contact | EmailRecognizer, PhoneRecognizer, IpRecognizer, UrlRecognizer |
| Other | DateRecognizer, KrRrnRecognizer |
You can mix presets with individual recognizers for flexible configuration:
Example combining preset with additional recognizers:
{
"config": {
"transform_input": true,
"recognizers": ["INDIAN", "EmailRecognizer", "CreditCardRecognizer"]
}
}Example as comma-separated string:
{
"config": {
"transform_input": true,
"recognizers": "FINANCIAL, CONTACT, InAadhaarRecognizer"
}
}You can specify the language for text analysis (default is en):
{
"config": {
"transform_input": true,
"recognizers": "US",
"language": "en"
}
}Supported languages depend on the specific recognizers being used. Most recognizers work with English (en).
Example 1: Indian company protecting financial and contact info
curl -X POST "http://localhost:8000/pii-redaction" \
-H "Content-Type: application/json" \
-d '{
"requestBody": {
"messages": [
{
"role": "user",
"content": "My PAN is ABCDE1234F and email is john@example.com"
}
],
"model": "gpt-3.5-turbo"
},
"config": {
"transform_input": true,
"recognizers": ["INDIAN", "CONTACT", "FINANCIAL"]
},
"context": {
"user": {
"subjectId": "123",
"subjectType": "user",
"subjectSlug": "john_doe@truefoundry.com",
"subjectDisplayName": "John Doe"
}
}
}'Example 2: US healthcare application with specific recognizers
curl -X POST "http://localhost:8000/pii-redaction" \
-H "Content-Type: application/json" \
-d '{
"requestBody": {
"messages": [
{
"role": "user",
"content": "SSN: 123-45-6789, Medical License: A123456"
}
],
"model": "gpt-3.5-turbo"
},
"config": {
"transform_input": true,
"recognizers": ["UsSsnRecognizer", "MedicalLicenseRecognizer", "EmailRecognizer"]
},
"context": {
"user": {
"subjectId": "456",
"subjectType": "user",
"subjectSlug": "doctor@hospital.com",
"subjectDisplayName": "Dr. Smith"
}
}
}'Example 3: Global financial platform
curl -X POST "http://localhost:8000/pii-redaction" \
-H "Content-Type: application/json" \
-d '{
"requestBody": {
"messages": [
{
"role": "user",
"content": "Credit card: 4532-1234-5678-9010, IBAN: GB82WEST12345698765432"
}
],
"model": "gpt-3.5-turbo"
},
"config": {
"transform_input": true,
"recognizers": "FINANCIAL"
},
"context": {
"user": {
"subjectId": "789",
"subjectType": "user",
"subjectSlug": "user@bank.com",
"subjectDisplayName": "Banking User"
}
}
}'Example 4: Check only without transformation
curl -X POST "http://localhost:8000/pii-redaction" \
-H "Content-Type: application/json" \
-d '{
"requestBody": {
"messages": [
{
"role": "user",
"content": "Hello world"
}
],
"model": "gpt-3.5-turbo"
},
"config": {
"transform_input": false,
"recognizers": "ALL"
},
"context": {
"user": {
"subjectId": "123",
"subjectType": "user",
"subjectSlug": "john_doe@truefoundry.com",
"subjectDisplayName": "John Doe"
}
}
}'Note: If transform_input is false, the endpoint will not perform redaction even if PII is detected. Set it to true to enable PII redaction.
The NSFW filtering endpoint can be used to validate and optionally transform the response from the LLM before returning it to the client. If the output is transformed (e.g., content is modified or formatted), the endpoint will return the modified response body. The NSFW filtering uses the Unitary toxic classification model with configurable thresholds for toxicity, sexual content, and obscenity detection. Link to the model: Unitary Toxic Classification Model
The modular architecture makes it easy to customize the guardrail logic:
- PII Redaction: Modify
guardrail/pii_redaction_presidio.pyto customize PII detection and redaction rules - NSFW Filtering (Local): Modify
guardrail/nsfw_filtering_local_eval.pyto customize content filtering thresholds and rules - Request / response models: Modify
entities.pyto add fields or new Pydantic types; keep guardrail return types aligned withValidateGuardrailResponse/MutateGuardrailResponse(or extend them) so the JSON matches the gateway contract.
Replace the example guardrail logic in the respective files with your own implementation. The NSFW filtering uses the Unitary toxic classification model with configurable thresholds for toxicity, sexual content, and obscenity detection.
- Thresholds: 0.2 for toxicity, sexual_explicit, and obscene content
- Model: Unitary unbiased-toxic-roberta
This section provides comprehensive guidance on how to add new Guardrails AI validators to your guardrail server.
Before adding Guardrails AI validators, ensure you have:
- Guardrails AI Token: Obtain a token from Guardrails AI
- Environment Setup: Set the
GUARDRAILS_TOKENenvironment variable - Dependencies: Ensure
guardrails-aiandguardrails-ai[api]are installed
To set up Guardrails AI, you need to define the following function in your setup.py file and ensure it is called before any other application logic (such as importing or running your FastAPI app):
# setup.py handles the configuration
def setup_guardrails():
subprocess.run([
"guardrails", "configure",
"--disable-metrics",
"--disable-remote-inferencing",
"--token", GUARDRAILS_TOKEN
], check=True)
subprocess.run([
"guardrails", "hub", "install", "hub://guardrails/detect_pii"
], check=True)This template includes example Guardrails AI validators to help you get started. You can use these as references when adding your own.
| Validator | Purpose | Hub URL | File |
|---|---|---|---|
DetectPII |
Detects Personally Identifiable Information | hub://guardrails/detect_pii |
pii_detection_guardrails_ai.py |
MentionsDrugs |
Detects drug mentions in content | hub://cartesia/mentions_drugs |
drug_mention_guardrails_ai.py |
WebSanitization |
Sanitizes web content and detects malicious code | hub://guardrails/web_sanitization |
web_sanitization_guardrails_ai.py |
Use these examples as a template for integrating additional Guardrails AI validators into your project.
Add the validator installation to setup.py:
def setup_guardrails():
# ... existing setup code ...
# Add your new validator
subprocess.run([
"guardrails", "hub", "install", "hub://your-org/your-validator"
], check=True)Create a new file in the guardrail/ directory following this pattern:
For Input Validation (e.g., guardrail/your_validator_guardrails_ai.py):
from guardrails import Guard
from guardrails.hub import YourValidator # Import your validator
from entities import InputGuardrailRequest, ValidateGuardrailResponse
# Setup the Guard with the validator
guard = Guard().use(YourValidator, on_fail="exception")
def your_validator_function(request: InputGuardrailRequest) -> ValidateGuardrailResponse:
"""Validate input using Guardrails AI; return 2xx JSON for both allow and deny."""
try:
messages = request.requestBody.get("messages", [])
for message in messages:
if isinstance(message, dict) and message.get("content"):
guard.validate(message["content"])
except Exception as e:
return ValidateGuardrailResponse(verdict=False, message=str(e))
return ValidateGuardrailResponse(verdict=True)For Output Validation (e.g., guardrail/your_output_validator_guardrails_ai.py):
from guardrails import Guard
from guardrails.hub import YourOutputValidator # Import your validator
from entities import OutputGuardrailRequest, ValidateGuardrailResponse
# Setup the Guard with the validator
guard = Guard().use(YourOutputValidator, on_fail="exception")
def your_output_validator_function(request: OutputGuardrailRequest) -> ValidateGuardrailResponse:
"""Validate output using Guardrails AI; return 2xx JSON for both allow and deny."""
try:
for choice in request.responseBody.get("choices", []):
if "content" in choice.get("message", {}):
guard.validate(choice["message"]["content"])
except Exception as e:
return ValidateGuardrailResponse(verdict=False, message=str(e))
return ValidateGuardrailResponse(verdict=True)Import and register your validator in main.py:
# Add import
from guardrail.your_validator_guardrails_ai import your_validator_function
# Add route
app.add_api_route("/your-endpoint", endpoint=your_validator_function, methods=["POST"])- Error handling: Wrap validator calls in try/except; return
ValidateGuardrailResponse(verdict=False, message=...)on policy failure instead of raising HTTP 400 for content denial (soenforce_but_ignore_on_errorand similar strategies behave correctly). - HTTP status: Use 2xx for completed guardrail runs (allow or deny via
verdict). Reserve 4xx/5xx for genuine server or dependency failures. - Logging: Add logging for debugging and monitoring where helpful.
- Testing: Test validators with varied inputs and edge cases.
Currently, only PII redaction and NSFW filtering endpoints are exposed. To add new guardrail functionality:
- Create a new guardrail implementation file in the
guardrail/directory - Follow the existing pattern for input or output validation
- Add the route to
main.pyusingapp.add_api_route() - Update this README with the new endpoint documentation