PII Redaction POC
A proof-of-concept .NET solution that redacts personally identifiable information (PII) from user prompts before sending them to a large language model (LLM). The design demonstrates enterprise-grade separation of concerns using SOLID principles, dependency injection, and the Microsoft.Extensions.AI abstractions.
Purpose
Financial and customer-service prompts often contain regulated data (names, government IDs, account numbers). This POC shows how to:
- Accept a console prompt
- Detect PII using Regex, ONNX NER, and domain rules
- Replace values with stable placeholders
- Send only the sanitized prompt to an LLM (mocked for now)
Documentation
Full solution reference (architecture, NER models, routing, Tamil/Tanglish, Git setup, improvement roadmap): docs/solution-guide.md
Why Three Detection Strategies?
| Strategy | Used For | Rationale |
|---|---|---|
| Regex | Email, phone, PAN, Aadhaar, credit card | Deterministic, format-bound identifiers with stable rules that are easy to audit and test |
| ONNX NER | Person names | Contextual entities without rigid formats; names vary widely in surface form |
| Domain rules | Loan number, customer ID, account number | Business-specific identifiers defined by internal systems, not inferable from generic models alone |
Why the LLM Receives Only Sanitized Text
The placeholder map (<PERSON_1> → original value) is kept in-process for audit or downstream de-tokenization. Only the sanitized prompt crosses the LLM boundary. This reduces data-exposure risk and supports compliance requirements for regulated workloads.
Project Structure
src/
├── PiiRedaction.ConsoleApp/ # Console demo: input/output, DI bootstrap
├── PiiRedaction.TestHarness.Wpf/ # WPF MVVM test harness for manual POC validation
├── PiiRedaction.Core/ # Business logic: detection, redaction, models
└── PiiRedaction.Infrastructure/ # Technical adapters: ONNX Runtime, mock LLM
models/ # Optional ONNX model files (gitignored)
| Project | Responsibility |
|---|---|
PiiRedaction.ConsoleApp |
Read prompt, call sanitizer, display results, call LLM service |
PiiRedaction.TestHarness.Wpf |
Desktop test harness: preset prompts, redact UI, batch validation |
PiiRedaction.Core |
PII detection abstractions, redaction, sanitization orchestration |
PiiRedaction.Infrastructure |
ONNX model runner, IChatClient mock implementation |
Prerequisites
- .NET SDK 10.x (or compatible SDK for
net10.0) - ONNX NER model for person-name detection (see ONNX Model Setup)
- Python 3.10+ (only for the model download script)
Note: This environment targets
net10.0because .NET 10 SDK is installed. The architecture is identical to the planned .NET 9 layout; changeTargetFrameworkin.csprojfiles if you use .NET 9 SDK.
Build and Run
For pushing this repository to Xenovex Git (xts.xenovex.com), see docs/solution-guide.md § Git remote setup.
From the repository root:
dotnet restore
dotnet build
dotnet run --project src/PiiRedaction.ConsoleApp
By default the console app runs 16 curated sample prompts covering English and Tamil/Tanglish/mixed person names, regex identifiers, domain IDs, combined scenarios, and a clean no-PII ticket. No flags are required for Tamil samples — they run in the default batch alongside English.
List available samples:
dotnet run --project src/PiiRedaction.ConsoleApp -- --list
Run a single sample by index or name:
dotnet run --project src/PiiRedaction.ConsoleApp -- --sample 2
dotnet run --project src/PiiRedaction.ConsoleApp -- --name MrTitlePerson
WPF Test Harness
A desktop MVVM application for interactive POC validation with English and Tamil prompts. Requires Windows (net10.0-windows).
Prerequisites: English and Tamil ONNX models downloaded (see ONNX Model Setup).
dotnet run --project src/PiiRedaction.TestHarness.Wpf
Workflow:
- Select a category from the dropdown (e.g. Career Guidance, Banking & Financial) or leave All to see every prompt. Use the search box for finer filtering.
- Click a test prompt in the left panel to load it into the input box (previous results are cleared automatically).
- Click Redact to run the full detection pipeline. The status bar shows model availability, script composition (LatinOnly / TamilOnly / Mixed), and elapsed time.
- Review Sanitized Output, detected entities, and the placeholder map in the right panel. A leak warning appears if any detected value remains in the sanitized text.
- Optionally click Send Mock LLM to send only the sanitized prompt to the mock LLM.
- Click Run All to execute scenarios in the selected category (or all when All is chosen) and view pass/fail results in the batch panel.
The harness uses the same DI registrations and IPromptSanitizer pipeline as the console app, with thin application services (IRedactionAppService, ITestPromptCatalog, IScriptAnalysisService, IModelStatusService) following SOLID principles.
Interactive mode (enter your own prompt):
dotnet run --project src/PiiRedaction.ConsoleApp -- --interactive
Console sample catalog
Samples are defined in SamplePromptCatalog.cs.
| # | Name | Category | NER / Person example |
|---|---|---|---|
| 0 | FullFinancialWithCustomer | NER + Regex + Domain | Customer Ravi Kumar + email, phone, loan, PAN |
| 1 | CustomerNameOnly | NER | Customer Anita Sharma |
| 2 | MrTitlePerson | NER | Mr. John Smith |
| 3 | MrsTitlePerson | NER | Mrs. Lakshmi Reddy |
| 4 | DrTitlePerson | NER | Dr. Jane Doe |
| 5 | TwoCustomersInOnePrompt | NER | Customer Ravi Kumar and Customer Priya Nair |
| 6 | PersonWithDomainIds | NER + Domain | Customer Meera Iyer + CID / ACC |
| 7 | PersonWithEmailNoPhone | NER + Regex | Customer Arjun Mehta + email |
| 8 | AllRegexTypes | Regex | email, phone, PAN, Aadhaar, card |
| 9 | AllDomainIds | Domain | LN, CID, ACC |
| 10 | TamilCustomerNameOnly | NER (Tamil) | வாடிக்கையாளர் ராஜேஷ் குமார் |
| 11 | TamilWithPhonePan | NER (Tamil) + Regex | Tamil person + phone + PAN |
| 12 | TanglishCustomer | NER (English/Tanglish) | Customer Senthil + phone |
| 13 | MixedTamilEnglish | NER (Mixed) | வாடிக்கையாளர் Ravi Kumar + phone |
| 14 | TamilFullFinancial | NER (Tamil) + Regex + Domain | Tamil canonical demo |
| 15 | NoPiiCleanTicket | Negative | no redaction |
Person names are detected via ONNX NER using dslim/bert-base-NER (or a compatible token-classification export). A real model is required for person-name detection; there is no regex or heuristic fallback.
NuGet Packages
# Infrastructure
dotnet add src/PiiRedaction.Infrastructure package Microsoft.ML.OnnxRuntime
dotnet add src/PiiRedaction.Infrastructure package Microsoft.ML.Tokenizers
dotnet add src/PiiRedaction.Infrastructure package Microsoft.Extensions.AI.Abstractions
dotnet add src/PiiRedaction.Infrastructure package Microsoft.Extensions.AI
dotnet add src/PiiRedaction.Infrastructure package Microsoft.Extensions.Logging.Abstractions
dotnet add src/PiiRedaction.Infrastructure package Microsoft.Extensions.Options
# Core
dotnet add src/PiiRedaction.Core package Microsoft.Extensions.Options
# ConsoleApp
dotnet add src/PiiRedaction.ConsoleApp package Microsoft.Extensions.Hosting
dotnet add src/PiiRedaction.ConsoleApp package Microsoft.Extensions.DependencyInjection
dotnet add src/PiiRedaction.ConsoleApp package Microsoft.Extensions.Configuration.Json
dotnet add src/PiiRedaction.ConsoleApp package Microsoft.Extensions.Configuration.EnvironmentVariables
Configuration
{
"PiiRedaction": {
"OnnxModelPath": "models/ner-model.onnx"
}
}
| Setting | Description |
|---|---|
OnnxModelPath |
Path to ONNX NER model (relative to working directory or discovered by walking up from the current directory) |
ONNX Model Setup
Person-name detection requires a token-classification ONNX model and companion tokenizer files in the models/ directory:
| File | Purpose |
|---|---|
models/en/ner-model.onnx |
English BERT NER model (or legacy models/ner-model.onnx) |
models/en/vocab.txt |
BERT WordPiece vocabulary |
models/en/ner-labels.txt |
One BIO label per line (O, B-PER, I-PER, etc.) |
models/ta/model.onnx |
Tamil IndicBERT NER model |
models/ta/sentencepiece.bpe.model |
SentencePiece tokenizer for Tamil model |
models/ta/ner-labels.txt |
Fine-grained Tamil NER labels |
Download scripts
From the repository root:
.\scripts\download-ner-model.ps1
.\scripts\download-tamil-ner-model.ps1
Or with Python directly:
python scripts/download-ner-model.py
python scripts/download-tamil-ner-model.py
The English script exports dslim/bert-base-NER via Hugging Face Optimum when Python is available. The Tamil script exports prachuryyaIITG/SampurNER_Tamil_IndicBERTv2. Otherwise each script downloads pre-exported ONNX assets from Hugging Face directly.
Set EnableTamilNer to false in appsettings.json to revert to English-only routing.
Inference pipeline
RoutingOnnxNerModelRunner classifies script composition and delegates to:
EnglishOnnxNerRunner— BERT WordPiece tokenization for Latin script and TanglishTamilOnnxNerRunner— SentencePiece tokenization for Tamil script (U+0B80–U+0BFF)
Both runners share OnnxTokenClassifierRunner for ONNX Runtime inference and BIO label decoding. Overlapping person spans from mixed-script prompts are merged (longer span wins).
Swapping Mock LLM for Azure OpenAI
The application depends on ILlmPromptService (Core) and IChatClient (Microsoft.Extensions.AI). To use Azure OpenAI later, replace the mock registration in ServiceCollectionExtensions.cs:
// Remove:
// services.AddSingleton<IChatClient, MockChatClient>();
// Add (example — package and API may vary by provider SDK version):
// services.AddAzureOpenAIChatClient(
// new Uri(configuration["AzureOpenAI:Endpoint"]!),
// configuration["AzureOpenAI:ApiKey"]!,
// configuration["AzureOpenAI:DeploymentName"]!);
services.AddSingleton<ILlmPromptService, MockLlmPromptService>(); // unchanged
MockLlmPromptService already uses IChatClient, so it works with any registered chat client implementation.
Sample Execution Output
=== PII Redaction POC ===
Original Prompt:
Customer Ravi Kumar with email ravi.kumar@gmail.com and phone 9876543210 has LoanNumber LN-456789 and PAN ABCDE1234F. Please summarize this customer issue.
Detected PII:
[PERSON ] Ravi Kumar (Ner)
[EMAIL ] ravi.kumar@gmail.com (Regex)
[PHONE ] 9876543210 (Regex)
[LOAN_NUMBER ] LN-456789 (Domain)
[PAN ] ABCDE1234F (Regex)
Sanitized Prompt:
Customer <PERSON_1> with email <EMAIL_1> and phone <PHONE_1> has LoanNumber <LOAN_NUMBER_1> and PAN <PAN_1>. Please summarize this customer issue.
Internal Placeholder Map (not sent to LLM):
<EMAIL_1> -> ravi.kumar@gmail.com
<LOAN_NUMBER_1> -> LN-456789
<PAN_1> -> ABCDE1234F
<PERSON_1> -> Ravi Kumar
<PHONE_1> -> 9876543210
Mock LLM Response:
[Mock LLM Response] Received sanitized prompt (146 chars). No original PII was transmitted.
SOLID Principles Applied
| Principle | Application |
|---|---|
| Single Responsibility | Each detector, redactor, and runner has one job; Program.cs only orchestrates |
| Open/Closed | Add new IPiiDetector implementations without changing merge logic |
| Liskov Substitution | All detectors are interchangeable via IPiiDetector |
| Interface Segregation | Separate interfaces for detection, redaction, sanitization, and LLM |
| Dependency Inversion | Core defines abstractions; Infrastructure implements them |
Testing
The solution includes an NUnit test suite across two projects:
| Project | Focus |
|---|---|
tests/PiiRedaction.Core.Tests |
Detectors, redactor, sanitizer, golden pipeline scenarios (fake NER), real-model integration tests |
tests/PiiRedaction.Infrastructure.Tests |
Mock LLM, ONNX runner unit tests, real-model NER runner tests |
Run tests
dotnet test
dotnet test --filter "FullyQualifiedName~GoldenPromptTests"
dotnet test --filter "Category=RealModel"
dotnet test --filter "Category=TamilNer"
dotnet test --logger "console;verbosity=detailed"
Fast CI runs without the ONNX model: fake-based tests always execute; tests marked Category=RealModel or Category=TamilNer are skipped when the corresponding ONNX models are absent. Download models first:
.\scripts\download-ner-model.ps1
.\scripts\download-tamil-ner-model.ps1
Test architecture
PromptScenarioCatalog— five focused end-to-end English scenarios (canonical demo, multi-regex, duplicate people, overlap stress, no-PII negative)TamilPromptScenarioCatalog— five Tamil/Tanglish/mixed golden scenarios (fake NER for person spans)ProductionPipelineFactory— builds the same Domain → Regex → OnnxNer composite stack as production DI;CreateWithRealModel(runner)wires a real runner;CreateWithRoutingRealModelswires English + Tamil routingFakeOnnxNerModelRunner— unit-test double for NER; golden tests inject person spans per scenarioGoldenPromptTests— end-to-end sanitization proof across the catalog (fake NER)RealNerModelFixture— shared fixture that loadsmodels/ner-model.onnxonce per class; skips when model missingRealNerModelRunnerTests— direct ONNX inference with span accuracy checksRealNerPipelineTests— full pipeline with real English NER (canonical, multi-person, clean-ticket negative)RealTamilPipelineTests— full pipeline with routed English + Tamil NER (Category=TamilNer)RealTamilNerModelRunnerTests— direct Tamil ONNX inference (Category=TamilNer)OnnxNerModelRunnerTests— unit tests for missing/invalid model paths (no download required)CompositePiiDetectorTests— overlap merge and source-priority rulesLlmBoundaryTests— verifies raw PII never appears in outbound LLM messages
Assertions use FluentAssertions for readable failures on long prompt strings.
Future Enhancements
- ASP.NET Core API host with request/response middleware
- Persistent audit log of redaction events (without storing raw PII)
- Secure vault for reversible tokenization
- Real Azure OpenAI / OpenAI provider registration