# PII Redaction POC — Solution Architecture ## Purpose This document describes the architectural design of the **PII Redaction POC**, a .NET proof-of-concept that intercepts user prompts containing regulated personally identifiable information (PII), redacts sensitive values into stable placeholders, and transmits **only sanitized text** across the LLM trust boundary. The solution is structured for enterprise adoption: clear layer separation, interface-driven composition, dependency injection, and swappable infrastructure adapters (ONNX NER, `Microsoft.Extensions.AI` chat clients). The POC validates a compliance-oriented pattern suitable for financial and customer-service workloads where raw PII must not leave the application process when invoking external language models. --- ## Canonical Example The console application ships with a **sample catalog** (16 prompts). The canonical demo is sample `FullFinancialWithCustomer`. Tamil script, Tanglish, and mixed-script samples run in the **default** `dotnet run` batch (no `--interactive` required). The table below shows the exact strings produced by the production pipeline when the ONNX NER models are loaded (run `scripts/download-ner-model.ps1` and `scripts/download-tamil-ner-model.ps1` first). | Stage | Value | |-------|-------| | **Input** | `Customer Ravi Kumar with email ravi.kumar@gmail.com and phone 9876543210 has LoanNumber LN-456789 and PAN ABCDE1234F. Please summarize this customer issue.` | | **Sanitized Output** | `Customer with email and phone has LoanNumber and PAN . Please summarize this customer issue.` | | **Mock LLM Response** | `[Mock LLM Response] Received sanitized prompt (146 chars). No original PII was transmitted.` | Detected entities for this prompt: | Type | Value | Detection Source | |------|-------|------------------| | PERSON | Ravi Kumar | Ner | | EMAIL | ravi.kumar@gmail.com | Regex | | PHONE | 9876543210 | Regex | | LOAN_NUMBER | LN-456789 | Domain | | PAN | ABCDE1234F | Regex | The internal placeholder map (`` → `Ravi Kumar`, etc.) is retained in-process and is **not** included in the outbound LLM request. --- ## Console Sample Catalog Running `dotnet run --project src/PiiRedaction.ConsoleApp` executes all samples sequentially. Use `--list`, `--sample N`, or `--name SampleName` to filter. ### NER / person-name samples These prompts exercise `OnnxNerPiiDetector` and `RoutingOnnxNerModelRunner`. Person names require ONNX models (`models/en/` for English, `models/ta/` for Tamil script). Without models, person spans are not detected. Legacy `models/ner-model.onnx` is still supported for English. | Sample | Input (excerpt) | Detected person | Sanitized (excerpt) | |--------|-----------------|-----------------|---------------------| | **CustomerNameOnly** | Customer Anita Sharma reported unauthorized… | Anita Sharma | Customer `` reported unauthorized… | | **MrTitlePerson** | Mr. John Smith called about a duplicate debit… | John Smith | `` called about a duplicate debit… | | **MrsTitlePerson** | Mrs. Lakshmi Reddy requested a callback regarding LN-112233. | Lakshmi Reddy | `` requested a callback regarding ``. | | **DrTitlePerson** | Dr. Jane Doe escalated a complaint… | Jane Doe | `` escalated a complaint… | | **TwoCustomersInOnePrompt** | Customer Ravi Kumar and Customer Priya Nair… | Ravi Kumar, Priya Nair | Customer `` and Customer ``… | | **PersonWithDomainIds** | Customer Meera Iyer holds CID-7070… | Meera Iyer | Customer `` holds ``… | | **PersonWithEmailNoPhone** | Customer Arjun Mehta wrote from arjun.mehta@company.in… | Arjun Mehta | Customer `` wrote from ``… | ### Tamil / Tanglish / mixed samples These prompts exercise `RoutingOnnxNerModelRunner` script routing. Tamil script uses `models/ta/`; Latin Tanglish uses `models/en/`. Mixed prompts may invoke both models. | Sample | Input (excerpt) | Detected person | Sanitized (excerpt) | |--------|-----------------|-----------------|---------------------| | **TamilCustomerNameOnly** | வாடிக்கையாளர் ராஜேஷ் குமார் சேமிப்பு… | ராஜேஷ் குமார் | வாடிக்கையாளர் `` சேமிப்பு… | | **TamilWithPhonePan** | வாடிக்கையாளர் ராஜேஷ் குமார் தொலைபேசி 9876543210 PAN… | ராஜேஷ் குமார் | `` … `` … `` | | **TanglishCustomer** | Customer Senthil phone 9876543210… | Senthil | Customer `` phone ``… | | **MixedTamilEnglish** | வாடிக்கையாளர் Ravi Kumar phone 9876543210… | Ravi Kumar | வாடிக்கையாளர் `` phone ``… | | **TamilFullFinancial** | வாடிக்கையாளர் ராஜேஷ் குமார் மின்னஞ்சல் ravi.kumar@gmail.com… | ராஜேஷ் குமார் | Tamil canonical — all placeholder types | ### Other sample categories | Category | Sample | Purpose | |----------|--------|---------| | NER + Regex + Domain | FullFinancialWithCustomer | End-to-end financial prompt (canonical) | | Regex only | AllRegexTypes | Email, phone, PAN, Aadhaar, credit card | | Domain only | AllDomainIds | Loan number, customer ID, account number | | Negative | NoPiiCleanTicket | Passthrough with no detected PII | Sample definitions live in [`SamplePromptCatalog.cs`](../src/PiiRedaction.ConsoleApp/Samples/SamplePromptCatalog.cs). --- ## High-Level Data Flow The diagram below traces the canonical example from console input through Core sanitization to the Infrastructure LLM adapter. Data labels reflect the canonical strings at each stage. ```mermaid flowchart TB subgraph consoleApp [PiiRedaction.ConsoleApp] program["Program.cs"] di["ServiceRegistration"] end subgraph core [PiiRedaction.Core] sanitizer["PromptSanitizer"] composite["CompositePiiDetector"] regexDet["RegexPiiDetector"] domainDet["DomainRulePiiDetector"] onnxDet["OnnxNerPiiDetector"] redactor["PlaceholderPiiRedactor"] end subgraph infra [PiiRedaction.Infrastructure] onnxRunner["RoutingOnnxNerModelRunner"] enRunner["EnglishOnnxNerRunner"] taRunner["TamilOnnxNerRunner"] mockLlm["MockLlmPromptService"] mockChat["MockChatClient"] end rawPrompt["Raw prompt with PII"] sanitizedPrompt["Sanitized prompt with placeholders"] llmResponse["Mock LLM acknowledgment"] program -->|"Customer Ravi Kumar ... PAN ABCDE1234F"| sanitizer sanitizer --> composite composite --> domainDet composite --> regexDet composite --> onnxDet onnxDet --> onnxRunner sanitizer --> redactor redactor -->|"Customer PERSON_1 ... PAN PAN_1"| sanitizedPrompt program -->|"SanitizedPrompt only"| mockLlm mockLlm --> mockChat mockChat --> llmResponse rawPrompt -.-> program di -.-> sanitizer di -.-> mockLlm ``` `RoutingOnnxNerModelRunner` selects English and/or Tamil ONNX models based on script composition in the prompt. See [Dual-Model NER Routing (Tamil + English)](#dual-model-ner-routing-tamil--english) for the routing decision tree. --- ## Detection to Redaction Detail `PromptSanitizer` orchestrates a two-phase pipeline: **detect** then **redact**. `CompositePiiDetector` aggregates spans from all registered detectors, resolves overlaps by registration order and source priority, and returns a merged entity list. `PlaceholderPiiRedactor` replaces spans right-to-left to preserve indices, assigns stable per-type counters, and builds the in-process placeholder map. ```mermaid flowchart LR inputText["Original prompt text"] subgraph detectPhase [Detection Phase] domainDet["DomainRulePiiDetector"] regexDet["RegexPiiDetector"] onnxDet["OnnxNerPiiDetector
(RoutingOnnxNerModelRunner)"] composite["CompositePiiDetector"] merge["Overlap merge and source priority"] entityList["PiiEntity list"] end subgraph redactPhase [Redaction Phase] redactor["PlaceholderPiiRedactor"] replace["Right-to-left span replacement"] placeholderMap["Placeholder map in-process"] sanitizedText["Sanitized text"] end inputText --> domainDet inputText --> regexDet inputText --> onnxDet domainDet --> composite regexDet --> composite onnxDet --> composite composite --> merge merge --> entityList entityList --> redactor inputText --> redactor redactor --> replace replace --> sanitizedText replace --> placeholderMap ``` **Overlap resolution rules** (applied by `CompositePiiDetector`): 1. Detectors run in registration order: **Domain → Regex → ONNX NER**. 2. On overlapping spans, the first registered detector wins. 3. Tie-breaking uses source priority: Domain (3) > Regex (2) > NER (1). The ONNX NER detector delegates to `RoutingOnnxNerModelRunner`, which routes inference to English and/or Tamil models by script composition. See [Dual-Model NER Routing (Tamil + English)](#dual-model-ner-routing-tamil--english). **Placeholder assignment** (applied by `PlaceholderPiiRedactor`): - Format: `<{TYPE}_{n}>` (e.g. ``, ``). - Duplicate values of the same type reuse the same placeholder. - Replacement proceeds from highest `StartIndex` to lowest to avoid index drift. --- ## Dual-Model NER Routing (Tamil + English) Person-name detection uses two ONNX token-classifier models: **English** (`models/en/`, BERT WordPiece) and **Tamil** (`models/ta/`, SentencePiece or WordPiece). `OnnxNerPiiDetector` calls `RoutingOnnxNerModelRunner`, which classifies prompt script via `ScriptRouter` and dispatches to `EnglishOnnxNerRunner` and/or `TamilOnnxNerRunner`. Both runners share `OnnxTokenClassifierRunner` for BIO decoding; only **PERSON** spans are emitted. The diagram below expands the detection and NER branches summarized in [High-Level Data Flow](#high-level-data-flow) and [Detection to Redaction Detail](#detection-to-redaction-detail). ### End-to-end pipeline (with NER branch) ```mermaid flowchart TB subgraph Entry["Console entry"] A["Program.cs
Host + AddPiiRedactionServices()"] B["PromptDemoRunner.RunAsync()"] A --> B end B --> C["SanitizationRequest(OriginalPrompt)"] C --> D["PromptSanitizer.Sanitize()"] subgraph Detect["CompositePiiDetector.Detect() — registration order"] direction TB E1["DomainRulePiiDetector
LOAN_NUMBER, CUSTOMER_ID, ACCOUNT_NUMBER"] E2["RegexPiiDetector
EMAIL, PHONE, AADHAAR, PAN, CREDIT_CARD"] E3["OnnxNerPiiDetector
PERSON (via IOnnxNerModelRunner)"] E1 --> MERGE E2 --> MERGE E3 --> MERGE MERGE["Merge overlapping spans
sort: StartIndex ↑, Length ↓, Source priority ↓
(Domain=3, Regex=2, Ner=1)
first candidate wins on overlap"] end D --> Detect MERGE --> F["IReadOnlyList<PiiEntity>"] F --> G["PlaceholderPiiRedactor.Redact()
replace spans right-to-left
dedupe by Type|Value → <TYPE_n>"] G --> H["SanitizationResult
SanitizedPrompt, DetectedEntities, PlaceholderMap"] H --> I["MockLlmPromptService.SendPromptAsync(SanitizedPrompt)"] I --> J["Mock LLM response
(sanitized text only)"] subgraph NerBranch["OnnxNerPiiDetector branch"] E3 --> N1{"RoutingOnnxNerModelRunner
.IsModelAvailable?"} N1 -->|no| N2["return []"] N1 -->|yes| N3["RoutingOnnxNerModelRunner
.PredictEntities()"] end ``` ### RoutingOnnxNerModelRunner decision tree `ScriptRouter.GetComposition` scans each character once. Tamil letters (U+0B80–U+0BFF) and ASCII Latin letters (`char.IsAsciiLetter`) determine the route. When both scripts appear, classification is **Mixed** (early exit). ```mermaid flowchart TB IN["text"] --> SR["ScriptRouter.GetComposition(text)
scan each char"] SR --> C1{"LatinOnly?"} SR --> C2{"TamilOnly?"} SR --> C3{"Mixed?"} SR --> C4{"NoLetters?"} C1 -->|yes| EN1{"EnglishOnnxNerRunner
.IsModelAvailable?"} EN1 -->|yes| EN_RUN["EnglishOnnxNerRunner.PredictEntities(text)"] EN1 -->|no| SKIP1["skip English"] EN_RUN --> ACC SKIP1 --> ACC C2 -->|yes| TA_GATE{"EnableTamilNer
&& TamilOnnxNerRunner
.IsModelAvailable?"} TA_GATE -->|yes| TA_RUN["TamilOnnxNerRunner.PredictEntities(text)"] TA_GATE -->|no| SKIP2["skip Tamil"] TA_RUN --> ACC SKIP2 --> ACC C3 -->|yes| EN2{"English available?"} EN2 -->|yes| EN_MIX["EnglishOnnxNerRunner.PredictEntities(text)"] EN2 -->|no| SKIP3["skip English"] EN_MIX --> TA_GATE2{"EnableTamilNer
&& Tamil available?"} SKIP3 --> TA_GATE2 TA_GATE2 -->|yes| TA_MIX["TamilOnnxNerRunner.PredictEntities(text)"] TA_GATE2 -->|no| SKIP4["skip Tamil"] TA_MIX --> ACC SKIP4 --> ACC C4 -->|yes| EMPTY["no NER inference"] EMPTY --> OUT_EMPTY["return []"] subgraph EN_Pipeline["EnglishOnnxNerRunner"] EN_RUN --> EN_ENC["BertWordPieceEncoder
(model dir vocab.txt)"] EN_ENC --> EN_OCR["OnnxTokenClassifierRunner
NerLabelConfig.English
B-PER / I-PER / B-PERSON / I-PERSON"] end subgraph TA_Pipeline["TamilOnnxNerRunner"] TA_RUN --> TA_ENC["TokenClassifierEncoderFactory.Create()
vocab.txt → BertWordPieceEncoder
else SentencePiece (*.bpe.model, spiece.model, tokenizer.model)"] TA_ENC --> TA_OCR["OnnxTokenClassifierRunner
NerLabelConfig.Tamil
label contains 'person' (case-insensitive)"] end subgraph SharedInference["OnnxTokenClassifierRunner (shared)"] ENC["Encode(text, max 128 tokens)"] ONNX["ONNX InferenceSession.Run
input_ids + attention_mask [+ token_type_ids]"] ARGMAX["Per-token argmax over logits"] BIO["BIO decode → PiiEntityType.Person
PiiDetectionSource.Ner"] ENC --> ONNX --> ARGMAX --> BIO end EN_OCR --> SharedInference TA_OCR --> SharedInference BIO --> ACC["accumulate entities"] ACC --> MERGE["MergePersonSpans()
sort: Length ↓, StartIndex ↑
drop overlapping spans
(longer span wins)"] MERGE --> OUT["return merged PERSON entities"] ``` ### Routing rules | Rule | Source | Behavior | |------|--------|----------| | **Script classification** | `ScriptRouter.GetComposition` | Single pass over characters. Tamil letter = U+0B80–U+0BFF. Latin letter = `char.IsAsciiLetter`. Both seen → `Mixed` (early exit). Neither → `NoLetters`. Tamil only → `TamilOnly`. Latin only → `LatinOnly`. | | **LatinOnly** | `RoutingOnnxNerModelRunner` | Run **English only** if `englishRunner.IsModelAvailable`. | | **TamilOnly** | `RoutingOnnxNerModelRunner` | Run **Tamil only** if `EnableTamilNer` (default `true` in `PiiRedactionOptions`) **and** `tamilRunner.IsModelAvailable`. | | **Mixed** | `RoutingOnnxNerModelRunner` | Run **both** models independently on the **full text** (English if available; Tamil if `EnableTamilNer` and available). | | **NoLetters** | `RoutingOnnxNerModelRunner` | No NER inference; returns `[]` from routing (before merge). | | **Model availability gate** | `OnnxNerPiiDetector` | If `RoutingOnnxNerModelRunner.IsModelAvailable` is false, NER detector returns `[]` (English OR Tamil available when Tamil enabled). | | **Post-route merge** | `MergePersonSpans` | After EN/TA results are concatenated, overlapping PERSON spans are deduped; **longer span wins**, then ordered by `StartIndex`. | | **Composite merge** | `CompositePiiDetector` | Domain → Regex → NER all run. Overlaps resolved globally: earlier registration order + longer span + higher source priority (Domain > Regex > Ner). | | **Encoder choice** | `EnglishOnnxNerRunner` vs `TamilOnnxNerRunner` | English always uses `BertWordPieceEncoder`. Tamil uses factory: `vocab.txt` → WordPiece; else first SentencePiece file found; fallback WordPiece with warning. | | **NER output scope** | `OnnxTokenClassifierRunner` | Only **PERSON** entities decoded from BIO tags; max sequence length **128** tokens. | --- ## Runtime Sequence ```mermaid sequenceDiagram participant User participant Program as Program.cs participant DI as ServiceProvider participant Sanitizer as PromptSanitizer participant Detector as CompositePiiDetector participant Redactor as PlaceholderPiiRedactor participant LlmSvc as MockLlmPromptService participant Chat as MockChatClient User->>Program: Start application Program->>DI: Resolve IPromptSanitizer, ILlmPromptService DI-->>Program: Sanitizer, LlmService alt Interactive mode User->>Program: Enter prompt via console else Default mode Program->>Program: Load canonical sample prompt end Program->>Sanitizer: Sanitize(SanitizationRequest) Sanitizer->>Detector: Detect(originalPrompt) Detector-->>Sanitizer: IReadOnlyList PiiEntity Sanitizer->>Redactor: Redact(originalPrompt, entities) Redactor-->>Sanitizer: RedactionResult Sanitizer-->>Program: SanitizationResult Program->>Program: Display detected entities Program->>Program: Display sanitized prompt Program->>Program: Display placeholder map in-process Program->>LlmSvc: SendPromptAsync(sanitizedPrompt) Note over Program,LlmSvc: Placeholder map never passed LlmSvc->>Chat: GetResponseAsync(user message) Chat-->>LlmSvc: Assistant response LlmSvc-->>Program: Mock LLM response string Program->>User: Write LLM response to console ``` --- ## Trust Boundary The LLM boundary is the point at which data leaves the application process via `ILlmPromptService` / `IChatClient`. Only the sanitized prompt crosses this boundary. Original PII values, detection metadata, and the placeholder-to-value map remain in-process. ```mermaid flowchart TB subgraph inProcess [In-Process Trust Zone] originalPrompt["Original prompt with raw PII"] detectedEntities["Detected PiiEntity list"] placeholderMap["Placeholder map"] sanitizationResult["SanitizationResult"] consoleDisplay["Console audit output"] end subgraph llmBoundary [LLM Trust Boundary] sanitizedOnly["Sanitized prompt text only"] end subgraph externalLlm [External LLM Provider] chatClient["IChatClient implementation"] modelInference["Model inference"] end originalPrompt --> sanitizationResult detectedEntities --> sanitizationResult placeholderMap --> sanitizationResult sanitizationResult --> consoleDisplay sanitizationResult -->|"SendPromptAsync"| sanitizedOnly sanitizedOnly --> chatClient chatClient --> modelInference originalPrompt -.-x|"Never transmitted"| chatClient placeholderMap -.-x|"Never transmitted"| chatClient detectedEntities -.-x|"Never transmitted"| chatClient ``` In the POC, `MockChatClient` simulates the external provider without network I/O. Replacing it with Azure OpenAI or another `IChatClient` implementation does not change the trust model: `MockLlmPromptService` (or a future production adapter) continues to accept only the sanitized string. --- ## Project Responsibilities | Project | Layer | Responsibility | |---------|-------|----------------| | `PiiRedaction.ConsoleApp` | Presentation | Application entry point; reads prompt (sample or interactive); bootstraps `IHost` and DI via `AddPiiRedactionServices`; orchestrates sanitization and LLM invocation; renders audit output (detected entities, sanitized text, placeholder map). | | `PiiRedaction.Core` | Domain / Application | Defines abstractions (`IPiiDetector`, `IPiiRedactor`, `IPromptSanitizer`, `ILlmPromptService`); implements detection strategies (`RegexPiiDetector`, `DomainRulePiiDetector`, `OnnxNerPiiDetector`, `CompositePiiDetector`); implements `PlaceholderPiiRedactor` and `PromptSanitizer`; owns domain models (`PiiEntity`, `SanitizationResult`, `RedactionResult`) and configuration (`PiiRedactionOptions`). Has no dependency on ONNX Runtime or LLM SDKs. | | `PiiRedaction.Infrastructure` | Infrastructure | Implements technical adapters: `RoutingOnnxNerModelRunner`, `EnglishOnnxNerRunner`, `TamilOnnxNerRunner` (ONNX Runtime inference), `MockChatClient` and `MockLlmPromptService` (`Microsoft.Extensions.AI`); depends on Core abstractions and is swappable without changing domain logic. | | `tests/PiiRedaction.Core.Tests` | Test | Unit and integration tests for detectors, redactor, sanitizer, overlap rules, golden prompt scenarios (`PromptScenarioCatalog`), and LLM boundary assertions. | | `tests/PiiRedaction.Infrastructure.Tests` | Test | Tests for mock LLM behavior and ONNX runner load semantics. | **Dependency direction:** `ConsoleApp` → `Infrastructure` → `Core`. Core references no outer layers, preserving the Dependency Inversion Principle and enabling future hosts (ASP.NET Core API, worker services) to reuse the same Core and Infrastructure assemblies. --- ## Key Abstractions and Extension Points | Abstraction | Defined In | Default Implementation | Extension | |-------------|------------|------------------------|-----------| | `IPiiDetector` | Core | `CompositePiiDetector` wrapping Domain, Regex, ONNX | Add new detector; register in composite order | | `IPiiRedactor` | Core | `PlaceholderPiiRedactor` | Replace with hashing, vault-backed tokens, etc. | | `IPromptSanitizer` | Core | `PromptSanitizer` | Unlikely to change; orchestrates detect + redact | | `ILlmPromptService` | Core | `MockLlmPromptService` | Production adapter with telemetry, retry, policy | | `IChatClient` | Microsoft.Extensions.AI | `MockChatClient` | Azure OpenAI, OpenAI, or other provider SDK | | `IOnnxNerModelRunner` | Core | `RoutingOnnxNerModelRunner` | Script-based routing to English (BERT WordPiece) and Tamil (SentencePiece) ONNX models | --- ## Configuration Surface Runtime behavior is controlled via `appsettings.json` under the `PiiRedaction` section: | Setting | Effect | |---------|--------| | `OnnxModelPath` | Legacy English model path (`models/ner-model.onnx`). Used as fallback when `models/en/` is absent. | | `EnglishOnnxModelPath` | Primary English ONNX model (`models/en/ner-model.onnx`). | | `TamilOnnxModelPath` | Tamil ONNX model (`models/ta/model.onnx`). | | `EnableTamilNer` | When `false`, routing uses English model only. Default `true`. | Download model assets with `scripts/download-ner-model.ps1` and `scripts/download-tamil-ner-model.ps1`. --- ## Related Documentation - [NER models](ner-models.md) — English/Tamil ONNX model IDs, assets, labels, and integration reference - [README](../README.md) — build, run, configuration, and testing instructions - [ServiceCollectionExtensions.cs](../src/PiiRedaction.ConsoleApp/DependencyInjection/ServiceCollectionExtensions.cs) — DI registration and detector ordering - [PromptScenarioCatalog.cs](../tests/PiiRedaction.Core.Tests/TestSupport/PromptScenarioCatalog.cs) — focused golden pipeline scenarios including the canonical example