# PII Redaction POC — Solution Guide Single reference for architecture, NER models, routing, operations, Tamil/Tanglish support, Git setup, and the post-POC improvement backlog. --- ## Table of contents 1. [Purpose](#1-purpose) 2. [Canonical example](#2-canonical-example) 3. [Detection strategies](#3-detection-strategies) 4. [Solution architecture](#4-solution-architecture) 5. [Trust boundary](#5-trust-boundary) 6. [NER models](#6-ner-models) 7. [NER routing — English vs Tamil](#7-ner-routing--english-vs-tamil) 8. [Configuration and dependency injection](#8-configuration-and-dependency-injection) 9. [Running and testing](#9-running-and-testing) 10. [Tamil and Tanglish support](#10-tamil-and-tanglish-support) 11. [Git remote setup (Xenovex)](#11-git-remote-setup-xenovex) 12. [Improvement roadmap](#12-improvement-roadmap) 13. [Key source files](#13-key-source-files) --- ## 1. Purpose This .NET proof-of-concept intercepts user prompts containing regulated personally identifiable information (PII), redacts sensitive values into stable placeholders, and transmits **only sanitized text** across the LLM trust boundary. The solution uses: - Clear layer separation (Core / Infrastructure / hosts) - Interface-driven composition and dependency injection - Swappable ONNX NER adapters and `Microsoft.Extensions.AI` chat clients It targets financial and customer-service workloads where raw PII must not leave the application process when invoking external language models. --- ## 2. Canonical example With English and Tamil ONNX models loaded (`scripts/download-ner-model.ps1`, `scripts/download-tamil-ner-model.ps1`): | Stage | Value | |-------|-------| | **Input** | `Customer Ravi Kumar with email ravi.kumar@gmail.com and phone 9876543210 has LoanNumber LN-456789 and PAN ABCDE1234F. Please summarize this customer issue.` | | **Sanitized output** | `Customer with email and phone has LoanNumber and PAN . Please summarize this customer issue.` | | **Mock LLM response** | `[Mock LLM Response] Received sanitized prompt (146 chars). No original PII was transmitted.` | | Type | Value | Source | |------|-------|--------| | PERSON | Ravi Kumar | NER | | EMAIL | ravi.kumar@gmail.com | Regex | | PHONE | 9876543210 | Regex | | LOAN_NUMBER | LN-456789 | Domain | | PAN | ABCDE1234F | Regex | The placeholder map (`` → `Ravi Kumar`, etc.) stays **in-process** and is never sent to the LLM. --- ## 3. Detection strategies | Strategy | Detects | Rationale | |----------|---------|-----------| | **Regex** | Email, phone, PAN, Aadhaar, credit card | Deterministic, format-bound, auditable | | **ONNX NER** | Person names | Contextual; no rigid format | | **Domain rules** | Loan number (`LN-`), customer ID (`CID-`), account (`ACC-`) | Business-specific identifiers | **Overlap resolution** (`CompositePiiDetector`): detectors run in order **Domain → Regex → NER**. On overlapping spans, candidates are sorted by start index, length, and source priority (**Domain 3 > Regex 2 > NER 1**); the first non-overlapping candidate wins. **Redaction** (`PlaceholderPiiRedactor`): format `<{TYPE}_{n}>`; duplicate values reuse placeholders; replacement is right-to-left to preserve indices. --- ## 4. Solution architecture ### Project structure ``` src/ ├── PiiRedaction.ConsoleApp/ # Console demo, DI bootstrap ├── PiiRedaction.TestHarness.Wpf/ # WPF MVVM manual test harness ├── PiiRedaction.Core/ # Detection, redaction, abstractions └── PiiRedaction.Infrastructure/ # ONNX runners, mock LLM models/ # ONNX assets (gitignored) tests/ # NUnit unit and integration tests ``` | Project | Layer | Responsibility | |---------|-------|----------------| | `PiiRedaction.ConsoleApp` | Presentation | Samples, interactive mode, audit output | | `PiiRedaction.TestHarness.Wpf` | Presentation | Category-filtered prompts, redact UI, batch runner | | `PiiRedaction.Core` | Domain | `IPiiDetector`, `IPromptSanitizer`, detectors, models | | `PiiRedaction.Infrastructure` | Infrastructure | `RoutingOnnxNerModelRunner`, `MockLlmPromptService` | | `tests/*` | Test | ~112 tests; `RealModel` and `TamilNer` categories | **Dependency direction:** `ConsoleApp` / `Wpf` → `Infrastructure` → `Core`. Core has no ONNX or LLM SDK references. ### Key abstractions | Abstraction | Default implementation | Extension | |-------------|------------------------|-----------| | `IPiiDetector` | `CompositePiiDetector` | Add detector; register in composite order | | `IPiiRedactor` | `PlaceholderPiiRedactor` | Hashing, vault tokens | | `IPromptSanitizer` | `PromptSanitizer` | Orchestrates detect + redact | | `IOnnxNerModelRunner` | `RoutingOnnxNerModelRunner` | Script-based EN/TA routing | | `ILlmPromptService` | `MockLlmPromptService` | Production adapter | | `IChatClient` | `MockChatClient` | Azure OpenAI, etc. | ### Data flow ```mermaid flowchart TB subgraph hosts [Hosts] console[ConsoleApp / WpfHarness] end subgraph core [PiiRedaction.Core] sanitizer[PromptSanitizer] composite[CompositePiiDetector] domain[DomainRulePiiDetector] regex[RegexPiiDetector] onnxDet[OnnxNerPiiDetector] redactor[PlaceholderPiiRedactor] end subgraph infra [PiiRedaction.Infrastructure] router[RoutingOnnxNerModelRunner] en[EnglishOnnxNerRunner] ta[TamilOnnxNerRunner] llm[MockLlmPromptService] end console --> sanitizer sanitizer --> composite composite --> domain composite --> regex composite --> onnxDet onnxDet --> router router --> en router --> ta sanitizer --> redactor console -->|"sanitized text only"| llm ``` **Pipeline:** `Sanitize` → `Detect` (all detectors) → `Redact` → `SanitizationResult`. Optional: `SendPromptAsync(sanitizedPrompt)` to LLM. --- ## 5. Trust boundary Only the **sanitized prompt string** crosses `ILlmPromptService` / `IChatClient`. Original PII, entity metadata, and the placeholder map remain in-process. ```mermaid flowchart LR subgraph inProcess [In-Process] raw[Original prompt] entities[Detected entities] map[Placeholder map] audit[Console / WPF display] end subgraph boundary [LLM boundary] sanitized[Sanitized prompt only] end subgraph external [External LLM] chat[IChatClient] end raw --> audit entities --> audit map --> audit sanitized --> chat raw -.->|never sent| chat map -.->|never sent| chat ``` --- ## 6. NER models Person-name detection is the **only** NER responsibility. Structured PII uses regex and domain rules. ### Routing summary | Script in prompt | Model(s) | Typical use | |------------------|----------|-------------| | `LatinOnly` | English | English names, Indian names in Roman script, **Tanglish** | | `TamilOnly` | Tamil | Tamil-script names | | `Mixed` | Both; merge spans | Code-mixed prompts | | `NoLetters` | Neither | Digits/symbols only | ### English model | Property | Value | |----------|-------| | Hugging Face ID | [`dslim/bert-base-NER`](https://huggingface.co/dslim/bert-base-NER) | | Tokenizer | WordPiece (`vocab.txt`, `BertWordPieceEncoder`) | | Person labels | `B-PER`, `I-PER`, `B-PERSON`, `I-PERSON` | | Primary path | `models/en/ner-model.onnx` | | Legacy fallback | `models/ner-model.onnx` | | Download | `.\scripts\download-ner-model.ps1` | ### Tamil model | Property | Value | |----------|-------| | Hugging Face ID | [`prachuryyaIITG/SampurNER_Tamil_IndicBERTv2`](https://huggingface.co/prachuryyaIITG/SampurNER_Tamil_IndicBERTv2) | | Tokenizer | WordPiece when `vocab.txt` present (export path); SentencePiece fallback | | Person labels | Any BIO tag containing `person` (case-insensitive) | | Path | `models/ta/model.onnx` | | Download | `.\scripts\download-tamil-ner-model.ps1` | **Why SampurNER over MuRIL:** Tamil-specific NER training, smaller footprint (~0.3B vs ~0.6B), validated ONNX export in this repo. MuRIL reserved as eval-driven fallback. ### Model assets (gitignored) | Directory | Key files | Approx. size | |-----------|-----------|--------------| | `models/en/` | `ner-model.onnx`, `vocab.txt`, `ner-labels.txt` | ~431 MB ONNX | | `models/ta/` | `model.onnx`, `vocab.txt`, `ner-labels.txt` | ~1 GB ONNX | Only `models/en/.gitkeep` and `models/ta/.gitkeep` are committed. ### Shared inference Both runners use `OnnxTokenClassifierRunner`: - Max sequence length: **128 tokens** (long prompts truncate silently) - BIO decode → `PiiEntityType.Person` only - Missing model → fail-open: `[]` from NER (person names not redacted) --- ## 7. NER routing — English vs Tamil ### Call chain ``` PromptSanitizer → CompositePiiDetector → OnnxNerPiiDetector → RoutingOnnxNerModelRunner.PredictEntities() → ScriptRouter.GetComposition(text) → switch (composition) { English / Tamil / both } → MergePersonSpans() ``` ### Step 1: `ScriptRouter` (Core) **File:** `src/PiiRedaction.Core/Detection/ScriptRouter.cs` Single pass over characters: - Tamil letter: Unicode **U+0B80 – U+0BFF** - Latin letter: `char.IsAsciiLetter` - Both seen → `Mixed` (early exit) - Neither → `NoLetters` - Otherwise → `TamilOnly` or `LatinOnly` ```csharp // ScriptComposition enum: LatinOnly, TamilOnly, Mixed, NoLetters public ScriptComposition GetComposition(string text) { /* scan chars */ } ``` **Tanglish** in Roman script → `LatinOnly` → **English model only**. ### Step 2: `RoutingOnnxNerModelRunner` (Infrastructure) **File:** `src/PiiRedaction.Infrastructure/Onnx/RoutingOnnxNerModelRunner.cs` This is the **branch decision**: | `ScriptComposition` | English | Tamil | |---------------------|---------|-------| | `LatinOnly` | Yes if available | No | | `TamilOnly` | No | Yes if `EnableTamilNer` and available | | `Mixed` | Yes if available | Yes if `EnableTamilNer` and available | | `NoLetters` | No | No | Both models run on the **full prompt text** for `Mixed` (no script segmentation). ### Step 3: Merge `MergePersonSpans`: overlapping PERSON spans → **longer span wins**; ordered by `StartIndex`. ### Routing diagram ```mermaid flowchart TD text[Prompt text] --> sr[ScriptRouter] sr --> latin[LatinOnly] sr --> tamil[TamilOnly] sr --> mixed[Mixed] sr --> none[NoLetters] latin --> en[EnglishOnnxNerRunner] tamil --> ta[TamilOnnxNerRunner] mixed --> en mixed --> ta en --> merge[MergePersonSpans] ta --> merge none --> empty[No NER] ``` ### Worked examples | Input style | Composition | Models | |-------------|---------------|--------| | `Customer Ravi Kumar…` | `LatinOnly` | English | | `வாடிக்கையாளர் ராஜேஷ் குமார்…` | `TamilOnly` | Tamil | | `Naan Suresh, phone 9003789456…` | `LatinOnly` | English (Tanglish) | | `வாடிக்கையாளர் Ravi Kumar phone…` | `Mixed` | Both | | `Callback on 9123456780…` | `NoLetters` | Neither (phone via Regex) | --- ## 8. Configuration and dependency injection ### `appsettings.json` ```json { "PiiRedaction": { "OnnxModelPath": "models/ner-model.onnx", "EnglishOnnxModelPath": "models/en/ner-model.onnx", "TamilOnnxModelPath": "models/ta/model.onnx", "EnableTamilNer": true } } ``` | Setting | Effect | |---------|--------| | `EnglishOnnxModelPath` | Primary English model | | `OnnxModelPath` | Legacy English fallback | | `TamilOnnxModelPath` | Tamil model | | `EnableTamilNer` | `false` = English-only routing | ### DI registration ```csharp services.AddSingleton(); services.AddSingleton(); services.AddSingleton(); // OnnxNerPiiDetector receives IOnnxNerModelRunner (the router) ``` Same pattern in `PiiRedaction.ConsoleApp` and `PiiRedaction.TestHarness.Wpf` `ServiceCollectionExtensions.cs`. Detector registration order in composite: **Domain → Regex → ONNX NER**. --- ## 9. Running and testing ### Build ```bash dotnet restore dotnet build ``` ### Console app ```bash # All 16 samples (English + Tamil/Tanglish/mixed) dotnet run --project src/PiiRedaction.ConsoleApp dotnet run --project src/PiiRedaction.ConsoleApp -- --list dotnet run --project src/PiiRedaction.ConsoleApp -- --name TamilCustomerNameOnly dotnet run --project src/PiiRedaction.ConsoleApp -- --interactive ``` Samples: `src/PiiRedaction.ConsoleApp/Samples/SamplePromptCatalog.cs` ### WPF test harness (Windows) ```bash dotnet run --project src/PiiRedaction.TestHarness.Wpf ``` - **Category dropdown:** Career Guidance, Banking & Financial, Negative, Edge & Harness - Click prompt → loads input; **Redact** runs pipeline - **Run All** executes scenarios in the selected category - Script badge shows predicted routing (`LatinOnly`, `TamilOnly`, `Mixed`) ### Download models ```powershell .\scripts\download-ner-model.ps1 .\scripts\download-tamil-ner-model.ps1 ``` If Tamil PowerShell download 404s, use Python 3.12+: ```powershell .\scripts\download-tamil-ner-model.ps1 -Python "$env:LOCALAPPDATA\Programs\Python\Python312\python.exe" ``` ### Tests ```bash dotnet test dotnet test --filter "Category=RealModel" dotnet test --filter "Category=TamilNer" dotnet test --filter "FullyQualifiedName~ScriptRouterTests|FullyQualifiedName~RoutingOnnxNerModelRunnerTests" ``` Tests skip gracefully when ONNX files are absent. ### Console sample categories | Category | Examples | |----------|----------| | NER (English) | `CustomerNameOnly`, `MrTitlePerson`, `TwoCustomersInOnePrompt` | | NER (Tamil) | `TamilCustomerNameOnly`, `TamilFullFinancial` | | Tanglish / Mixed | `TanglishCustomer`, `MixedTamilEnglish` | | Regex / Domain | `AllRegexTypes`, `AllDomainIds` | | Negative | `NoPiiCleanTicket` | --- ## 10. Tamil and Tanglish support ### Implementation status | Phase | Status | Scope | |-------|--------|-------| | **1** Generic ONNX token classifier | Done | `OnnxTokenClassifierRunner`, encoders, `NerLabelConfig` | | **2** Tamil download + config | Done | Scripts, dual paths, `EnableTamilNer` | | **3** Script routing + DI | Done | `ScriptRouter`, `RoutingOnnxNerModelRunner` | | **4** Tests, samples, docs | Partial | Tests/samples done; eval metrics open | | **5** Optional enhancements | Not started | See below | ### Remaining gaps (Phase 4–5) | Gap | Impact | |-----|--------| | Tamil numeral normalization (௦–௯ → 0–9) | May miss phone/Aadhaar in Tamil script | | Tanglish heuristics (`peru`, `enga peru`) | Better Latin-name recall in Tamil context | | Label-aware regex cues | Contextual name detection | | Fail-closed when NER missing | Compliance hardening | | MuRIL model swap | If Tamil recall insufficient | ### Tanglish expectations | Input | Handler | Expected recall | |-------|---------|-----------------| | Tamil script names | Tamil ONNX | High (with tuning) | | Latin Indian names (`Ravi Kumar`) | English ONNX | High | | Tanglish spellings (`Senthil`) | English NER + optional heuristics | Medium | | Code-mixed prompts | Both models + merge | Medium–high for IDs; names variable | ### Success metrics (eval set target) | Metric | MVP target | |--------|------------| | Tamil script person recall | ≥ 85% | | Tanglish person recall | ≥ 70% | | False positives on clean prompts | ≤ 5% | | Structured PII in Tamil prompts (regex) | ≥ 95% | | English canonical regression | 100% | --- ## 11. Git remote setup (Xenovex) Remote: `https://xts.xenovex.com/Bilal-Nazer-Ali/llm-pii-poc.git` ### Create and push ```powershell cd C:\Users\bilal.n\Projects\llm-pii-poc $git = "C:\Program Files\Git\bin\git.exe" & $git remote add origin & $git branch -M main & $git push -u origin main ``` Use an **empty** remote repository (no README) when pushing existing history. ### Committed vs gitignored | Committed | Gitignored | |-----------|------------| | `src/`, `tests/`, `scripts/`, `docs/` | `bin/`, `obj/`, `.vs/` | | `README.md`, `PiiRedaction.slnx` | `models/*.onnx`, `vocab.txt`, `models/en/*`, `models/ta/*` | | `models/en/.gitkeep`, `models/ta/.gitkeep` | `scratch/` | After clone, run model download scripts locally. --- ## 12. Improvement roadmap Post-POC items **not yet implemented**. Current maturity: strong architecture and tests; production needs fail-closed policy, API host, and observability. ### P0 — Security and correctness | Item | Proposal | |------|----------| | Fail-closed when NER unavailable | `RequireNerOnStartup`, `BlockLlmWhenNerUnavailable` options | | Outbound LLM guard | Verify sanitized text before `IChatClient` | | Truncation warning | Surface 128-token limit in `SanitizationResult` | | NER confidence | Populate `PiiEntity.Confidence` or hide UI column | | WPF leak check | Use placeholder map, not naive `Contains` | ### P1 — Platform | Item | Proposal | |------|----------| | `PiiRedaction.Composition` | Shared `AddPiiRedactionServices` (Console + WPF duplicate today) | | Unified prompt catalog | Single source for console, WPF, golden tests | | Async `IPromptSanitizer` | Replace WPF `Task.Run` wrapper | | `PiiRedaction.Application` | Shared orchestration for API/WPF | | Minimal API + health checks | `POST /v1/prompts/sanitize`, model readiness | | Tamil Phase 4 | Numeral normalization, Tanglish heuristics | ### P2 — Operations | Item | Proposal | |------|----------| | Placeholder audit store | TTL, encryption, correlation ID | | Observability | Metrics, OpenTelemetry traces | | ONNX session pool | Concurrency strategy for API load | | CI pipeline | Fast tests without models; nightly real-model job | ### Suggested phases 1. **Production hardening** — fail-closed, composition root, API, truncation metadata (~1–2 weeks) 2. **Detection quality** — normalizer, Tamil Phase 4, regex hardening (~1 week) 3. **Operations** — audit store, metrics, session pool (~1–2 weeks) --- ## 13. Key source files | Topic | Path | |-------|------| | Script classification | `src/PiiRedaction.Core/Detection/ScriptRouter.cs` | | **Model branch decision** | `src/PiiRedaction.Infrastructure/Onnx/RoutingOnnxNerModelRunner.cs` | | English NER | `src/PiiRedaction.Infrastructure/Onnx/EnglishOnnxNerRunner.cs` | | Tamil NER | `src/PiiRedaction.Infrastructure/Onnx/TamilOnnxNerRunner.cs` | | Shared ONNX inference | `src/PiiRedaction.Infrastructure/Onnx/OnnxTokenClassifierRunner.cs` | | NER detector | `src/PiiRedaction.Core/Detection/OnnxNerPiiDetector.cs` | | Sanitizer | `src/PiiRedaction.Core/Sanitization/PromptSanitizer.cs` | | Composite merge | `src/PiiRedaction.Core/Detection/CompositePiiDetector.cs` | | Options | `src/PiiRedaction.Core/Configuration/PiiRedactionOptions.cs` | | DI | `src/PiiRedaction.ConsoleApp/DependencyInjection/ServiceCollectionExtensions.cs` | | Console samples | `src/PiiRedaction.ConsoleApp/Samples/SamplePromptCatalog.cs` | | WPF catalog | `src/PiiRedaction.TestHarness.Wpf/Services/TestPromptCatalog.cs` | | Routing tests | `tests/PiiRedaction.Infrastructure.Tests/Onnx/RoutingOnnxNerModelRunnerTests.cs` | | Script tests | `tests/PiiRedaction.Core.Tests/Detection/ScriptRouterTests.cs` | --- *Last consolidated: July 2026. Replaces separate architecture, NER models, routing reference, Tamil plan, improvement roadmap, and Git setup documents.*