Observed Signal · Apr 26, 2026 · Technical Release · Source: DEV Community · Impact: 3/5 · Sentiment: Positive
PIIGhost: Open‑Source Python PII Anonymizer for LLM Agents
PIIGhost is an open-source Python library that provides detection, anonymization and deanonymization pipelines to hide personally identifiable information (PII) when interacting with large language models and agent frameworks. It composes detectors (regex, exact-match, GLiNER2 NER), span conflict resolvers, entity linking and merging, placeholder factories, and anonymizers into a reusable AnonymizationPipeline. PIIGhost includes a conversation-aware ThreadAnonymizationPipeline (per-thread memory) and a PIIAnonymizationMiddleware that integrates with LangGraph/LangChain to anonymize messages before LLM calls, deanonymize tool arguments, and restore user-facing text. The project includes a human-in-the-loop demo (piighost-chat) and is published on GitHub with documentation. The library targets practical PII risks in agent workflows and multi-message conversations.
Provides an open-source, practical toolkit for PII anonymization in LLM agent workflows and conversational UIs — important for privacy compliance and safe agent deployment, but not a major platform-level release.
Track LangChain Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- PIIGhost is an open-source Python library for detecting and anonymizing confidential PII for LLM agents.
- It supplies detectors (RegexDetector, ExactMatchDetector, GLiNER2-based Gliner2Detector, CompositeDetector) and a generic AnyDetector protocol to plug custom detectors.
- PIIGhost provides span conflict resolvers (e.g., ConfidenceSpanConflictResolver), entity linkers (ExactEntityLinker), entity resolvers (MergeEntityConflictResolver, FuzzyEntityConflictResolver), and anonymizers with multiple PlaceholderFactory options (LabelCounter, LabelHash, FakerCounter, Mask).
- A ThreadAnonymizationPipeline adds conversation-scoped memory to maintain consistent placeholders across messages; a PIIAnonymizationMiddleware integrates the pipeline into LangGraph/LangChain agent workflows.
- Source code and demo repositories are available: github.com/Athroniaeth/piighost and github.com/Athroniaeth/piighost-chat, with documentation at athroniaeth.github.io/piighost/fr/.
Connected Companies & Entities
7 Entities mappedOntology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
piighost-proofreader: Anonymized CV Proofreading with LLMs
The article describes piighost-proofreader, an open-source project that lets users have their CVs proofread by large language models without exposing personal data. The tool anonymizes sensitive fields locally (using the piighost anonymization service with a per-CV thread_id), sends the redacted Markdown to an LLM, then streams structured mistake objects back to the frontend using the instructor library. Corrections are remapped onto the original PDF using PyMuPDF word streams and a sequence of four fallback locator strategies to handle tokenisation and multi-column drift. The author provides repository links and a live demo URL for the project.
Weekend-built PII Firewall Blocks LLM Data Leaks
An author built and open-sourced a pre-request governance stack that prevents personally identifiable information (PII) from being sent to LLM providers. Motivated by an incident where a real credit card number was accidentally sent to GPT-4o during benchmarking, the project implements a FastAPI enforcement dependency that scans prompts with Microsoft Presidio before any model call, evaluates YAML-defined policies (block/warn/alert), and short-circuits requests (HTTP 403) when blocking rules fire. The system logs every inference to a PostgreSQL audit vault, sends CloudEvents-compatible webhook alerts (Slack/Teams/PagerDuty), and supports multi-provider routing (OpenAI, Groq, Google Gemini, Anthropic, local Ollama). Code is published on GitHub (sochaty/llm-governance-engine) and the stack runs via docker compose.
Deterministic Tokenization Preserves LLM Output; NoPII Launched
The author ran 109 structured tests across GPT-4o, Claude 3.5 Sonnet, and Gemini 1.5 Pro to measure how PII protection methods affect LLM output quality. They compared raw prompts, generic placeholder masking (e.g., [PERSON]), and deterministic opaque tokenization (unique tokens per entity). Tokenization preserved 91–96% of raw output quality while placeholder masking reduced quality to 54–68%, with steep drops in entity consistency and reasoning coherence for multi-entity prompts. Tests also revealed a safety refusal mode: models declined ~15–20% of tokenized prompts when labels like “SSN” remained next to opaque tokens; replacing labels (context phrase neutralization) resolved the refusals. Based on the research, the team built NoPII, a reverse-proxy privacy layer that tokenizes PII in transit, detokenizes responses, supports streaming, uses a PCI Level 1 / SOC2 vault, and integrates with major LLM providers with one base_url change.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
