Observed Signal · Mar 29, 2026 · Technical Release · Source: DEV Community · Impact: 2/5 · Sentiment: Positive
Open-source ShadowAudit Stops PII Leaks to LLMs
ShadowAudit is an open-source tool that inspects prompts leaving an application to any LLM API and blocks or flags personal data before it reaches the model. It detects items such as email addresses, phone numbers, API keys, and Indian national IDs (Aadhaar and PAN), and can be integrated with two lines of code as a wrapper around existing LLM clients. ShadowAudit also produces GDPR Article 30 compliance reports automatically from its audit logs. The project is published on GitHub by the author as part of an open-source portfolio and the author requests community feedback.
Tool helps developers prevent PII leakage into LLMs and automate GDPR Article 30 reporting, reducing regulatory and security risk for applications that use chatbots/LLMs.
Track OpenAI Signals & Market Shifts in Real-Time
Polaris7 autonomous intelligence agents track regulatory filings, primary sources, executive changes, and deal flow 24/7. Create your free Explorer workspace to monitor these entities.
Key Takeaways & Evidence Grounding
- ShadowAudit is an open-source tool that scans prompts between an app and any LLM API to detect potential personal data leaks.
- It detects emails, phone numbers, API keys, and Indian national IDs including Aadhaar and PAN numbers.
- Integration example shows a two-line wrapper: ShadowAudit.from_config("shadowaudit.yaml") and client = sa.wrap(openai.OpenAI()).
- ShadowAudit can automatically generate GDPR Article 30 compliance reports from its audit log.
- The project's repository is github.com/Jeffrin-dev/ShadowAudit.
Connected Companies & Entities
1 Entity mappedOntology Mapping & Concepts
Related Market Signals & Shifts
Recent verified developments and strategic activity across this market segment.
Weekend-built PII Firewall Blocks LLM Data Leaks
An author built and open-sourced a pre-request governance stack that prevents personally identifiable information (PII) from being sent to LLM providers. Motivated by an incident where a real credit card number was accidentally sent to GPT-4o during benchmarking, the project implements a FastAPI enforcement dependency that scans prompts with Microsoft Presidio before any model call, evaluates YAML-defined policies (block/warn/alert), and short-circuits requests (HTTP 403) when blocking rules fire. The system logs every inference to a PostgreSQL audit vault, sends CloudEvents-compatible webhook alerts (Slack/Teams/PagerDuty), and supports multi-provider routing (OpenAI, Groq, Google Gemini, Anthropic, local Ollama). Code is published on GitHub (sochaty/llm-governance-engine) and the stack runs via docker compose.
LLMs Leak Personal Data, Raising Doxxing Risks
Generative large language models can inadvertently expose sensitive personal data by aggregating dispersed public records and user-contributed content. The article reports examples where models returned private phone numbers, past addresses, and employer links; Reddit users said Google Gemini returned a private number as a service hotline, and security researchers found chatbots suggesting manipulated support numbers placed by fraudsters. Model behaviour is inconsistent: ChatGPT, Gemini and Claude often refuse or limit sensitive outputs, while Grok (xAI) was observed to be more permissive. The piece warns that automated aggregation by LLMs lowers the bar to doxxing, highlights limited consumer options for removing third-party data (especially in German-speaking markets), and calls for stronger legal/political safeguards and technical controls over which personal data LLMs may reveal.
Developer Audits 1,000+ AI Coding Prompts
A developer who sent over 1,000 prompts to AI coding tools built an open-source scanner, reprompt, to analyze what was actually sent. The audit found accidental leaks (three API keys, one JWT, 12 emails, 47 internal file paths), a 35% agent error-loop rate, and that 50–70% of conversation turns were low-information filler. reprompt reads local session files from tools (Claude Code, Codex CLI, Cursor, Aider, Gemini CLI), runs regex-based scans locally with zero network calls, and offers analyses for privacy, agent repetition, and turn importance. The project is MIT-licensed, supports nine AI tools, runs quickly, and is available on GitHub (reprompt-dev/reprompt). The author frames the tool as relevant to compliance concerns under the EU AI Act and as a way for developers to surface credential leakage and inefficient agent behaviors.
Track Real-Time Market Signals & Shifts
Set up custom watchlists to receive automated, evidence-grounded executive digests whenever material signals or shifts occur across your tracked landscape.
