secure-llm-auditor: personal data scanner
secure-llm-auditor
A single-file Python script that finds personal data in text with regexes, anonymizes it and writes a report before a prompt goes to an AI
Medium risk
We rate an entry medium when the tool runs code, makes network calls or reads project files. Check what exactly it does before installing.
Why this level
- Reads local files that may contain real personal data and saves fragments of it into reports
- Detection relies on simple regexes and can miss data, so it should not be trusted as the only safeguard
Install
Manual install
git clone https://github.com/alinalutz/secure-llm-auditor && cd secure-llm-auditor && pip install -r requirements.txtClone and install the colorama and pdfkit dependencies.
This is third-party code. Review the repository files before installing.
What it does
The audit.py script reads a text file and searches with regular expressions for full names, passport numbers, SNILS, INN and mentions of biometrics. Based on what it finds, it assigns a protection level from Russian government decree No. 1119 and a personal data category, replaces the matches with markers like "[NAME REMOVED]" and checks whether anything is left after the replacement. The result is saved as a dated markdown report in the reports folder, and pdfkit, if installed, also produces a PDF. The report prints reference lines about the GOST 28147-89 and GOST R 34.10-2012 algorithms, which the code does not actually check, just prints as a reminder.
Who it is for. For developers who manually check text for personal data before sending it to an LLM and want a quick rough scan rather than a finished legal opinion.
Good fit when
- You want a quick check of a text file for obvious names, passport numbers, SNILS or INN before feeding it to a model
- You need a rough anonymized version of the text for an experiment
- You want a markdown report of what was found for an internal review
Not a fit when
- You need a legally meaningful FZ-152 or GOST R 57580.2 compliance assessment: the script does not check encryption or signatures, it only prints algorithm names
- You need accurate PII detection: the regexes are simple and produce both false positives and misses
- You need a packaged CLI with flags and tests: this is one file without argparse
Example request
Check draft.txt for personal data before I send it to a modelLimitations
The repository has no stated license. It is a single audit.py file with no packaging, no tests and no error handling beyond a missing-file check. The regexes match Russian full names, 10-digit passport numbers, SNILS and 12-digit INN by shape only, without checksum validation or context, so both misses and false positives on ordinary words are likely. The GOST R 57580.2 section in the report is static text, not a check result. The README cuts off at the clone command and references a different repository name. PDF output needs pdfkit and a system wkhtmltopdf; without them only markdown is produced.
How to disable. Delete the local repository copy; nothing is installed system-wide besides the requirements.txt dependencies.
Security check
- Reads local files that may contain real personal data and saves fragments of it into reports
- Detection relies on simple regexes and can miss data, so it should not be trusted as the only safeguard
README in short
The README calls the project a CLI tool for auditing text against FZ-152, decree No. 19/2024 and GOST R 57580.2, listing PII detection, protection-level classification, anonymization, a pre-LLM safety check, and markdown and PDF reports. The listed stack is Python 3.10+, regular expressions, FastAPI for a future web version and Docker for containerization, but the repository only contains the script itself and accumulated test reports. The install instructions cut off at a git clone command with a placeholder instead of the real repository name.
FAQ
Does the script actually verify GOST encryption?
No. The GOST 28147-89 and GOST R 34.10-2012 lines in the report are fixed text, there is no encryption check in the code.
What happens if colorama and pdfkit are missing?
colorama does not affect functionality, pdfkit only controls PDF output: without it only the markdown report is saved.
Related
A code security audit skill by Cloudflare: the agent runs recon, coverage-led hunting and independent verification of findings, then produces a structured repor
NVIDIA's open stack for running OpenClaw, Hermes and LangChain Deep Agents in OpenShell sandboxes with network policy and managed inference
Security scanner for agent skills and MCP servers: finds prompt injection, data exfiltration and supply chain risks before install
Static code analysis with rules that look like source code, plus a built-in MCP server for AI agents