The Short Answer
Redaction breaks enterprise AI because it deletes the relationships, not just the identifiers. The fix is transformation before training plus runtime filtering on live chat and uploads. On ibl.ai you own all the code and the data, so both layers run inside your perimeter.
Every regulated enterprise hits the same wall. The data that would make an agent genuinely useful — support tickets, CRM histories, claims files, contract workflows — is the data privacy rules will not let into a pipeline unprotected.
The standard answer has been to strip it. The standard result has been a model that no longer knows anything worth knowing.
Why does redacting PII make a model worse, not just smaller?
Because identifiers are what hold a record set together. Redaction removes data points and leaves gaps where correlations used to be.
Take a fraud model trained on transactions. Redact the customer names and behaviour can no longer be tracked across accounts.
Mask the account numbers and a ring spanning several identities becomes five unrelated customers. Strip the timestamps and the temporal shape that separates ordinary activity from suspicious activity goes with them.
The pattern repeats in every vertical: the compliance review finishes, the approved extract lands, and it is too degraded to train on. Months are spent getting permission to use data that no longer answers the question.
What is PII transformation, and how is it different from masking?
Transformation replaces each real entity with a consistent synthetic one across the entire corpus, so the relationships survive. Masking replaces values with generic tokens and destroys uniqueness; redaction deletes them outright.
| Approach | What happens to the identity | What happens to the relationships |
|---|---|---|
| Redaction | Deleted | Broken — gaps where the correlations were |
| Masking | Replaced with a generic token | Collapsed — every person becomes the same person |
| Transformation | Replaced with a consistent fictional identity | Preserved — same structure, different people |
On 24 September 2026, micro1 released flow-transform 1.0, a model built for exactly this.
Rather than anonymising each record on its own, it gives each real-world entity one synthetic counterpart that holds across the dataset — what the company describes as a privacy-preserved digital twin of the enterprise.
Ali Ansari announced it on X; the benchmark table below is as reported by RuntimeWire.
Treat it as a signal about where the field is going rather than something to put in a plan this quarter: the launch materials do not say how it is sold or where it sits in micro1's existing products.
| Measure | Score |
|---|---|
| PrivacyBench (Tonic.ai) — detection F1 | 96.0% |
| Identity synthesis accuracy | 98.28% |
| Combined detection + synthesis | 95.46% |
| Enterprise De-Identification Bench TQI — NVIDIA NeMo Anonymizer | 74.9 |
| …flow-transform 1.0, no agentic review | 84.1 |
| …flow-transform 1.0, with agentic review | 88.9 |
Two things to hold on to when reading that table. 74.9 is NeMo Anonymizer's score, not flow-transform's own pre-review baseline — flow-transform without agentic review is 84.1, and 88.9 with it, so agentic review is worth roughly five points, not fourteen.
And the Enterprise De-Identification Bench is micro1's own: it built the benchmark, defined the TQI metric and chose the weights.
Its own description says the corpus was generated from templates, represents one fictional company's records, and covers structured tabular data — it "does not establish performance on a live company's files or on complex documents and other formats."
That caveat lands directly on the argument this post is making. The data enterprises actually want is messy prose and scanned paper, and a vendor beating NVIDIA on a metric the vendor designed is a sanity check, not a scoreboard.
Where does PII enter an AI system after the training data is clean?
Everywhere a person types or uploads. A clean training corpus says nothing about the conversation happening right now.
An employee pastes a customer record into a chat to ask a question about it. A claims handler uploads the claim PDF. Someone drops in a screenshot of an internal console with account numbers on screen.
None of that passed through the compliance review, and all of it reaches a model.
On 25 September 2026 we extended PII and PHI filtering on the ibl.ai platform to cover in-chat file uploads across every path a file can take to a model — text, Office documents, images and PDFs, on the OpenAI and Google multimodal routes, the graph and deep-agent routes, Claw, and the code interpreter.
Images go through OCR and are redacted at the pixel level. PDFs are rasterised, redacted and rebuilt. Each agent's own privacy and PHI settings then decide what happens: block, redact, or allow.
How do you prove to a regulator what the filter actually did?
With an audit trail that records the detection without recording the data.
Every detection — file upload, prompt input, model output, code-interpreter extraction, memory retrieval — is written to a read-only, platform-admin-scoped privacy-flags endpoint, filterable by rail, action taken, source, agent, session and date.
The design decision worth copying: it stores entity types only, never raw values.
A log that captured what it detected would be a second copy of the sensitive data, sitting in a system with different access controls and a different retention policy than the one it is auditing. Most audit tooling gets this wrong.
What does an enterprise actually have to build?
Three layers, and most programmes only build one.
1. Pipeline transformation. Models like flow-transform 1.0, so training and analytics run on statistically valid data that contains no real identity.
This helps satisfy GDPR, HIPAA and CCPA obligations; it does not discharge them.
Consistent cross-corpus identity replacement is pseudonymisation-shaped, and under GDPR pseudonymised data is still personal data — the consistency that makes it useful is exactly what preserves linkage.
HIPAA de-identification still requires Safe Harbor or Expert Determination, and no benchmark score confers either.
2. Runtime filtering. Multi-modal detection across text, images, PDFs and structured data, applied per interaction rather than per batch, configurable by data type and regulatory framework.
3. Audit infrastructure. Every privacy action logged, queryable and exportable, holding metadata rather than a duplicate of the content.
Buying only the first leaves the live conversation unprotected. Buying only the second leaves the model untrained. Buying neither is the status quo that makes regulated AI programmes stall.
Why does ownership decide whether any of this is enough?
Because a privacy control you cannot inspect is a promise, not a control. If the filter, the audit log and the records live in a vendor's account, the strongest statement you can make to a regulator is that a third party says it handled your data correctly.
On ibl.ai you own all the code and the data. The platform runs under a perpetual licence on your own infrastructure — your cloud, your VPC, on-premise, or fully air-gapped — so the detection rules, the privacy flags and the records they describe never leave your perimeter.
It is model-agnostic, so a regulated deployment can run an open-weight model entirely inside its own network with no external API call to reason about at all, and there is no per-seat pricing to make the compliant path the expensive one.
More than 1.6M users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.
ibl.ai is family-owned and operated from New York, NY — a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.
The teams that fix the PII pipeline stop having the same meeting every quarter. The ones still redacting will keep wondering why the model does not know anything.
Related: HIPAA-Compliant AI: Keeping PHI on Your Own Infrastructure — the same argument where the regulator is HHS and the data is clinical.