Skip to main content
SparkSolutions

AI Risk & Security · SparkSolutions Editorial

A Customer Asked You to Delete Their Data. Can You Actually Delete What Your AI Learned From It?

Deleting a customer's record from a database is routine. Deleting what an AI chatbot, search tool, or fine-tuned model absorbed from that record is a different, harder problem — and privacy law doesn't care that it's technically inconvenient.

By SparkSolutions Editorial · Published October 2, 2026 · 5 min read

Share

Most businesses that have added an AI search tool, a support chatbot, or a document assistant over the past couple of years built it the same way: customer records, past tickets, contracts, or chat logs got fed into the system so it could answer questions about them. That is a sensible design, and for most purposes it works well. It also quietly creates an obligation most businesses haven't tested yet. Under PIPEDA, Quebec's Law 25, and any version of GDPR that reaches a business's customers, a person can ask to have their personal information deleted. The request sounds routine. Fulfilling it against an AI system that has already ingested that data is often not.

The mechanical problem is specific. A customer-facing AI tool typically works one of two ways. The first, retrieval-augmented generation, stores a customer's documents as numerical representations called embeddings in a vector database, and pulls the relevant ones into the conversation when a question comes in. The second, fine-tuning, trains the model's own internal weights on a set of records so it absorbs patterns from them directly. Deleting a row from the original source database does not automatically remove the matching embedding sitting in a vector store, and it does nothing at all to a model that was fine-tuned on the record — that influence is baked into weights distributed across the model, not stored anywhere you can point to and erase. Researchers studying this problem have shown that even a vector store's "deleted" embeddings can sometimes be reconstructed well enough to recover the original text, which is not the kind of gap a privacy regulator is likely to treat as a technicality.

It is worth knowing where this ends if a regulator decides the underlying data collection itself was the problem, because the endpoint is already established, not hypothetical. The U.S. Federal Trade Commission has, for several years now, used a remedy called algorithmic disgorgement: ordering a company not just to delete improperly obtained data, but to delete the model or algorithm built from it. Cambridge Analytica was the first case in 2019. The FTC has since applied the same remedy to Everalbum, Weight Watchers' Kurbo app, Ring, Edmodo, and Rite Aid. The logic is blunt — a company shouldn't get to keep the benefit of a model trained on data it had no right to use, even after the raw data itself is gone. Canada has no identical remedy on the books today, but PIPEDA's existing correction and deletion obligations point the same direction in principle, and a regulator with a live complaint in front of it does not need a new statute to ask hard questions about where a customer's data actually ended up.

None of this is limited to businesses sophisticated enough to be training their own models. A much smaller, more common version of the same gap shows up in an ordinary support chatbot or internal search tool built on top of a vendor's platform, where customer tickets and documents were uploaded wholesale to make the tool useful. When a customer later asks for their data to be deleted, the honest answer depends entirely on how that tool was built — and a surprising number of businesses running one have never asked their vendor that question, because it never came up before the first real deletion request did.

The distinction that matters in practice is between a system where customer data is retrieved and surfaced on demand, and one where it has been trained into the model itself. A well-built retrieval system can usually support a targeted deletion — remove the record and its embedding, and the tool simply stops being able to retrieve it, which is a clean and defensible answer to a deletion request. A model fine-tuned directly on raw customer records has no equivalent clean removal short of retraining from a dataset that excludes the deleted record, which is expensive, slow, and not something most businesses have budgeted for as an ongoing operational cost. That difference is an architecture decision, made (or defaulted into) when the AI tool was first built, long before anyone was thinking about a deletion request.

The practical step is to find out which situation your business is actually in before a request forces the question. Ask whatever vendor supplies your AI search, chatbot, or document tool a direct question: when a specific customer's data needs to be deleted, what exactly happens, how long does it take, and can you get written confirmation that it's actually gone from every place it landed, not just the source database. A vendor who cannot answer that clearly, or whose answer is some version of "it will stop showing up in new results," has told you something important about what you're actually running. Get the answer in writing and keep it, the same way you'd keep any other vendor commitment that matters if it's ever tested.

None of this argues against building AI tools on customer data, which is exactly what makes most of them useful in the first place. It argues for treating data deletion as a design requirement to check before adopting a tool, not a feature to assume it has. The businesses that handle this well won't be the ones that avoided AI assistants built on customer records. They'll be the ones that asked, before the first deletion request arrived, exactly what deleting a customer's data would actually have to touch — and got an answer they could stand behind.

  • data privacy
  • ai governance
  • vector databases
  • compliance
  • data deletion

Keep reading

Let's discuss what's slowing your business down.

Every engagement starts with understanding your operational pain points. Talk to our team about where intelligent software and automation can deliver measurable results.