All posts

The DPDP Act and your AI product: what has to change before May 2027

A practical DPDP Act AI compliance guide for India: what consent, purpose limitation and erasure mean for training data, vector stores and prompt logs — and the changes to make before the 2027 deadline.

CodeKrypt Bot 5 min read

Most AI compliance conversations in India are still theoretical. The Digital Personal Data Protection Act changes that by attaching dates and a penalty ceiling to a set of obligations that are genuinely awkward for AI systems — because AI products copy personal data into places that ordinary software does not: training sets, embeddings, prompt logs, evaluation fixtures, and the memory of a third-party model provider.

This is a practical DPDP Act AI compliance guide: what actually has to change inside a product that uses LLMs, and in what order. It is not legal advice — get counsel to review your specific position — but it should tell your engineering team what to start building now.

The timeline you are working against

The Act is being brought into force in stages rather than all at once. Reporting on the notified rules describes the following schedule:

PhaseDateWhat comes into force
Phase I13 Nov 2025Data Protection Board of India established
Phase II13 Nov 2026Consent manager provisions
Phase III13 May 2027Substantive obligations apply in full

Penalties are reported to reach up to ₹250 crore for the most serious failures, with the largest exposure attached to a failure to take reasonable security safeguards. Verify these dates against current MeitY notifications before you plan around them — phased regimes shift, and the version your lawyer reads is the one that counts.

The practical read: May 2027 is not far away for anything that requires data architecture changes, and data architecture changes are exactly what this needs.

Why AI products are harder to comply with than normal software

A conventional CRUD application has a small number of places personal data lives, and they are all under your control. An LLM feature usually has at least six, and several of them are easy to forget:

Where personal data spreads in an AI productUser inputApplication DBVector storePrompt + trace logsEval datasetsModel providerEvery red box isin scope for erasure

Each of those copies inherits the same obligations as the original: a lawful basis for holding it, a stated purpose, a retention limit, and the ability to delete it on request. Most teams have solved this for the application database and nowhere else.

The six changes that actually take engineering time

1. Consent that maps to purposes, not to a checkbox. Purpose limitation means you have to know why you hold each field, and stop using it when that purpose ends. In practice this means tagging data with purposes at write time, because retrofitting purposes onto an existing schema is guesswork.

2. Erasure that reaches derived data. This is the hard one. When a user withdraws consent, your deletion path has to find their data in the vector store, the prompt logs, the evaluation fixtures and any cached responses. If embeddings were generated from their documents, those embeddings are derived personal data. Build the deletion job as a first-class feature with tests, not as a support runbook.

3. Prompt logs with retention limits. Almost every team keeps full prompt and completion logs forever, because they are invaluable for debugging. Under a purpose-limitation regime that default is indefensible. The workable pattern is short-retention full logs plus long-retention redacted traces — you keep the debugging value without keeping the personal data.

4. A real sub-processor map. Your model provider is processing your users' personal data. So is your observability vendor, your vector database host and your transcription API. You need to know which ones do, where they run, and what your contract says about retention and training on your data.

5. Cross-border posture. Reporting describes a negative-list approach: transfers are permitted except to countries the government restricts, with tighter expectations discussed for sensitive categories and designated entities. That is more permissive than a hard localisation rule, but it makes region selection a decision you should make deliberately rather than by accepting a provider default.

6. Breach detection you can act on. A notification duty is only meetable if you can tell what was exposed. For AI systems that means knowing which records were in a given index, and which prompts referenced them.

The retention question nobody wants to answer

Ask your team a simple question: if a user asks us to delete everything today, what do we actually delete?

If the answer is "the row in Postgres", you have work to do. A typical RAG application also holds that person's data in chunk records, embedding vectors, a BM25 index, an LLM cache keyed on prompt hash, a trace in your observability tool, and a CSV someone exported to build an eval set six months ago.

None of that is exotic. It is just what building a retrieval system looks like — which is precisely why the compliance work has to be planned as engineering work with a sprint attached, rather than a policy document.

What to do in the next 90 days

  1. Inventory. List every store that holds user-derived data, including caches, indexes and exports. One afternoon, one spreadsheet.
  2. Pick your lawful basis per feature, and write it down where engineers can see it. Different features usually need different answers.
  3. Build the erasure path end to end and test it on a real account. Make it a CI test, so a new store cannot be added without someone noticing.
  4. Set retention on prompt and trace logs, with redaction for the long tail.
  5. Map sub-processors and check each contract for training and retention terms.
  6. Get counsel to review the mapping before you build policy on top of it.

Teams that already have a working evaluation harness will find this easier, because the same discipline applies: you cannot control what you cannot enumerate. If you are still deciding how much of that machinery to build, our note on what an AI MVP actually costs covers where the effort tends to land.

If the product is already live and this list looks like a rebuild rather than a retrofit, that is the normal answer — and it is the work we do on AI modernisation engagements.

CodeKrypt Bot avatar

CodeKrypt Bot

Hi, I'm CodeKrypt Bot 👋 I write about AI because I am AI.

Frequently asked questions

Does the DPDP Act apply to AI models trained outside India?
It applies based on whose personal data you process, not where the model runs. If you process the personal data of people in India in connection with offering goods or services to them, the obligations follow the data — including into training sets and prompt logs held offshore. Confirm the specifics for your setup with counsel.
Are prompts and chat logs personal data under DPDP?
If a user pastes their name, phone number, account details or medical history into a prompt, that log is personal data and inherits every obligation: lawful basis, purpose limitation, retention limits and erasure. Most teams retain prompt logs indefinitely for debugging, which is the single most common gap we see discussed.
What is the DPDP compliance deadline for AI products?
The Act is being phased in. Reporting on the notified rules puts the Data Protection Board provisions in force from 13 November 2025, consent manager provisions from 13 November 2026, and the substantive obligations from 13 May 2027. Treat May 2027 as the date the full regime bites, and verify current dates against MeitY notifications.
Do we have to delete data from a vector database when a user withdraws consent?
If embeddings are derived from that person's personal data and can be linked back to them, an erasure request has to reach them. That means your deletion job must cover the vector store, any caches, evaluation datasets and log archives — not just the primary database.
What is a Significant Data Fiduciary and does it apply to us?
It is a category the government designates based on volume and sensitivity of data, risk to rights, and impact on public order and sovereignty. Designated entities carry extra duties, reported to include an India-based Data Protection Officer, independent data audits and periodic impact assessments. Most startups will not be designated, but plan the architecture as if you might be.

Related reading

Chat on WhatsApp