Dpdp act ai model extraction india

India processes personal data at a scale few nations can match. Over a billion Aadhaar-linked identities, hundreds of billions of UPI transactions a year, and a digital services economy that has made "India stack" a genuine global reference point — this is the backdrop against which the Digital Personal Data Protection Act, 2023 was written. It is also the backdrop against which a quieter, less-discussed risk is beginning to surface: what happens when the AI models trained on all of this data can be made to reveal, or reconstruct, fragments of what they learned?

 

This isn't a hypothetical concern borrowed from somewhere else. It is a specific, technical vulnerability class — commonly referred to as model extraction and distillation attacks — that lets an outsider systematically query a deployed AI model and, over enough interactions, either clone its functional behaviour or coax it into reproducing pieces of the very data it was trained on. Global research has demonstrated this against real production systems, not just laboratory prototypes. For Indian businesses, the more pressing question is what this means under the country's own data protection law — a law with real enforcement machinery now standing behind it.

 

 

A Law That Arrived With Teeth, Not Just Intentions

The DPDP Act is India's first comprehensive digital privacy law, passed by Parliament in August 2023. For two years it existed largely on paper, awaiting the operational detail needed to enforce it. That changed on 13 November 2025, when the Ministry of Electronics and Information Technology notified the DPDP Rules, 2025 and formally constituted the Data Protection Board of India. The Act's provisions are now rolling out in three deliberate phases: the Board's establishment took effect immediately; the consent manager framework becomes operational in November 2026; and the remaining substantive provisions — the core rights, obligations, and security requirements most businesses will actually feel day to day — come into full force by 13 May 2027.

 

That timeline matters enormously for how Indian organizations should be thinking about AI risk right now. This is not a law on the distant horizon. It is a law with a regulator already seated, a penalty regime already defined, and a countdown already running toward full applicability. Any organization fine-tuning or deploying AI models on personal data today will be operating under this law's full weight well within the working life of those very models.

 

 

Reach That Extends Well Beyond India's Borders

One detail regularly surprises businesses encountering the Act for the first time: its jurisdiction is not confined to companies incorporated in India. The DPDP Act applies to the processing of digital personal data of individuals located in India, and it extends to processing carried out entirely outside India, where that processing relates to offering goods or services to people within the country. A global SaaS company serving Indian users from servers in another continent is within scope. So is an Indian company that has quietly routed its model fine-tuning through an overseas cloud provider. Geography offers little shelter here — what matters is whose data is being processed, not where the processing physically happens.

 

 

Where AI Model Risk Meets Statutory Obligation

The Act structures its protections around two roles: the Data Principal, the individual to whom personal data belongs, and the Data Fiduciary, the organization that decides how and why that data is processed. Several of the rights and duties built into this relationship intersect directly with the risk that model extraction and training-data leakage represent.

 

The Right to Correction and Erasure, under Section 12, allows a Data Principal to demand that inaccurate data be corrected or that data no longer necessary for its original purpose be deleted. This sounds like a straightforward database operation — until the data in question has already been used to train or fine-tune a model. Deleting a row from a table is trivial. Removing the influence of that row from a trained model's parameters is not; the information has, in effect, been diffused into the model's behaviour rather than stored in a single retrievable location. An organization that believes it has honoured an erasure request in good faith may not have actually erased anything from the model's memory at all — a gap that only becomes visible if someone specifically goes looking for it, whether that someone is a regulator, a researcher, or an attacker.

 

Rule 8 of the DPDP Rules imposes an even broader, proactive obligation: personal data must be erased once the purpose it was collected for is no longer being served, regardless of whether the individual ever asks. For any organization running large-scale consumer platforms — fintech apps, e-commerce marketplaces, digital lending products — this creates a continuous erasure duty that any model trained on that underlying data quietly inherits, whether or not the engineering team building the model was thinking about the Act at all when they did so.

 

Breach notification duties, under Section 8(6) and Rule 7, require a Data Fiduciary to notify the Data Protection Board without delay, and to inform affected Data Principals within 72 hours, including a plain-language account of what data was exposed. This provision was almost certainly drafted with conventional breaches in mind — a compromised server, a leaked database — but its language does not exempt an AI system. If a model can be shown, through inference or extraction, to have leaked personal information it was trained on, a defensible legal reading suggests that qualifies as exactly the kind of unauthorized disclosure this provision was built to catch — no stolen credentials or breached firewall required.

 

 

The Financial Weight Behind These Obligations

India's regulator has been given genuinely significant enforcement power, and the numbers involved deserve boardroom attention, not a footnote. Failure to notify the Board or affected individuals of a data breach can attract penalties of up to INR 200 crore per incident. Failure to implement reasonable security safeguards — a category an unmonitored, easily-extractable AI model arguably falls squarely within — can draw penalties of up to INR 250 crore. The Act treats repeated violations as an aggravating factor, meaning penalties compound rather than simply repeat for organizations that don't correct course. These are not advisory fines; they are enforceable as a decree of a civil court, giving them the same legal weight as a judicial order.

 

To put that in perspective: a single INR 250 crore penalty runs into tens of millions of US dollars — for one category of failure, arising from one incident, entirely separate from the reputational fallout, customer trust erosion, and potential downstream litigation that would follow.

 

 

A Higher Bar for India's Largest Data Handlers

The Act creates a distinct category — Significant Data Fiduciaries (SDFs) — for organizations the government designates based on the volume and sensitivity of personal data they process, and the potential risk their processing poses to individuals' rights. Any Indian organization training large-scale AI systems on substantial volumes of customer data is a natural candidate for this classification. SDF status brings additional weight: appointing a India-based Data Protection Officer, conducting periodic Data Protection Impact Assessments, and submitting to independent data audits.

 

This raises a question most compliance functions haven't yet asked themselves: if a Data Protection Impact Assessment is already a mandatory exercise for how your organization handles personal data broadly, does that assessment currently extend to whether your trained AI models can leak that same data back out — or does it stop at the database and infrastructure layer, leaving the model itself unexamined?

 

 

What Makes This a Distinctly Indian Challenge

A few dynamics specific to India's regulatory design and digital economy sharpen this risk in ways that don't simply mirror privacy regimes elsewhere.

 

A consent-first regime leaves less room for justification after the fact. Unlike some data protection frameworks that recognize multiple lawful bases for processing, the DPDP Act leans heavily on consent as the primary ground on which personal data may be processed. That means the data flowing into a training pipeline was almost certainly collected under a specific, stated purpose — narrowing an organization's ability to later argue that an unexpected downstream consequence, like a model memorizing and later exposing that data, was within scope of what a user actually agreed to.

 

The regulator is new, and its posture on AI-specific harms is still forming. The Data Protection Board of India is a young institution, still establishing precedent on how it will treat novel categories of harm. That is a genuine double-edged reality for businesses: there is no settled case law yet defining how an AI extraction incident would be assessed or penalized — but that absence of precedent also means the organizations that take this risk seriously now, ahead of any enforcement action, are the ones best positioned when precedent does eventually get set.

 

AI adoption in India's most data-sensitive sectors has outpaced governance maturity. Fintech, digital lending, healthtech, and insurtech companies across the country have adopted AI and begun fine-tuning models on customer data at a pace that has, in a great many cases, moved faster than formal data governance structures could keep up with. These are precisely the sectors holding the most sensitive categories of personal data — financial histories, health records, identity information — and precisely where a successful extraction or inference attack would carry the sharpest regulatory and reputational consequences under the Act's security and breach-notification requirements.

 

Cross-border AI infrastructure adds a layer many teams haven't fully mapped. A large share of Indian AI deployments rely on foreign-hosted models, cloud infrastructure, or third-party fine-tuning services. Every one of those handoffs is a point at which personal data — and the statutory obligations attached to it — travels along with it, even when the vendor providing the AI capability sits entirely outside India's borders.

 

 

The Practical Takeaway

The DPDP Act does not name model extraction, membership inference, or distillation attacks — no data protection statute anywhere in the world currently does. But the obligations the Act creates were never written to require that specificity. If personal data enters a model, and that model can later be shown to leak, reconstruct, or reveal traces of that data, the Act's existing machinery around erasure, breach notification, and security safeguards is already built to respond — regardless of whether the organization involved had ever previously considered an AI model to be a place where personal data could reside, and from which it could later resurface.

 

That gap is far better closed on an organization's own initiative than discovered by the Data Protection Board, a journalist, or an attacker. For any Indian business fine-tuning or deploying AI models on customer data, the useful next step isn't a legal memo produced in isolation from the technical reality — it's finding out, concretely and technically, whether the models already in production are actually exposed to extraction and inference risk, and whether the organization's current DPDP compliance posture accounts for that exposure at all.



Comments

No Comments Found.