Opinion

Good AI needs good, well governed data

In this exclusive interview, Ashok Krishnan, VP – Commercial Business and Revenue, Centre of Data for Public Good (CDPG), IISc, discusses why India’s next Digital Public Infrastructure (DPI) must focus on trusted data exchanges. Sharing his views on the country’s AI ambitions, challenges of institutional data sharing, need for robust data governance and how platforms such as IUDX could transform healthcare, public services and the broader data economy. He also explains why strong data infrastructure; not just advanced AI models will be critical to India’s digital future.

Aadhaar solved identity. UPI solved payments. Both worked because India built shared, open, rule-based infrastructure rather than leaving every institution to solve the same problem alone. The next frontier is data itself, enabling different systems, departments and organisations to exchange data securely, on the data owner’s terms, without every party having to build bespoke pipes to every other party.

At CDPG, we build exactly this, trusted AI data exchange platforms like IUDX, ADeX and GDI, so that data can move safely between hospitals, cities and within cities and other agencies the same way money now moves between banks. As India’s AI ambitions grow, this layer becomes even more important: good AI needs good, well-governed data and that’s the infrastructure we’re focused on.

Technology isn’t the barrier. Open standards exist and work. And regulation, under the DPDP framework, is maturing sensibly, DPDP exists to ensure every individual’s personal data is handled with care and accountability. That’s exactly what it should do.

The real friction is institutional willingness and it shows up in more than one way.

In some places, DPDP has become a convenient shield for departments that were never inclined to share data in the first place, an excuse for unwillingness, not a reason for it.

But there’s a second, more legitimate concern we take just as seriously: department heads who genuinely want to open up public data, but are worried about the liability if that data is misused downstream, a breach, an unlawful use, an incident that becomes their legal exposure and their career risk. That caution is rational. Nobody should have to bet their job on infrastructure they don’t control.

This is where CDPG’s IUDX platform could aid powering National & State level Trusted Data Exchanges. India doesn’t yet have a formal legal safe harbour for institutions sharing public data responsibly and it should. Until that exists in statute, IUDX is built to give institutions the next best thing: consent enforced by design, purpose-limitation built in, every access logged and auditable, with the data owner retaining control throughout. When something can be shared safely, verifiably, with a complete audit trail, the liability question stops being a reason to say no.

DPDP was never designed to bring public data initiatives to a grinding halt. It was designed to make sure access happens safely. Whether the barrier is unwillingness or legitimate caution about risk, the answer is the same: build the infrastructure that removes the excuse and de-risks the decision and make the case for the law to catch up to what the infrastructure already makes possible. That’s not a technology problem and it’s increasingly not a regulation problem either. It’s a problem of giving institutions the confidence to act. That’s exactly what we’ve built.

And, we’d urge the Indian Parliament to take up the Digital India Act with real urgency, it’s the natural place to formalise exactly this kind of protection, giving institutions a genuine legal safe harbour for responsible data sharing, rather than leaving them to rely on platform-level safeguards alone. And CDPG stands ready to help facilitate that conversation, bringing together the full spectrum of institutions, regulators and industry who need a seat at that table. At our Data For Good Symposium to be held at IISc 25–26 September, 2026, we plan to organise multiple panels and roundtables specifically built to debate these exact questions and more. If even one outcome of that Symposium is a set of concrete inputs that help shape a regulatory framework for a genuinely thriving public data ecosystem, one that works for citizens, institutions and industry alike, we’ll have done our job

Yes. A model is only as good as the data it learns from and that layer has received a fraction of the attention and investment the models themselves have. It’s building faster cars without investing in roads.

This is exactly the gap CDPG’s IUDX can bridge. We haven’t just built a trusted data exchange; we’ve built the ‘Trusted AI Data Platform’: consent enforced by design, anonymisation built in, full data governance with auditable trails on every access. And we’re now adding a confidential compute layer on top of that, which changes the guarantee we can offer data owners in a fundamental way.

With ‘confidential-compute’, an AI model can train and infer on sensitive data without that data ever leaving the owner’s control. Not shared. Not copied. Not exposed, even to us. The model comes to the data, learns what it needs inside a secure enclave and leaves; the data owner retains full sovereignty throughout. That’s the difference between data being shared and data being unlocked.

But here’s where it gets genuinely interesting. The same trusted enclave lets us bring multiple parties’ data together for a single model, without any of them handing their data to each other. Take throat cancer detection. Train a model purely on anonymised, annotated medical imaging and clinical history and you get one level of signal. Now bring in population-level dietary and lifestyle data and macroeconomic data on the sale and consumption of the products that actually drive risk, tobacco, alcohol, specific dietary patterns, inside that same secure enclave, as context and controlled noise around the core clinical data. Does that make the model materially stronger? We don’t fully know yet. But we now have the infrastructure to actually find out, safely, without any institution ever having to hand its data to another. That’s the kind of question this platform exists to let India ask.

Until the industry treats infrastructure like this as core AI infrastructure, not a compliance afterthought, the models will keep outrunning the foundations they stand on.

Healthcare. Not because the other sectors matter less, but because healthcare is where trusted data exchange stops being a convenience and becomes the difference between life and death.

Start with epidemics. Right now, we largely detect outbreaks after they’ve already become outbreaks, after enough people have shown up sick enough, often across siloed hospital systems that never talk to each other. A functioning health data exchange changes the timeline. It means catching the signal early enough to stop an epidemic before it breaks out, not just respond faster once it has.

Public health policy improves the same way. Right now, too much policy is built on estimates, extrapolation and data that’s months or years old by the time it reaches a policymaker’s desk. Imagine Public Health Policies built instead on real, current data, from patients who’ve willingly contributed their anonymised health information specifically because they want to advance research, not because they had no choice. That’s a fundamentally different and much stronger, foundation for public health decisions.

But the example I keep coming back to is the simplest one and the most human. Picture a road accident. You’re rushed to the nearest hospital, not the one that knows you, with family members who are panicked and may not even be with you and who couldn’t reliably recall your allergies or your current medications even if they were. In that moment, a functioning health data exchange means a doctor can pull your medical history from your driver’s license number or your ABHA ID and know, within seconds, what you’re allergic to and what you’re taking. That’s not a convenience. In an emergency, that’s the difference between the right treatment and a preventable mistake. If we build nothing else, building the infrastructure that makes that moment possible for every Indian is worth building.

Pharma R&D is bottlenecked by the same problem, sensitive health data locked inside individual institutions, too fragmented to power real discovery. Trusted data exchange changes that: AI models trained on properly governed, pooled health data can meaningfully accelerate drug discovery, including for rare diseases that pharma economics has long underserved, because no single hospital’s patient population was ever large enough to justify the research. Pool that data safely across institutions and borders and it finally is. That matters for the millions of Indians living with a rare disease and very few options.

Through IUDX Health and our work with ICMR-NIE on ADARV, we’re already proving pieces of this, faster outbreak detection, evidence-based public health planning. The next five years are about connecting those pieces into the system I’ve just described.

I do. The timeline will be longer than UPI’s, data is more varied and more sensitive than money, but the logic is identical. UPI worked because it was a shared, neutral layer any bank or app could plug into. A national data exchange layer does the same thing for data: one trusted way for any authorised system to request and receive it, instead of a thousand one-off arrangements duplicating the same work.

Get this right and the gains go well beyond efficiency, welfare delivery that actually reaches who it’s meant for, public health response measured in hours instead of weeks, services that anticipate what citizens need instead of making them prove eligibility over and over again.

It’s a real risk and one we take seriously. Any infrastructure that assumes large in-house technical teams will naturally favour large players. That’s not a hypothetical, it’s just what happens when nobody designs against it.

That’s why CDPG builds on open standards and open-source software. A small organisation integrates using the same public specifications as a large one, with no proprietary access to buy and no lock-in to escape. We invest deliberately in documentation, reference implementations and low-cost onboarding because a data exchange is only a public good if SMEs and the institutions without dedicated engineering teams can use it too, not just the ones that can afford to.

By 2030, I’d like data-sharing to feel as unremarkable as digital payments feel today; something citizens and institutions do routinely, safely and without a second thought. 

Concretely, I’d hope to see trusted data exchanges operating at District, State & National, scale across health, finance, urban and agricultural systems; a generation of AI applications in India built on well-governed, high-quality public data rather than scraped or siloed sources; and smaller cities, states and enterprises participating in this ecosystem on equal footing with the largest players. 

The real measure of success won’t be the platforms we’ve built; it will be whether citizens experience faster, fairer, more responsive public services because of them.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button