Substack

Showing posts with label Public Policy. Show all posts
Showing posts with label Public Policy. Show all posts

Friday, July 31, 2026

Forward guidance and Kevin Warsh

The market reaction to Kevin Warsh’s second FOMC meeting decision to stay the course on interest rates despite rising inflation has triggered a debate on the Fed’s credibility in fighting inflation. 

This is how the $31 trillion US Treasury market reacted.

Long-term government borrowing costs shot higher as Mr. Warsh spoke, with the 30-year bond notching its largest one-day increase in more than a year. Trading around 5.22 percent, it is at the highest level since 2007. The 10-year Treasury yield, which serves as the benchmark for borrowing costs around the world, also rose alongside expectations about inflation over a longer time horizon.

Warsh has made it clear that he does not agree with the policy of forward guidance, and intends to discontinue it. He has argued that forward guidance distorts market incentives and comes in the way of market discipline. This blog is in agreement with it. 

The reversal of forward guidance comes abruptly, at a time when market expectations of the Kevin Warsh era are vitiated by the circumstances surrounding his appointment. Markets clearly think that Warsh may buckle down to pressure from the White House or prefer to please it and delay raising interest rates. 

In this context, I created the table below that evaluates forward guidance and central bank credibility. 

The high credibility and low forward guidance Volcker era is the golden quadrant. The central bank has earned trust through demonstrated willingness to act decisively, but it doesn’t tell you what it’s going to do next. Participants must do their own work to assess fundamentals, price risk properly, and build in uncertainty premia. Volcker’s Fed famously targeted money supply, not interest rates, and didn’t telegraph moves. The result was that markets had to think for themselves, and the discipline that emerged was real. Risk was priced because nobody could assume a backstop.

The Bernanke-Yellen era of high credibility coupled with a high degree of forward guidance ensured that the latter became the market’s mirror. When participants believe the central bank will do what it says, and the central bank tells them exactly what it will do, the rational strategy is to front-run the guidance rather than analyse fundamentals. The “dot plot” era, the “considerable period” language, the “whatever it takes” formulation created a world where the trade was to decode the Fed rather than decode the economy. The moral hazard is structural insofar as investors are rewarded for taking risk they don’t understand because the central bank has told them the floor exists.

The worst combination for policy effectiveness is that of low credibility and high forward guidance. The central bank issues guidance but nobody believes it. Participants pursue their own interests, and policy actions run counter to market expectations. This produces whipsaw where the central bank says one thing, markets price another, and when the policy actually arrives it dislocates rather than stabilises. One could argue the late-stage BOJ falls here (decades of forward guidance that markets progressively stopped believing).

Finally, the Warsh era tries to limit forward guidance but in a period of low central bank credibility, inducing a period of “greatest uncertainty”. Worse still, the transition itself amplifies the shock - moving from a regime of high credibility andhigh forward guidance (the post-2008 consensus) to one where both are removed simultaneously. The question to analyse is whether the Warsh Fed currently sits in the fourth quadrant.

There are four points to be noted.

First, the transition path matters enormously. Moving from the Bernanke-Yellen quadrant (high-high) to the Volcker quadrant (high-low) requires maintaining credibility while withdrawing guidance. This means the central bank must demonstrate through actions (not words) that it will do the right thing even when it doesn’t tell you what the right thing is. The Warsh challenge is that he has withdrawn guidance and may not yet have established credibility through a demonstrated willingness to act decisively against inflation. The risk is that he lands in the bottom-left quadrant (low-low) rather than the top-left (high credibility, low guidance).

Second, while I had blogged here listing forward guidance as an important channel in eroding market discipline, it must be noted that it’s not forward guidance per se that erodes discipline, but forward guidance combined with high credibility. Low-credibility forward guidance is merely useless. And high-credibility forward guidance is actively dangerous because it works too well, by substituting the central bank’s judgment for the market’s own.

Third, there’s a temporal asymmetry in credibility. Credibility is earned slowly (Volcker needed 2-3 years of painful rate hikes) but lost quickly (a single capitulation can destroy it). Warsh’s current position - rates unchanged while prices rise - coupled with the circumstances of his appointment (and the credibility problem it engenders) risks a credibility test analogous to Arthur Burns in 1972-73. If inflation accelerates and the Fed is seen as having waited too long, the move from “low guidance with emerging credibility” to “low guidance with low credibility” could be swift and self-reinforcing.

Fourth, this also highlights the less discussed aspect of policy making that involves shaping expectations. Once collective beliefs are formed and incentives aligned, it is very hard to reshape them. The Bernanke-Yellen response to the market tumult during and after the Global Financial Crisis, guided by the research works of Gauti Eggertsson and Michael Woodford, and Paul Krugman, may well have been required in its immediate aftermath. The mistake was to continue and institutionalise it as part of the regular central bank policy toolkit. Bernanke or Yellen, with their credibility, ought to have exited forward guidance once normalcy was restored. But instead, they chose to please the markets with this new crutch. It is a bit like subsidies - once offered, there’s no sunset. 

Monday, July 13, 2026

Workers and startups are helping train AI to replace them

Data annotation work is increasingly moving up the value chain, from tagging and labelling data to replicating the work of semi-skilled (on the factory floor) and skilled (consultants, analysts, lawyers, engineers, and doctors) workers. 

Startups sell data to AI labs, which use it to train and refine their AI algorithms and develop software products/solutions that replicate the work of these workers. In other words, the semi-skilled and skilled workers, or at least some among them, are feeding their time and skills into the AI algorithms that seek to replace them and their kind. Both the training startups and those workers offering their services to them are basically helping make themselves redundant. 

On this, the FT has a very good film about how Indian startups are paying factory floor workers and gig workers (and even people in their homes doing regular household chores) to use cameras and record their work. Data annotation is becoming the new BPO for India’s IT industry. 

The Ken has an article that raises the possibility that for all the attention and hype around robotics, Indian startups might remain stuck at the lowest end of the robotics value chain - data collection. 

India was the back office for the IT boom. It became the annotation and reinforcement-learning labour pool for the generative AI boom. It is now emerging as the behavioural data factory for the physical AI boom... Building the robot is only half the problem. Building the intelligence behind it is much harder. That requires data. Vast amounts of it. Unlike large language models, which were trained on the equivalent of hundreds of years of human reading scraped from the internet, robotics companies are working with barely a fraction of that in video... What they need is meticulous, first-person recordings of humans interacting with the physical world, carefully collected, annotated, and painstakingly structured. 

So the industry turned to India. Across the country, workers are recording themselves doing everyday chores for data-collection firms, which then sell that footage to companies such as Tesla, Figure AI, and Agility Robotics to train their humanoids. Indian startups see this as a moment to claim a seat in the global AI value chain. The country has over 260 robotics startups, and investors are beginning to pay attention... The footage being recorded by Indian workers becomes proprietary once it leaves the country. The datasets assembled from it are accumulating on foreign servers. The foundation models trained on them are owned by foreign companies.

The NYT has an article about how startups like Handshake, Mercor, and Surge in the US are paying skilled workers to collect data on their work. 

Mercor and a handful of similar start-ups are the primary middlemen in a supply chain of “human data” that may power the next generation of A.I. As OpenAI, Anthropic and other major ventures compete to become the industry’s dominant platform, the market for premium data that has been vetted by experts is exploding…They need mathematicians to annotate proofs, lawyers to mark up briefs and professors to grade essays… To use the parlance of the industry, data labeling has moved up the “value chain,” and the start-ups that offer this service have become some of the fastest growing in Silicon Valley… The data-training start-ups see a lucrative opportunity in recreating workplaces in miniature: controlled environments in which their gig workers can evaluate and reproduce emails, memos and slide presentations in context. The information emerging from such a setup, the companies boast, will help shrink the gap between what A.I. models can accomplish and what office workers actually do from one minute to the next, as ideas and instructions flow between meetings, documents and applications…

To keep improving their models — to make them more useful, more sophisticated, less prone to hallucination and mistakes — A.I. companies heavily refine what goes into them. That’s post-training, and it includes buying data from vendors like Handshake and its competitors… Deeptune, a start-up that makes “training environments” with simulations of the software programs, like Slack and Salesforce, that many workers toggle between all day long to get their work done. The idea is to painstakingly create a mirror image of, say, an investment bank so that A.I. can observe every interaction…

It may turn out that once OpenAI, Anthropic and others have taught their models to perform a certain job, their need for more training data in that area could sharply decline. In this way, Mercor, Scale, Handshake and their peers are much like the elite freelancers they employ: making money today, but in danger of being dropped tomorrow… People sign up for data-training gigs for a variety of reasons. The main one is, of course, money… Though the labor is unpredictable and rates vary… the workers who cobble together enough shifts can generate meaningful income. People might sign up because they have been laid off, or because they can’t find enough work in their field. They might do it because they’re eager to get “A.I.” on their résumé, or because they need extra cash in retirement… Many people who contract for these companies understand that this is a short-term opportunity, a brief chance to train the models to automate jobs before they themselves are automated out of the job of training models.

I asked Claude to generate a visualisation of this market landscape, including an assessment of the Indian landscape. The numbers are clearly estimates and must be validated (though at a ballpark they appear alright). 

The unit economics of the data chain shown below for a garment worker in India is instructive. She gets roughly ₹400 a day to wear the camera (or $0.60 per hour); the startup pays the factory ₹450–500 per hour; US-based startups like Human Archive price data at $1–10 per hour; and once annotated and packaged, it sells to global robotics labs at $15–50 per hour. That is a 25–85 times markup, and every rung above the worker is owned outside India.

The graphic also shows that the vast majority of AI workers are doing the BPO equivalent, whereas the vast majority of funding is going to those building the data centres. India has 170-odd AI startups that have raised $2.6 billion in total and over 260 robotics startups, but the genuine model/product builders are a tiny set, and the majority have rebranded annotation as an AI line of business (iMerit, Objectways, Awign, Karya, Deccan AI, Human Archive, Egolab, Neo Cambrian, Humyn Labs, RoBoEra, etc.). 

It must also be highlighted that in the majority of cases in India, the data goes from the garment worker to an Indian data aggregator to a robot-brain lab in San Francisco, and comes back as a robot/humanoid. The frontier LLM labs are not in that loop. This also means that none of the emerging governance conversations about frontier models - safety frameworks, export controls, model-access negotiations - touches the mainstream data collection work being done in India. India is negotiating hard for access to frontier language models while simultaneously handing over, for ₹400 a day, the training substrate for the physical models that will actually displace its manufacturing workforce. Those are two different conversations, and only one of them is being had.

Further, as the Times article highlights, while these annotation startups are flourishing now, they may not be sustainable ventures. Once experts teach the models to do something, their services are no longer needed in the same way, and the vendors themselves need the models to keep improving to show they add value, while needing them to remain imperfect so clients keep coming back. 

I asked Claude for historical precedents and got this:

Frederick Winslow Taylor’s explicit programme, from the 1890s, was for management to “gather in all of the great mass of traditional knowledge which in the past has been in the heads of the workmen.” Skilled machinists were stopwatched; the Gilbreths filmed them with chronocyclegraphs — a literal 1910s head-camera. Workers cooperated because they were paid piece-rate bonuses to do so. The tacit craft was decomposed into instruction cards and handed to cheaper, unskilled labour. Outcome: enormous productivity gains, the collapse of the craft wage premium, a machinists’ revolt, congressional hearings in 1911–12, and Taylorism banned in US government arsenals by 1915. It took roughly fifty years and the postwar labour accord before the gains were broadly shared… 

In the 1990s American hospitals routed physician dictations to transcriptionists in Bengaluru and Chennai; it was unglamorous work, but India was good at it. That corpus is precisely what trained speech recognition. The industry peaked and then largely evaporated. Compensation to the transcriptionists: zero… most startups in this space risk meeting the same fate as the transcription companies of the 1990s.

In this context, I am reminded of the claim made by Daron Acemoglu and Simon Johnson in their book Power and Progress that the trajectory of technological progress is a political choice made by society and should not be left to corporations and technocrats. Their central claim is that the direction of technology is a social choice, not a technical destiny, and that redirecting it requires countervailing power rather than better-intentioned technocrats. 

The problem, though, is that globally, and especially due to the Trump 2.0 regime, the rule makers have surrendered agenda-setting to Big Tech and AI Labs. Closer home, India has almost no leverage over the direction of frontier AI. Instead, its leverage is confined to the terms on which its labour and data enter the supply chain, and not to bending the technology’s arc. 

In the circumstances, what can a country like India do?

Here are some thoughts for consideration. One, a statutory floor rate for training-data contribution and an industry-led collective licensing body for data work are both administratively feasible and could increase value capture (from the worker’s current share of 1-2% of the value created) without banning anything. A comparator is the model of SoundExchange (US) or PRS (UK) in the music industry, which acts as a government-designated clearinghouse that collectively licenses music, collects usage fees, and distributes royalties to creators, effectively removing the burden of individual licensing. This model would also subtly frame the market in terms of treating data as labour, and not as mere raw material. 

Second, on the regulatory side, it may be useful to revisit the DPDP Act provision that permits employers to process worker data without explicit consent under “employment purposes”. Instead, there should be purpose limitation, or restrictions on repurposing training data for other activities, and consent requirements of all involved. 

Third, public spending on AI innovation and procurement preference could be made conditional on the recipient retaining licensing rights to datasets collected from Indian workers rather than doing work-for-hire. This would frame the collection of data as an input and not a product, and industrial policy could price it appropriately. Fourth, there is the argument about extending statutory instruments like the gig worker welfare boards or the Code on Social Security present in some states (Rajasthan, Karnataka, etc.) to cover data work. It could help build countervailing power. 

But pursuing these agendas can be costly. This being a global market, prohibiting or putting too onerous terms on value capture and the entry of data into the supply chain will backfire by moving the work to Vietnam, Ethiopia, or the Philippines. Besides, for the Indian workers, already facing an acute scarcity of jobs, the choice isn’t really on offer, and ₹400 a day is ₹400 a day. There is a collective action problem here which calls for multilateral engagement through a forum like the ILO. 

But this should not mean that we sit back helplessly and allow the market dynamics to play out. Instead, before enacting any of them, there must be a public debate on the merits or otherwise of these proposed measures. What are their respective costs, and what can be done to mitigate them? What versions, if any, of these measures should be enacted? Such debates are essential to make informed and collective social and political choices.

The public debate is important since the agenda-setting process here, like with any technology change, pushes certain considerations to the forefront while also marginalising certain others. Almost always, the former represents the interests of the corporations and elite beneficiaries of the change, and the latter represents those of the vulnerable and voiceless. Therefore, such agenda-setting debates are a purely political activity, with profound social implications. 

It is also important since there is the distinct likelihood that India could spend the next five years as the world’s back office for the third time, and when the juice has been sucked out and value captured, there could be nothing left standing that India owns. 

PS: In this context of collective action problems, it is worth taking inspiration from one very impressive and encouraging breakout (which has not received the level of attention it deserves) from South Korea. It is a tribute to the maturity and wisdom of the country’s corporate and political system and the robustness of its democracy that Samsung and SK Hynix agreed to share 10% of their windfall profits from memory chip sales, with no ceiling on payouts, with their employees for the next ten years. Sample this.

Samsung Electronics... agreed last month for employees to share the chipmaker’s blockbuster profits from an AI-led boom... SK Hynix... handed employees a similar profit-sharing deal last year... Samsung is also going to give Won500mn loans at low rates to employees... Samsung and SK Hynix together control much of the market for the advanced memory chips used in AI servers. Employees at both companies are in line for average annual bonus payouts of Won600mn, which compares with a national average salary of about Won50mn... district of Hwaseong... expected to gain corporate income tax receipts of Won1tn to Won1.3tn from Samsung alone this year, an extraordinary sum for a city authority whose annual budget is about Won3.5tn.

Monday, June 22, 2026

A framework for the application of AI in public systems

I have not blogged about AI in development. One reason is that, notwithstanding several claims, I have not come across promising examples that have been successfully applied at a reasonable scale within public systems in the Indian context. 

I have blogged here, urging caution on the expectations of the impact of AI on development in countries like India. The central argument is that the binding constraint on AI impact in the public sector is rarely technical capability. It is the gap between what AI can do and what institutions are prepared to validate, adopt, and integrate, a gap shaped by the validation cost, regulatory frameworks, incumbent system stickiness, and the political economy of transition.

As an analytical framework for the application of AI in development, I can think of three lenses: the layer where the AI intervention is proposed, the nature of the problem to be addressed, and the constraints on deployment and adoption. 

In the first lens, AI can provide structured information to feed into decision-making; or directly enable decision-making by synthesising information into ranked options or risk scores, thereby augmenting human judgment; or deliver specified outputs in the form of drafts, transaction alerts, assessments, etc. 

The second lens can distinguish between complicated/complex and wicked problems. The former has many variables but is tractable with better data and modelling (crop yield prediction, supply chain optimisation, credit risk scoring, traffic signal timing, etc.), whereas the latter involves contested goals, adaptive actors, and systemic feedback (improve learning outcomes or skills, institutional reform, manage urban growth, etc.). While there are limits to AI, especially in the latter, complicated sub-problems within the former are amenable to AI. 

The third lens covers the gap between technical feasibility and deployed impact. This gap is shaped by validation cost (regulatory standards must be met), the incumbent system stickiness (when the AI solution demonstrates value that helps overcome the systemic inertia), trust and legitimacy (system or society must accept the efficacy and reliability), and political economy (overcome the entrenched vested interests). 

The three lenses are analytical dimensions, or orthogonal axes on which any individual intervention can be located. Then we have the horizontal or vertical use cases, which are a typological cut across the universe of interventions. It sorts use cases by their organisational footprint (does this sit in every department or only one) rather than analysing the nature of any single use case. 

The graphic below applies the three lenses to each item in an illustrative catalogue of horizontal/vertical use cases. 

As a prudent strategy, it may be useful to move first in low-friction domains - where there are no regulatory incumbents, where information asymmetry is high, and where the beneficiary is a motivated adopter - and use the evidence and trust built there to lower the political cost of adoption in high-friction domains. In this reading, as is the trend in the private sector, it will be some time, if ever, before vertical sectoral use cases of AI become efficacious and reliable enough to be adopted at scale. 

Based on the above, what are the most valuable and widest spanning productivity and efficiency increasing uses of LLMs by governments in developing countries (with low state capabilities)? 

I can think of two in particular - application in the generation of quasi-judicial and adjudication orders of all kinds, and in the recording of M-Book in engineering works. The breadth of their use, the stakes involved, the low current baseline, and the manageability of their deployment gaps make them areas with potentially transformative impacts. 

The first one ranges from charge sheets to orders in service matters and disciplinary cases, to regulatory and licensing decisions, assessment and adjudication by tax, land, and other authorities exercising statutory powers, including court and tribunals. 

Each of these has a standard structure - factual matrix, applicable legal provisions, consideration of arguments, findings, and operative order - that AI can populate with high-quality given structured inputs. The official provides the facts and the decision, and AI generates the legally coherent reasoning and formal language.

The value of this comes from the low baseline of generally poor quality of orders issued that suffer both from basic procedural/hygiene deficiencies and more substantive application of judgment. These deficiencies and lapses, especially of the former kind, immediately invite disputes. In fact, such poorly crafted orders become the starting point for a long series of administrative processes that manifest in disputes and harassment, appeals and litigations, that clog administrative bandwidth, lock up scarce capital, impair balance sheets, and generally waste effort and resources. 

These kinds of orders also contribute a very large share of the litigation against the government that clogs the judiciary. The major reason why governments lose such cases is the quality of orders in terms of poor drafting and adherence to basic administrative compliance and procedures. Poorly drafted orders are often set aside, not because the underlying decision was wrong, but because the reasoning was inadequately articulated. AI can have a significant impact in this area in both improving the quality of orders and even forcing systems to comply with procedural requirements.

The volume of such orders is enormous. A mid-level income tax officer may need to issue hundreds of assessment orders annually; a district collector may handle thousands of revenue proceedings. The drafting burden is a critical bottleneck in disposal rates. Further, AI-generated drafts can be trained to flag inconsistencies, missing procedural steps (e.g., failure to give opportunity to be heard), and citation of superseded legal provisions, thereby acting as a compliance check before the order is issued.

An AI application that generates the first draft of the order, for the adjudicating officer to revise and issue, can have a transformative cascading impact down the chain. The first draft can be hard-coded to serve as a forcing function to ensure compliance with a checklist of processes; the order itself could be validated, and its logic could also help with tightening the substantive exercise of judgment itself. 

On the face of it, this is a low-hanging fruit. All it requires is a library of orders of all kinds bundled together, and a template for different kinds of orders. An enterprise or team subscription to Claude or OpenAI, that has a Zero Data Retention (ZDR) configuration for their API (thereby ensuring customer data is not retained and used for training the algorithm), and which allows access to Claude through Amazon Bedrock or Azure OpenAI within a private Virtual Private Cloud, can address any privacy and security concerns on sharing internal data. 

This would be a quick deployment and superior to any indigenous or sovereign LLM model development (which should continue and could perhaps be deployed in parallel to learn and get refined). 

The latter work can be taken up in a mission mode by the National Informatics Centre or a public think tank (say, the National Centre for Good Governance, NCGG) by deploying a team of competent experts. It can consolidate the library of orders and develop standardised LLMs for a few high-volume and value use cases - disciplinary cases, GST and income tax adjudication orders, and orders on a few categories of land claims settlements. They could work with a few state governments, the CBIC (state indirect tax departments), and the CBDT.

The other high-value and promising application can be in the recording of the Measurement Book (or M-Book) for engineering works - roads, buildings, canals, drains, electricity lines, etc. It records the quantity of work done, certified by a junior engineer, against which payment is authorised. It is also one of the most fraud-prone documents in public works administration, susceptible to over-measurement, fictitious entries, and post-facto alteration. Besides, it is also a procedurally burdensome activity. Its importance arises from the fact that it is the basis on which contract payments are made. 

For certain categories of works, AI-assisted quantity and quality measurement from photographs and drone-based progress tracking and volume estimation can be very reliable. Smartphones equipped with computer vision applications can estimate dimensions and quantities from site photographs (lengths of pipes laid, areas of surface plastered, volumes of earthwork excavated). Computer vision models can be trained to assess the quality of construction work from photographs by detecting visible defects in concrete, checking alignment of masonry, and identifying sub-standard finishing. For large civil works (earthwork, embankments, reservoirs, large buildings), drone imagery processed through photogrammetry software can generate accurate volumetric estimates and three-dimensional progress models.

At the very least, AI-assisted photographs and drone outputs can be used to validate M-Book entries, surface quality concerns and other discrepancies. Once M-Book entries are digitised, AI can generate contractor bills from certified measurements, cross-check against contract rates, identify arithmetic errors, flag unusual patterns (e.g., a sharp increase in claimed quantities near bill submission deadlines), and route completed bills for authorisation. Apart from improving the quality and increasing accuracy of payments, the time savings and reductions in delays in payments will be significant. 

The examples of order-drafting and M-book opportunities are likely to be particularly appealing for officials because they also significantly reduce their workload and drudgery. They are deployable at medium friction, directly legible to senior officials, and capable of generating rapid, measurable impact that builds institutional confidence in AI tools. 

Another area of promise is the categorisation or triaging of cases in several public domains. 

The Ken has an article on the application of AI by courts in India. A promising application of AI is in the categorisation and triaging of cases.

Half of India’s backlog—25 million cases—could disappear quickly if AI were used not to decide cases, but to triage them. The opportunity lies in the huge volume of matters that are effectively already dead on arrival: cases filed beyond limitation, petitions missing mandatory disclosures, matters rendered moot by changes in law, or disputes where precedent has already settled the question, leaving only a narrow technical point to close. Courts identify these only in occasional manual clean-ups—slow, inconsistent, and impossible to scale. 

An AI system, by contrast, could scan filings in bulk, flag statutory defects, detect outdated or defective pleadings, classify time-barred matters, and surface cases appropriate for summary disposal. These could then be bundled and placed before judges for quick orders, clearing the undergrowth so courts can focus on disputes that actually require judicial time.

Like with court cases, triaging of outpatient (OP) cases in primary health centres (PHC), community health centres (CHC), district hospitals, and medical colleges is an area where AI can play a potentially significant role. The daily OP load in these hospitals (at least the better ones among them) is multiples of what a doctor can manage, leaving them overburdened and stressed. The result is inefficient use of the doctor’s time, inadequate diagnosis time, incorrect diagnosis, wrong OP referrals, and so on. 

In all these hospitals, OP cases come to doctors with limited or no triaging. An AI-based triaging application where the symptoms are entered at the OP-registration, nurse and doctor-level, can significantly improve work conditions, increase hospital productivity, and enhance the quality of treatment. 

Similarly, citizen grievances or consumer complaints received in any office, especially public-facing ones, can be triaged for routing them to the right desks/officials, escalating to senior officials, analysing repeat complaints, and so on. 

Triaging is already one of the early emerging successes of AI, with examples like Bank of America’s digital assistant “Erica”, which handles billions of client interactions and has reduced call centre volumes by 40 per cent.

In general, the judiciary’s case load management is an area where, in theory, AI applications can have a transformative impact, especially in categorising cases, transcription of witness statements, and the preparation of orders. The Ken article has a good description of an AI application in a courtroom in Kerala.

A witness spoke; an AI-powered tool listened. A clean, searchable transcript appeared in real time—punctuated, structured, permanent. Stenographers were not clambering to catch up, litigants were not begging for readable copies. This split-screen view—one courtroom running on memory, the other on machine comprehension—is not a metaphor. It is India’s judiciary in 2025: a system where 19th-century workflows and modern AI systems operate side by side, neither quite replacing the other. Earlier in 2025, the Kerala High Court issued an office memorandum, making AI-assisted live transcription mandatory across all district courts… 

“Today, the AI-transcribed witness testimony is uploaded as soon as the proceedings are completed. Earlier, the witness testimony would be handwritten by the judge, and a party’s lawyer would apply for a readable copy and seek adjournment on that basis,” said Joseph Rajesh, the IT Registrar of the Kerala High Court. “All this time delay is now cut down to nil.”… A court order that once took four or five hours to type can now be generated in under an hour. A witness deposition that once required handwritten dictation, shorthand transcription, and final review can be captured in a single digital stream. 

The article also has this graphical exploration of the various possible use cases for AI within the judiciary.

There can also be daunting practical challenges to its adoption

Consider defect detection—the clerical step where filings are checked for completeness. In Delhi’s Tis Hazari courts, the rules are so complex that Nyaay’s engineers joked: If we can crack defect detection here, we can crack it anywhere. The joke holds true. AI had to be customised court by court. And transcription? In many urban courts, there are stenographers, but in district courts, they are scarce. AI fills a vacuum, but only if the court has electricity, microphones, and a judge willing to trust the output.

So how can these AI applications be deployed?

For all these use cases, the development of robust AI applications that can be deployed across institutional levels nationwide is an industrial engineering endeavour. It requires diligent, long-drawn, and high-quality problem-solving, including pilots (or beta testing), iterative adaptation and refinement. It cannot be achieved by an experiment undertaken by a district or state-level entity (or even as ad-hoc officer-driven initiatives within central government departments), as is often the case today. 

The Ken article cited above points out the problems of not having one entity in charge of the creation of these national public goods.

The downside is that every state is effectively running its own lab experiment: non-standard, improvised, and dependent on the priorities (or disinterest) of whoever happens to be leading the High Court. Some innovate aggressively. Some slow-walk. Some judges independently approach AI companies, but without formal approval, nothing can enter the courtroom workflow… High Courts and district courts continue testing tools piecemeal. Language models multiply while vendors proliferate and workflows diverge… India could end up with 25 distinct judicial AI ecosystems, none interoperable, standardised, or scalable nationally… 

India’s AI push is happening largely outside the Supreme Court’s flagship e-Courts project—the initiative meant to modernise the judiciary. Phase III of the programme still focuses on the basics: creating digitally readable records, enabling e-filing, and automating service of summons. A February Press Information Bureau note pegs the Phase III budget at Rs 7,200 crore, but only a little over Rs 50 crore is reserved for “integration of AI and blockchain technologies across High Courts”. In effect, the national blueprint is still laying the plumbing while the states are already experimenting with smart faucets…

If India manages to fuse Kerala’s mandate with Karnataka’s case clustering, Delhi’s defect detection, the Supreme Court’s governance frameworks, and the entrepreneurial urgency of tools like Adalat and Nyaay, the judiciary could undergo a systemic leap.

It would require a dedicated team at the national level (government department, Supreme Court, CBIC/CBDT, etc.) engaging single-mindedly on the endeavour - formulating the problem, collecting and cleaning data, developing models for different levels of government, undertaking pilots with close oversight, iterating with tight feedback loops, documenting processes, and scaling solutions gradually.

The models required for lower courts, High Courts, and the Supreme Court would vary and must therefore be customised, just as those required for hospitals of different kinds, tax and land adjudicating and appellate officials at different levels, and disciplinary authorities for different kinds of functional entities. In each case, there would be a need to do rigorous pilots to be able to refine and finalise robust enough solutions that can be deployed at scale.

Monday, June 8, 2026

Some thoughts on the likely impact of AI in countries like India

This post will argue that the long-term impact of AI innovations on the typical household’s daily life in a developing country like India will be far smaller than the discourse suggests. Much the same could be said for basic human development and public services delivery in general. 

While there will be some overlap, I make the distinction between the macroeconomic impact in terms of automation and resultant job losses, and the material impact on the lives of people. Also, while much has been written about the former, this post will focus on the latter. 

As a note of caution, as is the case with any such debates, evidence and data to substantiate claims are hard to come by. So there is a lot of judgment in the argument below. And I’ll only be too happy if the judgment is proved wrong. 

The argument rests on four observations.

First, the bulk of the economy and daily life in these countries lies in domains where AI can barely reduce costs. The two main production modes, agriculture and manufacturing, are unlikely to be significantly affected (see the third point). The main products that we buy - homes, vehicles, jewellery, and consumer durables - and the everyday services we buy - haircuts, repairs, domestic help, transport, retail, eating out, etc. - are dominated by physical inputs of materials, energy, land, labor, and transport. 

AI can reduce logistics and coordination overhead costs, which form a marginal share of the total production cost of these goods and services. The share of any of these prices that AI can potentially influence is likely a fraction. Of India’s roughly 470–565 million workforce, around 85% is informal, ~45% is in agriculture, and the formal IT/BPM/GCC sector employs only about 5.8 million people. The set of workers directly exposed to AI displacement is small - plausibly 5–8% over a fifteen-year horizon - and the consumption basket of the median Indian is composed almost entirely of goods and services whose prices AI is unlikely to meaningfully alter.

Second, the obvious rebuttal, that AI’s impact will come not through cost reduction but through dramatically expanded access to healthcare, education, finance, and government services for India’s 800-million-plus internet users, runs against some hard evidence. The last twenty years of EdTech, Medtech, SkillTech, and AgTech across India and comparable low-income contexts amount to a rich natural experiment. The result is essentially zero examples of even significant, much less transformational, district-scale impact in any of these domains. 

Despite massive amounts and efforts on EdTech, aggregate learning outcomes, as measured in the likes of ASER scores, have hardly moved. In fact, the share of Class V children in rural India who can read a Class II text has actually declinedover much of the EdTech era. Like Diksha for school education, eSanjeevani logs vast consultation volumes with no measurable population health effect. It is hard to find meaningful signatures of MedTech in primary or secondary healthcare. In skill development, a succession of schemes has trained tens of millions with dismal placement outcomes, including using technology extensively. Digital Green and dozens of AgTech pilots have generated good papers, but no district has measurably transformed agricultural productivity because of them. Outside India, the picture is similar - Kenya's M-Health pilots, Brazil's rural EdTech, Indonesian AgTech apps. The hit rate on "scaled, measurable, transformational" is essentially zero.

The same could be said about most areas of public services delivery - primary health care and school education; municipal government services like tax assessment, building permissions, utility service connections; and the services of regulatory agencies. While there are pilots and small slivers of some success, aggregate impacts across all these realms attributable to digital technologies, notwithstanding numerous and repeated initiatives, have been minimal. 

Third, the reason this pattern matters for AI is that it points to perhaps a wrong diagnosis being made about why previous efforts, including using technology, failed. The standard story that “the tech wasn’t good enough yet” implies AI will finally break through because it is qualitatively better. However, I’m inclined to argue that the diagnosis is wrong. 

The binding constraint in these domains has never been information quality or delivery. The child in the village school does not fail to read only because she lacks access to a good pedagogical sequence; she fails also because she has accumulated large antecedent learning lags, the teacher is over-burdened, indifferent, or absent, the system has no consequence for non-learning, she is hungry and has chores in the evening, her parents cannot reinforce at home, and so on. A perfect AI tutor in Hindi and Math does not change those facts.

The villager seeking treatment doesn't suffer only because no one can diagnose her condition; she suffers also because the doctor is indifferent or isn't there, the PHC has deficient diagnostics or medicines, transport to the CHC costs a day's wage, and the prescribed drug regimen is incompatible with her work and food situation. It is not only the lack of plumbing knowledge that holds back the aspiring plumber. Instead, he also lacks an apprenticeship network, a credential the contractor trusts, and tools. The farmer is also constrained by the inertia to change long-standing practices, water, fertiliser subsidies, and the price he gets from the mill, not only by ignorance of best agronomic practice or market information.

In all four cases, the recipient is operating in a bound system where information is, at best, the fifth-binding constraint. Solving the fifth-binding constraint produces no visible improvement because the first four still hold. Even with this constraint, the ability of AI to make significant improvements at scale in these difficult contexts is questionable. More than two decades of the internet and digital technologies have made little or no impact on actual outcomes, except for a few oft-repeated pilots. 

This is exactly what the vast majority of development economics literature has been telling us for two decades, and it’s why “information-delivery” technologies have a flat impact curve regardless of which generation of tech is doing the delivery. AI is, fundamentally, a much better information-delivery technology. By the logic above, it should be expected to have roughly the same impact profile - better demos and pilots, but similar real outcomes - unless something about AI breaks the pattern.

Having said this, it is also logical to argue that AI can address and relax all these constraints just enough to enable outcomes that, while not the best, are far better than those achieved now. While appealing and comforting, I am not inclined to agree. 

Fourth, the genuine exceptions exist but are narrower than the transformational claims that mainstream discourse suggests. AI is plausibly different in three specific places: supply-side augmentation that flows through existing institutions (Qure.ai’s tuberculosis screening integrated into state programs is the cleanest example); voice and vernacular interfaces that break the literacy ceiling text-based apps could never cross; and AI built on top of India Stack to alter citizen-state interactions. These are real but bounded effects, not transformations.

So what’s the final assessment?

AI radiates a wide beam of capability (advice, diagnosis, tutoring, prediction); the beam hits a wall of thick structural constraints, each labelled with the human reality it represents and the domain it blocks (say, education, health, livelihood, farming). Only a thin sliver of information makes it past the wall to reach the villager with the phone below. AI delivers information, but not transformation. 

The mainstream discourse on AI is built on the worldview that assigns outsized importance to knowledge-based services over the production of goods. This is a real blind spot that obscures the reality of the vast majority of non-AI-influenced interfaces and interactions in the daily lives of most Indians. 

In conclusion, I’ll stick out my neck and argue that AI’s footprint on median Indian life will look much more like the mobile phone’s did - ubiquitous, individually useful, but hardly transformational on people’s daily lives. The mobile phone did not move India’s Human Development Index; it moved convenience and communication. AI is on track to do something similar: meaningful at the margin, but layered on top of structural conditions it cannot itself relax. 

The global discourse, calibrated on knowledge-work economies where AI strikes the dominant production input, badly overstates the implications for a country where physical and institutional constraints set the floor. 

Unfortunately, like with the internet and digital technologies, this is likely to be a costly distraction for development. Instead of working to get the plumbing right, it is likely to displace resources and efforts towards getting AI solutions to address these problems. I have blogged herehere, and here on this.

Wednesday, June 3, 2026

Deploying public finance to derisk private capital in innovation and infrastructure

As the heading suggests, this post points to a few thoughts on the challenge of deploying public funds to derisk and crowd-in private capital into those areas of innovation and infrastructure that are not attractive enough for commercial capital. 

I have blogged about public funding of innovation here and here, and this working paper is about public funding of infrastructure projects. The additionality with public finance arises from its risk-tolerantpatient, and concessional nature. 

An urban water supply or sewerage, or an industrial bulk water supply, or an electricity distribution, or a solid waste management, or a streetlight energy saving project, or a mass transit project, will not attract commercial capital on its own. Similarly, a fledgling startup making transceivers, or display and camera modules, or high precision resistors/inductors/capacitors, or compressors, or brushless DC motors, or anode materials, or aluminium extrusions, or designing a narrow band IoT chip or some other mid-value chip, will generally struggle to attract risk capital. 

But they can be derisked by blending with a layer of public finance. The challenge is how to do this derisking effectively. Specifically, the challenge is how public funds can be channelled to derisk these projects or sectors. 

This is less of a problem with grant funding involving smaller amounts, which is simple enough to be done through public entities. In infrastructure, such grants come in the form of viability gap financing (VGF), and in innovation, they are given to those in TRL 1-5/6 stages. The problem lies in the deployment of risk capital (debt and, especially, equity and structured instruments) by public entities. 

Given its administrative inflexibility and constraints, direct deployment of funds by the government itself is not only inefficient but also creates incentive distortions.

In the circumstances, the commonly suggested option is arms-length financing through Development Finance Institutions (DFIs). But this approach is seriously hampered by unreasonable expectations (about returns or at the least capital preservation) arising from a deficient understanding and acknowledgement that de-risking, by its very nature, entails a strong likelihood of losing money. It would involve investing in projects that would not attract commercial investors, by insuring for the additional risk borne by them. 

Further, even with an arm's-length institutional structure, a fully government-owned entity is subject to constraints that prevent efficient deployment of funds. 

The response to this problem has been to build institutional structures by partnering with private investors, even having majority private investors. This, it has been argued, will free them from the fetters and requirements faced by public entities. 

However, India’s disappointing experience with infrastructure DFIs starting from IDFC, IIFCL, and NIIF, as documented in detail here, raises questions about this response. In all these cases, instead of complementing private capital, the DFI has ended up competing with private investors in their choice of investments. Instead of funding those risky sectors, the DFIs chase the derisked sectors like power generation, renewables, transmission, highways, ports, and airports. 

In this backdrop, a commonly cited option, especially in the context of innovation financing, is to transfer public funds to commercial investment vehicles (or the Fund of Funds, FoF, strategy) and let the latter manage those investments. This looks great in theory, insofar as it aligns incentives and brings in private sector efficiencies. 

But it has one problem. The private investors will be primed to invest, at best, in those marginally risky projects rather than in genuinely risky projects (or sectors) that sorely need public finance to derisk them. So, instead of deriskingprojects or sectors, public funding will do returns amplification for private capital. 

Here, a big problem, a market friction, is the absence of a pipeline of such risky projects and sectors that investors can draw from. Their search costs are a significant enough deterrent for investors. In contrast, commercial investors have access to a widely known pipeline of investible projects or innovations. 

It is also the case that the envelope of such risk capital available to fund infrastructure and innovation is much smaller than the envelope of investible projects. Therefore, there is little incentive to go beyond the confines of the mainstream and search out and fund the riskier projects. 

In the circumstances, I can think of three options for the deployment of public funds such that we are able to realise its additionality, and not compete with and crowd-out private capital or end up being leveraged primarily for returns amplification.

1. Invest in FoFs, but with sharply defined funding mandates, almost prescribing the specific nature of projects to invest in, at least a part of their portfolio. However, this can be unsettling for the commercial investors and may turn away the GPs who sponsor the fund from accessing public funds. 

2. The DFI could announce its offering as a set of financing instruments that meet the derisking objective. They could include credit guarantees (in the form of first-loss buffers), longer tenor, lower interest rate or hurdle rate, lower liquidation preference and a lower charge on the waterfall, subordinate debt, and so on. The DFI should market these instruments and possible investment projects to commercial investors. 

3. The DFI could co-invest with private investors. This would entail the public entity scouting the project or entrepreneur, doing due diligence on it/them, and then shopping it to commercial investors with an offer of an attractive enough derisking financing layer. This would also require an acknowledgement of the fact that the role of public finance is to derisk and not maximise returns. This is perhaps the most ideal approach, one which mature entities like NIIF in infrastructure finance ought to be mandated to do.

The second and third options require highly capable and incentive-aligned institutions. Given weak state capability, that’s a demanding requirement. It is for this reason that even in developed countries, risk capital funding in infrastructure and later-stage innovations is largely deployed through FoFs, notwithstanding its aforesaid failings. 

But this reality should not be a reason to ignore the failings of the FoF strategy and step away from pursuing the second or third options.

Tuesday, May 26, 2026

The missing link in India's FAR market - a trading platform

I blogged here on nine low-hanging fruits in urban planning in India, and here on the actual use of land value capture instruments in India. This post will discuss the Floor Area Ratio (FAR) and Transferable Development Rights (TDR) transactions in Indian cities. 

Fundamentally, FAR, beyond a certain limit that comes with property rights, can be purchased at a defined rate (up to the master plan limit). Further, when land is given up for a public purpose, instead of receiving financial compensation, the landowners can get additional FAR that can be traded (as TDR) and availed in certain predefined locations. Finally, the TDRs can be transacted through an institutionalised platform. 

The table below shows the status of FAR and TDR transactions across Indian states.

Clearly, while all states have the requisite policy and legal frameworks in place, the actual implementation (in terms of FAR sold and revenues realised, and TDRs issued and transacted) has been limited, except in Maharashtra, and to some extent in Hyderabad. 

Mumbai is the most complete illustration of FAR and TDR. It allows for purchaseable FAR (or premium FSI) of 0.5 to 0.84 over the base FSI, depending on the road width abutting the plot (nothing below a road width of 9 m). It is sold at 35-50% of the Ready Reckoner Rate (RRR), varying for residential, industrial, and commercial uses. The revenues are shared between the Brihanmumbai Municipal Corporation (BMC), the state government, and the State Road Development Corporation (MSRDC). BMC itself has so far earned over Rs 14,000 Cr from premium FSI sales. 

TDRs are issued (in the form of a Development Rights Certificate, DRC) by BMC for defined instances where land is surrendered for public purposes at the rate of 2.5X for the island city and 2X for suburbs against the surrendered land. It can be used across the municipal corporation limits (with some specific exclusions) in the range of 0.17 to 0.83, depending on the road width (cannot be used below 9 m roads). The TDR generated is indexed at the RRR, and an equivalent floor space is allotted at the receiving plot. Since April 2026, Mumbai has moved into a fully electronic trading platform for TDRs, and manual record keeping has been dispensed with. 

An electronic trading platform is essential for unlocking the full potential of TDRs. It is particularly important since it significantly enhances TDR liquidity, enables efficient market clearing, brings transparency and certitude for landowners, and also prevents fraud and abuse. Mumbai is the only place with a blockchain-secured, workflow-automated e-trading platform that allows for buyer-seller bidding and matching. 

A major reason for TDR trading failing to operate across states is the manual and opaque nature of the records and the problems with trading them. It forces cities to impose restrictions like limiting TDR to the same locality or zone, thereby sharply diminishing liquidity and deterring land owners from accepting TDRs. Until a state moves to a properly dematerialised exchange, secondary-market price discovery stays opaque, and TDR rates trade at significant discounts to face value. This is the main reason why so many landowners say, “We want money, not land.”

An illustration of its value comes from the experience of Hyderabad’s TDR Bank, which is currently only a depositary and coordination platform, though integrated with its building approval system. In the absence of a trading platform and the associated transparency and liquidity (and also the misguided unlimited FAR provision), in order to trigger TDR transactions, the state government was forced to come up with a government order (GO Ms No 95, March 2026) to mandate TDR utilisation for high-rise buildings. 

All told, even highly urbanised states like Tamil Nadu, Gujarat, AP, Telangana, Karnataka, and Haryana have struggled to realise the potential of TDRs due to thin secondary markets that deter efficient market matching and price discovery. 

An electronic TDR trading platform can significantly address these constraints and simplify TDR issuance by bringing complete transparency and allowing its trading across the city, thereby increasing liquidity and raising its attractiveness and credibility for land owners (both sellers/owners and buyers). 

How do other countries and cities undertake TDR trading?

The table below captures how TDR trading happens across a few cities/countries. 

Brazil's CEPAC is the gold standard for an exchange. It is the only system globally where development rights are explicitly structured as securities. As an aside, this also makes Mumbai’s success especially remarkable. CEPACs are auctioned by the Federal Bank of Brazil and can only be used in designated urban operation areas; they must be authorised by the CVM (the Brazilian equivalent of the U.S. SEC) to be traded on the B3 stock market, and have raised over USD 2.7 billion across 15 years from Faria Lima and Água Espraiada Urban Operations, with funds reinvested in transit, roads and parks. In Rio's Porto Maravilha, all CEPACs were wholesale-auctioned to Caixa Económica Federal, which financed all infrastructure improvement and 20 years of service costs without further public outlay. 

In this context, I have a co-authored working paper here, where we propose the establishment of a similar exchange for trading FARs with mechanisms similar to those in Mumbai. 

So what should be the design for a trading platform for TDRs? 

The wide diversity across states and cities, coupled with the weak state capabilities, necessitates a qualified approach in designing TDR trading platforms for Indian cities. Besides, the TDR is a local commodity, whose trading is inherently local. A three-level design may therefore be prudent. 

At the first level, as discussed here, there’s a need for a standardised national protocol layer, authorised by SEBI and MoHUA, and operated through a SEBI-approved entity. This layer could do four things: (a) define a Development Rights Certificate (DRC) standard with a unique national ID and QR code; (b) provide a centralised dematerialised depository with its eKYC, escrow, etc.; (c) set eligibility, disclosure, and dispute-resolution norms; and (d) supervise a national audit trail. This layer would enhance the credibility and efficiency of the platform. 

Each municipal authority (BMC, GHMC, BBMP, CMDA, APCRDA, GMDA, DDA) could operate a city-bounded marketplace on the common rails, and trading within a master-planned area only. It could borrow from the Mumbai e-TDR, which has a single nodal bank (SBI) and one blockchain-secured ledger. This would be supported with a municipal TDR bank inside each metro exchange. It could be designed as a public buyer-and-lender of last resort that holds inventory, sets a price floor, and even guarantees loans collateralised by DRCs. This solves the landowner objection of “we want money, not land” by giving every certificate-holder a guaranteed cash exit at a reserve price.

This three-layer structure also matches India’s existing regulatory plumbing in trading platforms. Mutual funds, government securities, and electricity markets are organised with a national legal and depository spine, with state-specific trading layers on top. 

The smaller cities in each state could have a pooled market, administered by an appropriate state government entity. Or it may be useful to start with the larger cities, and expansion to be done based on the emerging feedback.