<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator><link href="https://yashtambawala.com/feed.xml" rel="self" type="application/atom+xml" /><link href="https://yashtambawala.com/" rel="alternate" type="text/html" /><updated>2026-09-01T17:07:32+05:30</updated><id>https://yashtambawala.com/feed.xml</id><title type="html">Yash Tambawala</title><subtitle>I&apos;m Yash Tambawala, a technology professional based out of Bengaluru.</subtitle><author><name>Yash Tambawala</name></author><entry><title type="html">India’s Industrial Policy Buys Activity, Not Capability</title><link href="https://yashtambawala.com/2026/09/01/indias-industrial-policy-buys-activity-not-capability.html" rel="alternate" type="text/html" title="India’s Industrial Policy Buys Activity, Not Capability" /><published>2026-09-01T09:00:00+05:30</published><updated>2026-09-01T09:00:00+05:30</updated><id>https://yashtambawala.com/2026/09/01/indias-industrial-policy-buys-activity-not-capability</id><content type="html" xml:base="https://yashtambawala.com/2026/09/01/indias-industrial-policy-buys-activity-not-capability.html"><![CDATA[<p>C2i Semiconductors was founded in Bengaluru in June 2024, approved for support under the Design Linked Incentive scheme five months later, and last week agreed to be acquired by Infineon. It designs power management chips for AI data centres, where volatile GPU workloads place unusually hard demands on the power delivery stack. It raised from Yali Capital and then a $15 million round led by Peak XV, taking private funding to roughly ₹170 crore. Its first silicon had only recently come back from the fab.</p>

<p>The reaction was immediate and predictable. An Indian chip company, backed by public money, sold to a foreign multinational. Taxpayers had underwritten a German company’s R&amp;D.</p>

<p>The money argument doesn’t survive contact with the numbers. DLI reimburses up to half of eligible expenditure with a ceiling of ₹15 crore per application, and disbursement is milestone linked, so against ₹170 crore of private capital the public contribution to a company at C2i’s stage was a small fraction of what it spent. C2i was not living on government money and did not sell because government money ran out.</p>

<p>It sold because there is no domestic customer for AI data centre power IP at Infineon’s scale, no deep pool of Indian late-stage capital willing to fund a decade of silicon toward a listing, and no Indian company with the balance sheet and global distribution to have entered a competing bid. Those are gaps in capital markets, customer markets and industrial depth. Reimbursing design expenditure touches none of them.</p>

<p>So the interesting question is not whether C2i broke faith. It’s what the Indian state asked for in return for its money, and the answer is close to nothing.</p>

<h2 id="one-clause-and-what-it-wasnt">One clause, and what it wasn’t</h2>

<p>DLI’s only real protection is a requirement that beneficiaries maintain domestic status for three years after approval. C2i sits inside that window, which makes the near-term question administrative: enforce it, waive it, or accept a transaction structure that satisfies the government’s own reading of domestic status. Roughly twenty other DLI-backed companies are watching, because whatever happens here sets the precedent the rest of the cohort will plan around.</p>

<p>But notice what that clause is. It is a restriction on ownership for a fixed period. It is not a requirement to do anything. It does not oblige C2i to license its IP to Indian firms, to keep its design team in India after a sale, to train engineers who then disperse into the wider ecosystem, or to give the state any equity that would let the public balance sheet participate when an acquisition happens.</p>

<p>The clause restricts without requiring, which is the worst of both worlds. Enforce it strictly and you obstruct exactly the successful exits that make investors willing to fund the next hard-tech company. Waive it when an attractive offer appears and it never meant anything. Either way, India ends up with no claim on the capability it helped create.</p>

<h2 id="every-layer-of-abstraction-hides-the-same-trap">Every layer of abstraction hides the same trap</h2>

<p>There’s a pattern in how software engineers describe what happens when tools get better. An IDE removes the need to remember how compilation works. A framework removes the need to understand the infrastructure underneath it. An AI coding assistant removes the need to type the implementation yourself. Each layer makes you faster. Each layer also lets you become productive before you become competent, because the tool absorbs exactly the difficulty that would otherwise have forced you to learn something.</p>

<p>The danger isn’t the abstraction. It’s that the gap between productive and competent is invisible for as long as the abstraction holds. You only discover what you don’t actually know at the moment the layer fails: the framework hits a case it wasn’t built for, the IDE can’t explain the error, the assistant generates something that compiles and is wrong. By then you’re the one holding the problem, with none of the underlying understanding the tool had been quietly standing in for.</p>

<p>An industrial subsidy is the same kind of layer laid over a country instead of a codebase. It lets a sector look productive, exports rising, factories running, chips taped out, without the country having done the harder thing underneath: learning to design the next node, financing the next decade of a firm’s life, building the customer base and capital markets that let a company grow up at home instead of getting acquired abroad. The subsidy absorbs the difficulty. It does not transfer the difficulty into anyone’s hands.</p>

<p>Nobody at a ribbon cutting can tell productive from competent by looking. The distinction shows up only when the abstraction runs out, when a company needs something the scheme was never designed to provide and finds nothing underneath. For C2i, that moment was a term sheet from Infineon. For Indian handset manufacturing, it’s a value-addition number stuck for years despite record output.</p>

<p>So the question worth asking of any industrial policy is not whether it makes a sector productive. Almost all of them do, that’s the easy part. It is whether the policy forces competence to form underneath.</p>

<h2 id="capability-as-hope-indias-own-evidence">Capability as hope: India’s own evidence</h2>

<p>DLI is a small instrument inside a much larger family. The Production Linked Incentive programme spans fourteen sectors with an outlay approaching ₹2 lakh crore, and its showpiece, mobile phone manufacturing, is now far enough along to be judged on results.</p>

<p>On productivity, PLI delivered. Mobile phone exports rose from around ₹27,000 crore in FY20 to more than ₹1.2 lakh crore by FY24. Production under the large-scale electronics PLI crossed ₹5 lakh crore by mid-2024. Ninety-nine percent of phones sold in India are now made in India. India became a serious node in the global handset supply chain inside four years, which almost nobody predicted in 2020.</p>

<p>On competence, the picture is different. Domestic value addition in mobile manufacturing reached 23 percent in FY24, against an original ambition of 35 to 40. The government’s own more recent figure puts it at 25 to 28 percent as of this month, so the number is moving, but the distance to target has barely closed in years of trying. Raghuram Rajan and co-authors pointed out in 2023 that as imports of assembled phones fell, imports of components rose sharply: printed circuit boards, displays, cameras, batteries, semiconductors. India replaced importing finished phones with importing the parts of phones. That is a real gain in jobs and trade balance. It is the abstraction holding. It is not an electronics industry underneath it.</p>

<p>Two design details explain why the layer never collapsed into learning.</p>

<p>The first is where the money went. The largest beneficiaries have been Foxconn, Wistron and Pegatron, all Taiwanese contract manufacturers assembling for Apple, along with Samsung. That is not a scandal, attracting global manufacturers was an explicit goal, and their presence built supplier networks and trained workers that would not otherwise exist. But it clarifies what the instrument is: a payment for manufacturing activity physically located in India, made largely to firms whose core technology and customer relationships remain elsewhere. On the domestic side, an independent assessment by The India Forum found that four fifths of Indian applicants in mobile manufacturing failed to meet their thresholds at all, with only one domestic firm clearing both the investment and sales bar. The scheme was better at renting productivity from abroad than at building competence at home.</p>

<p>The second is that value addition was tracked throughout the scheme and never made a binding condition of payment. The government measured it, reported on it, and did not make the money depend on it. Payment triggered on incremental sales, so incremental sales is what the scheme bought. The harder problem was never the price of the money.</p>

<p>The government now appears to have absorbed the lesson. The next phase of smartphone incentives is reportedly being designed to link payouts to domestic value addition targets rather than treating value addition as a monitored but non-binding outcome. That is a deliberate attempt to force the abstraction to collapse earlier, closer to when the money is spent rather than a decade later when a policy review finds the gap. It is both the right correction and an admission about the first version.</p>

<p>The fair rebuttal, made by Ashwini Vaishnaw, is that value addition takes time, and that China needed roughly 35 years to reach around 38 percent domestic value addition on iPhones while India reached comparable levels in five. True, and it should temper the criticism. But it proves the argument rather than answering it, because China’s 38 percent was not the passive result of waiting. It came from joint venture requirements, technology transfer conditions and local sourcing mandates that forced learning into domestic firms whether they wanted it or not. Time alone doesn’t make an abstraction collapse into competence. Something has to force it.</p>

<p>There is a longer Indian record here too. Before PLI came M-SIPS, offering capital subsidies of 20 to 25 percent to electronics manufacturers, with over ₹10,000 crore in incentives approved by the time it closed in 2018, much of which never materialised beyond paper. The instruments keep changing. The habit of paying for the productive layer and hoping competence follows underneath it has outlasted all of them.</p>

<h2 id="capability-as-condition-what-everyone-else-did">Capability as condition: what everyone else did</h2>

<p>The countries India cites as models built the collapse into the policy itself. They made competence the entry fee.</p>

<p><strong>Taiwan refused the turnkey plant.</strong> In 1976 ITRI signed a technology transfer contract with RCA for a 7-micron CMOS process, on the order of four million dollars, and insisted on a full manufacturing transfer rather than a working factory handed over ready to run. A turnkey plant is exactly the kind of abstraction a country can hide behind indefinitely: it produces chips without anyone inside the country understanding how. Taiwan refused it. It sent an initial cohort of 19 engineers to the United States for intensive training across design, process, verification and equipment handling, and then rebuilt the process at home rather than importing the output. By late 1977 ITRI’s demonstration fab in Taiwan was running at 81 percent yield, better than the RCA plant the process came from. Between 1976 and 1980 ITRI spent around $120 million acquiring foreign technology, then began handing it to the private sector. UMC was spun out in 1980 with the upgraded fab and its staff. TSMC followed in 1987 with fabs, equipment, process technology and 98 people, all of whom already understood what they were operating.</p>

<p>What that bought showed up years later, in a negotiation. Philips initially wanted half of TSMC in exchange for its technology. ITRI talked it down to a 27.5 percent cash investment by demonstrating that Taiwan had already mastered parts of what Philips was offering. That is the payoff: not independence from foreign partners, but leverage across the table from them, because you can no longer be sold something you already know how to build.</p>

<p><strong>Japan forced five rivals to share the hard part.</strong> MITI’s VLSI project ran from 1976 to 1980 on roughly ¥70 billion, about ¥29 billion of it public. The condition of funding was that five bitter competitors, Fujitsu, Hitachi, Mitsubishi Electric, NEC and Toshiba, staff a joint laboratory and divide the microfabrication problem between them, rather than each quietly buying a shortcut and calling it proprietary. It produced over a thousand patents, around 16 percent of them joint inventions filed by engineers from rival firms. When it dissolved, the equipment was split among participants and the knowledge went home in people’s heads rather than staying locked inside one company’s black box. Japanese firms then took over 64K and 1M DRAM.</p>

<p><strong>Korea made the exam un-gameable.</strong> Support to the chaebol during the heavy and chemical industry drive carried export performance requirements, and firms that missed them lost access to subsidised credit. An export target is peculiarly hard to fake, because the examiner is a foreign buyer with no stake in Korean industrial policy. You cannot satisfy it by relabelling imports or lobbying the ministry that set it, the way a firm can quietly satisfy an incremental sales target at home. Economists studying the programme’s long-run effects have since found that subsidised firms didn’t just sell more, they showed measurably higher productivity and better post-subsidy export performance than firms that weren’t subsidised, which is the closest thing to direct evidence that the discipline, not just the money, was doing the work. This is precisely what the mobile PLI declined to do when it monitored value addition without attaching a rupee to it.</p>

<p><strong>China ran the experiment both ways, which settles the argument.</strong> The Special Economic Zone era from 1980 looks like an activity scheme on the surface, tax holidays, land, market access. But the incentives sat inside a structure with conditions: joint venture requirements, technology transfer as the price of entry, rising local sourcing obligations. Over three decades that forced proximity moved Chinese firms from assembling other people’s products to owning large parts of the value chain in electronics, batteries and solar. The abstraction was made to collapse early, on purpose, while the state still had leverage to insist on it.</p>

<p>The semiconductor Big Fund, launched in 2014, attached no comparable conditions and is the closer analog to DLI. Across three rounds it raised on the order of hundreds of billions of dollars, roughly $100 billion in the first phase, $41 billion in the second, a further $47 billion in 2024. Wuhan Hongxin promised a leap straight to 14 and 7 nanometre nodes on a $19 billion budget and collapsed in 2021 without shipping a commercial chip. Dehuai, HiDM, Tacoma and Quanxin stalled at land preparation or pitch deck stage. Jiangsu Advanced Memory went bankrupt in 2023. A GlobalFoundries joint venture fab in Chengdu was abandoned as an empty shell in 2018 and sat untouched for five years. In 2022 the fund’s own chief executive and several fund managers were arrested on corruption charges. The money was not entirely wasted, China’s ecosystem is genuinely larger than a decade ago, but its integrated circuit trade deficit nearly doubled between 2010 and 2020 while the subsidies flowed, and its leading fabs remain years behind at advanced nodes. Hundreds of billions bought a great deal of visible activity. Nothing forced it to become the thing it was named for.</p>

<p>Same country, same state capacity, same willingness to spend. The era that made competence the price of entry got competence. The era that paid for output on trust did not.</p>

<h2 id="but-shouldnt-we-fix-land-and-labour-first">But shouldn’t we fix land and labour first?</h2>

<p>The standard objection is that all of this is downstream of something more basic, and India should fix land, labour, capital and contract enforcement before engineering conditionality into subsidy schemes. That is half right, and the wrong half matters.</p>

<p>It fails as a sequencing claim. None of the countries above had working factor markets when they started. Taiwan built ITRI under martial law. Korea’s financial system in the 1970s was deliberately repressed, credit allocated by the state rather than by price, and liberalisation largely followed industrialisation rather than preceding it. China in 1980 had no land market, no labour market in any recognisable sense and no commercial legal system worth the name. If working factor markets were a precondition, none of these industries would exist.</p>

<p>What those countries built instead were enclaves where scarce factors could be directed at chosen sectors, and then used the enclave as a laboratory for reforms that spread outward. Shenzhen was not primarily a tax haven. It was a testing ground for land leasing, foreign ownership, labour contracting and price liberalisation, run in a controlled space so failures stayed local. Hsinchu did the same on a smaller scale, co-locating ITRI, its spinoffs and their suppliers so people, knowledge and capital circulated within a few square kilometres.</p>

<p>India tried this and got the design wrong instructively. The SEZ Act of 2005 created enclaves defined almost entirely by tax treatment rather than regulatory experimentation, then destabilised even that: minimum alternate tax exemptions withdrawn from 2011-12, the dividend distribution tax exemption for developers terminated, a sunset clause limiting income tax benefits to units operational by March 2020. Of 564 formally approved SEZs, only 192 were operational by 2014. All Indian SEZs together cover roughly 61,600 hectares. Shenzhen alone covers about 49,300. China’s SEZs now account for something like 22 percent of GDP and 60 percent of exports. India built many small tax enclaves and revoked the tax benefit. China built a few large reform laboratories and let the reforms spread.</p>

<p>Where the objection holds is on retention rather than creation.</p>

<p>Capability can be created in an enclave. It cannot compound in one.</p>

<p>A company that succeeds inside a protected zone still has to grow outside it, into an economy where land takes years to assemble, labour law makes scale manufacturing risky, power is unreliable, and a commercial dispute takes a decade to resolve. Those conditions decide whether a firm that succeeds puts its next plant and its next thousand engineers here or elsewhere. A subsidy offsets a bad factor market for exactly as long as the subsidy lasts, which is the same abstraction wearing different clothes.</p>

<p>It also matters which factor markets bind for which problem, since the phrase gets used as though it were one thing. For assembly and fabrication, land, labour, power and logistics are decisive. For fabless chip design they barely register. C2i needed almost no land and employed a few dozen engineers. What it lacked was deep late-stage risk capital and a domestic customer, which are capital and product market failures. Reforming labour codes would not have kept C2i Indian. A domestic pool of capital willing to fund a decade of silicon might have.</p>

<p>Subsidies and factor markets solve different problems. India has been using the first as a substitute for the second.</p>

<h2 id="turning-the-hope-into-a-condition">Turning the hope into a condition</h2>

<p>None of this argues for scrapping these schemes. You have to fund productivity first, because there is no route to competence that skips getting companies founded and factories built, and India in 2020 had little of either. The mobile PLI created an assembly base that did not exist, and that base is the raw material any component ecosystem would stand on.</p>

<p>What forces the collapse is not mysterious, because other countries already wrote the clauses. Make value addition and technology localisation binding conditions of payment rather than monitored outcomes. Take equity in strategic programmes so the public balance sheet participates when an acquisition happens, the way ITRI’s stakes in UMC and TSMC gave Taiwan a return on the knowledge it had built. Attach perpetual domestic licensing rights to IP developed with public money, so the technology stays available to Indian firms regardless of who owns the company. Fund shared design infrastructure, IP libraries, EDA access, multi-project wafer runs, so knowledge pools in an institution the way it pooled inside MITI’s joint laboratory, rather than dispersing when a single startup is sold. Use public procurement to create the domestic customer that does not yet exist, since defence, railways, power utilities and telecom all buy silicon and none of them buy Indian. Treat enclaves as reform laboratories rather than tax arbitrage, which is the lesson from China that India copied in form and missed in substance.</p>

<p>Every one of these does the same thing an export target did for Korea or a shared lab did for Japan. It removes the option of staying productive without becoming competent, by making the money itself depend on the harder thing happening.</p>

<h2 id="the-real-risk">The real risk</h2>

<p>Activity is politically attractive because it is immediate and countable. Capability runs on a slower clock, and its absence surfaces only years later: when value addition plateaus a decade short of target, when a promising company cannot find a domestic customer, when no Indian investor can fund the next stage, when the only credible buyer is foreign. By then the credit for the original announcement has been banked and nobody is accountable for the gap.</p>

<p>That is the danger in treating these schemes as ends rather than as scaffolding. Lowering risk while deeper capital markets, customer bases and industrial networks develop behind them is a legitimate function. But scaffolding is supposed to come down, and if nothing is being built to stand on its own once it does, the country discovers what it doesn’t know at the worst possible moment, in public, the way C2i’s shareholders discovered that no amount of activity had produced a domestic buyer willing or able to compete with Infineon.</p>

<p>C2i did not fail. It built real intellectual property, attracted serious capital, and became valuable enough for a global leader to buy, which for a hard-tech startup is success by any normal definition. Nor did Foxconn do anything wrong by assembling phones in India and collecting an incentive it was contractually entitled to. In both cases the private actors did exactly what the policy paid them to do.</p>

<p>The question is what India asked for in exchange. It did ask for things. PLI required incremental sales over a base year. DLI requires design milestones and turnover thresholds. These are real conditions, and firms that missed them went unpaid, which is why four fifths of domestic mobile applicants collected nothing. But look at what those conditions test. Every one of them is satisfiable by productivity alone. You can hit an incremental sales target by assembling imported components. You can hit a turnover threshold on a chip that is designed here and then leaves. Nothing in either metric requires you to know something at the end that you did not know at the start.</p>

<p>Now look at what the others demanded. Taiwan asked RCA for the process, not the plant. Japan asked five rivals to sit in one laboratory. Korea asked its champions to win orders from foreign buyers who owed them nothing, and pulled their credit when they lost. China asked for joint ventures and local content. Each of those is a condition you cannot satisfy without acquiring something you didn’t have before. They are tests of competence. India’s are tests of productivity, and productivity is the one thing you can buy without learning anything.</p>

<p>Activity is the instrument. Capability is the hope. And a country, like an engineer leaning on a tool it has never looked underneath, only finds out how much it was hoping for at the exact moment the tool stops being enough.</p>]]></content><author><name>Yash Tambawala</name></author><summary type="html"><![CDATA[What the C2i sale reveals about the design flaw in PLI and DLI: activity is the instrument, capability is the hope, and no clause makes it a condition.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://yashtambawala.com/assets/og/indias-industrial-policy-buys-activity-not-capability.png" /><media:content medium="image" url="https://yashtambawala.com/assets/og/indias-industrial-policy-buys-activity-not-capability.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">What Actually Improved</title><link href="https://yashtambawala.com/2026/09/01/what-actually-improved.html" rel="alternate" type="text/html" title="What Actually Improved" /><published>2026-09-01T07:30:00+05:30</published><updated>2026-09-01T07:30:00+05:30</updated><id>https://yashtambawala.com/2026/09/01/what-actually-improved</id><content type="html" xml:base="https://yashtambawala.com/2026/09/01/what-actually-improved.html"><![CDATA[<p>Two claims about AI get made constantly, and they fail for the same reason. One says the last nine years were “just scaling.” The other says nothing fundamental happened at all. Both treat capability as a single number that either went up a lot or didn’t.</p>

<p>Capability isn’t a single number. Progress happened in <strong>nine distinct areas</strong>, each with its own metrics, its own bottleneck, its own literature and largely its own people. They advanced in parallel, on different schedules, for different reasons. At any given moment one of them was the binding constraint: the thing everything else was waiting on. Much of the field’s history is the story of that constraint moving from one area to the next.</p>

<p>This article does three things. It names the nine areas and the metrics that define each. It places the work that moved those metrics on a shared timeline, so you can read across a year and see what was happening everywhere at once. And for every area it asks a question that gets asked far too rarely: <strong>how much room is actually left?</strong></p>

<p>That last question is the useful one. Some areas are running into limits that can be proven: thermodynamics, information theory, computational complexity. Some are running into limits that are merely economic and will fall. Some have no known limit at all, and that absence is itself a finding, not a gap. Where the remaining distance is large, progress keeps going. Where it’s small, progress stalls no matter how much money arrives. Where nobody has drawn the wall, we genuinely don’t know — and those happen to be the areas everything now depends on.</p>

<p>This is the third of three connected pieces. <a href="/2026/08/30/ai-system-design-for-product-managers.html">AI System Design</a> covers how workload shape decides architecture. <a href="/2026/08/31/how-an-ai-model-turns-input-into-output.html">How an AI Model Turns Input Into Output</a> covers the mechanics inside a single request. This one covers where all of it came from and how much further it goes.</p>

<h2 id="1-the-nine-areas">1. The nine areas</h2>

<p>The areas below are ordered the way a capability actually travels: from the silicon it runs on, through the training that creates it, to the scoreboard that tells us whether anything improved. They group into four planes.</p>

<table>
  <thead>
    <tr>
      <th>Plane</th>
      <th>Areas</th>
      <th>The question it answers</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>The machine</td>
      <td>Substrate · Architecture</td>
      <td>What is it made of?</td>
    </tr>
    <tr>
      <td>The learning</td>
      <td>Pretraining · Post-training · Run-time reasoning</td>
      <td>How does it become capable?</td>
    </tr>
    <tr>
      <td>The delivery</td>
      <td>Serving · Agents · Distribution</td>
      <td>How does it reach a person?</td>
    </tr>
    <tr>
      <td>The instrument</td>
      <td>Evaluation</td>
      <td>How would we know any of this?</td>
    </tr>
  </tbody>
</table>

<p><strong>Evaluation sits across all of them deliberately.</strong> It tells the other eight whether they moved, and its failure modes — saturation, contamination, label noise — corrupt every claim made anywhere else. Treating it as a peer rather than a caveat is part of being honest about the rest.</p>

<p>Here is what each area actually is, in plain terms.</p>

<ol>
  <li><strong>Substrate</strong> — the physical machine: chips, memory, the wires between them, and the power they draw. Everything else runs on it.</li>
  <li><strong>Architecture</strong> — the design of the model itself: how it is wired, how attention works, how much of it runs for any given token.</li>
  <li><strong>Pretraining</strong> — the long, expensive run that sets the weights by predicting text. Where raw capability comes from.</li>
  <li><strong>Post-training</strong> — turning a raw predictor into something that follows instructions and refuses sensibly. Small compute, enormous visible effect.</li>
  <li><strong>Run-time reasoning</strong> — extra computation spent per question <em>after</em> training is over. The model works a problem through before answering, and you pay for that thinking.</li>
  <li><strong>Serving</strong> — running the finished model for many users at once: how fast the first word arrives, how fast the rest follow, what a token costs.</li>
  <li><strong>Agents</strong> — the system wrapped around the model: tools, memory, loops. Where a single answer becomes a multi-hour task.</li>
  <li><strong>Distribution</strong> — who can obtain and run a given capability: open weights, price, and whether it fits on a laptop or a phone.</li>
  <li><strong>Evaluation</strong> — how anyone knows whether the other eight moved. The instrument, and it has its own failure modes.</li>
</ol>

<h2 id="2-five-ways-one-area-moves-another">2. Five ways one area moves another</h2>

<p>The areas are separate but not independent. Five patterns recur, and naming them makes the timeline readable as a mechanism rather than a list of events.</p>

<p><strong>Unlock.</strong> A gain in one area removes a wall in another. FlashAttention changed how much memory traffic attention required, not the mathematics, just the traffic. That made long context economically serveable, which in turn made multi-hour agents plausible.</p>

<p><strong>Handoff.</strong> One area stiffens, so effort migrates. As the returns from scaling pretraining data got harder to buy, run-time reasoning opened as an entirely new axis.</p>

<p><strong>Substitution.</strong> A gain in one area buys what another can’t supply. Quantization and KV-cache compression substitute for memory capacity you can’t purchase.</p>

<p><strong>Loop.</strong> A downstream gain refinances an upstream one. Post-training produced a usable product. The product produced revenue. The revenue bought substrate, which funded larger pretraining runs.</p>

<p><strong>Constraint-shift.</strong> An external cap on one area redirects effort into others. Export controls on chips produced a disproportionate share of the world’s architecture and serving-efficiency research. More on that below.</p>

<h2 id="3-nine-lanes-on-one-timeline">3. Nine lanes on one timeline</h2>

<p>Read a row to follow one area. Read a column to see what was happening in all nine at once. The parallelism is the point, and it’s what makes the progress feel sudden from outside. The highlighted staircase marks the binding constraint as it moves. The right-hand margin shows how much room each area has left.</p>

<div class="lanes">
  <div class="lanes-bar">
    <button class="lanes-btn" id="lb1" aria-pressed="true">Binding constraint</button>
    <span class="lanes-hint">Click a headroom gauge to open its bound &rarr;</span>
  </div>
  <div class="lanes-shell"><div class="lanes-scroll"><div class="lanes-grid band" id="lg1"></div></div></div>
  <div class="lanes-key">
    <span class="lanes-k"><i class="lanes-sw sw1"></i><span><b>Milestone</b>Work that moved a metric in that lane</span></span>
    <span class="lanes-k"><i class="lanes-sw sw2"></i><span><b>Imposed cap</b>A limit placed on a lane from outside it</span></span>
    <span class="lanes-k"><i class="lanes-sw sw3"></i><span><b>Binding constraint</b>The lane everything else was waiting on</span></span>
    <span class="lanes-k"><i class="lanes-sw sw4"></i><span><b>At the bound</b>Headroom effectively spent</span></span>
    <span class="lanes-k"><i class="lanes-sw sw5"></i><span><b>No bound drawn</b>Nobody has proved a ceiling here</span></span>
  </div>
</div>

<h2 id="4-the-machine">4. The machine</h2>

<h3 id="substrate">Substrate</h3>

<p><strong>Metrics:</strong> FLOP/s per chip at a given precision · HBM capacity and bandwidth · joules per FLOP · dollars per FLOP · FLOPs per watt.</p>

<p>The line runs V100 to A100 to H100 to H200 to B200, and the important changes go beyond raw arithmetic. The H100 introduced practical 8-bit training, which halved the bytes moved for every number in a training run. Memory bandwidth rose from 2.0 TB/s on the A100 to 3.35 on the H100 to 7.7 aggregate on the B200. As section 6 shows, bandwidth rather than arithmetic is what actually limits generation.</p>

<p>This is also the only area in the chart with a substantial number of <strong>imposed caps</strong>: the October 2022 export controls, their widening in October 2023, restrictions reaching memory itself in 2024, a partial relaxation in December 2025, and the Chip Security Act in March 2026. A lane can be limited by policy long before it’s limited by physics.</p>

<p><strong>Distance to the bound: roughly seven orders of magnitude.</strong> Landauer’s principle sets the energy cost of erasing one bit at <code class="language-plaintext highlighter-rouge">kT ln2</code>, about 2.9 × 10<sup>-21</sup> joules at room temperature. Charged per 16-bit operation that’s around 5 × 10<sup>-20</sup> J. Current accelerators spend about 5 × 10<sup>-13</sup> J per FLOP.</p>

<p>That gap runs to roughly ten million, and reversible computing sits below even that floor. The conclusion cuts against most energy commentary: <strong>physics is not the constraint on compute efficiency, and won’t be for decades.</strong> Power delivery, fab capacity and capital are the real constraints. It’s exactly why the caps that bind this lane are drawn by governments rather than by nature.</p>

<h3 id="architecture">Architecture</h3>

<p><strong>Metrics:</strong> loss at fixed compute · active parameters versus total · KV bytes per token · supported context length.</p>

<p>The 2017 transformer’s first contribution wasn’t accuracy but <em>training parallelism</em>: recurrence forced sequential processing, and attention didn’t. Everything downstream follows from that. The lane then splits into two long campaigns: making position information work at length (RoPE, and the interpolation methods that stretched it), and making attention cheaper per token. Grouped-query attention cut the number of key/value heads. Sliding windows bounded what each position could see. Multi-head latent attention compressed the KV cache into a smaller latent representation. Sparse mixture-of-experts made the number of parameters that run per token far smaller than the number that exist.</p>

<p>That last idea reshaped the economics. A model can hold 671 billion parameters and run 37 billion of them for any given token.</p>

<p><strong>Distance to the bound: at it.</strong> Alman and Song proved a sharp transition in 2023: with head dimension of order log <em>n</em>, truly subquadratic attention is possible <strong>if and only if</strong> the matrix entries stay bounded below a threshold of order √log <em>n</em>. Above that threshold it’s impossible under the Strong Exponential Time Hypothesis.</p>

<p>This is the theoretical reason no linear-attention scheme is a free lunch. Every subquadratic method buys its speed by changing the problem: approximating, bounding entries, or imposing sparsity. Those are good trades worth making. They aren’t escapes.</p>

<h2 id="5-the-learning">5. The learning</h2>

<h3 id="pretraining">Pretraining</h3>

<p><strong>Metrics:</strong> tokens · parameters · training FLOPs · the compute-optimal ratio between them · model FLOPs utilisation · dollars per run.</p>

<p>Two results define this lane, and the second corrects the first. The 2020 scaling laws established that loss falls predictably with compute. That sounds academic, but it was arguably the most consequential finding in the field’s financial history: it converted “will a bigger model be better” from a research gamble into an engineering forecast, and a forecast is something you can raise capital against.</p>

<p>Chinchilla then showed in 2022 that everyone had been undertraining. For a given compute budget, far more data and fewer parameters was the better trade. That single correction explains why a 7-billion-parameter model in 2024 outperformed a 175-billion-parameter model from 2020, and it redirected the whole industry from parameter counts toward data.</p>

<p><strong>Distance to the bound: approaching, on the resource rather than the theory.</strong> Two different limits apply here. The scaling law’s irreducible term, around 1.69 nats per token in published fits, is the entropy of the data itself: a floor no compute budget crosses. Separately, Epoch AI estimates the effective stock of public human text at roughly 300 trillion tokens, with models on track to consume it somewhere between 2026 and 2032.</p>

<p>The first bound is still distant. The second is arriving now; we’re inside the front edge of that window this year. But a stock isn’t a law, and stocks get substituted. That’s precisely what synthetic data and the pivot to run-time reasoning are doing.</p>

<h3 id="post-training">Post-training</h3>

<p><strong>Metrics:</strong> preference win-rate · instruction adherence · refusal correctness in <em>both</em> directions · jailbreak robustness · calibration.</p>

<p>This lane produced the most visible jump in the field’s history while using comparatively little compute. Reinforcement learning from human feedback turned a text predictor into something that answers the question asked. Constitutional AI replaced much of the human labelling with written principles. Direct preference optimisation later showed the same result was reachable without training a separate reward model at all.</p>

<p>The fact worth sitting with: <strong>GPT-3 existed in 2020.</strong> What arrived in late 2022 was largely a post-training and interface result, not a capability one — the clearest available proof that the visible jump and the underlying curve are different things.</p>

<p><strong>Distance to the bound: nobody has drawn one.</strong> There is no impossibility theory for alignment. Adjacent results gesture at the difficulty (Arrow’s theorem on aggregating preferences, the formal characterisation of Goodhart’s law under proxy optimisation) but none of them bounds how well a model can be made to do what was meant. The honest answer is that we can’t say how much of this lane is left.</p>

<h3 id="run-time-reasoning">Run-time reasoning</h3>

<p><strong>Metrics:</strong> accuracy as a function of thinking-token budget · pass@1 versus pass@k · <strong>cost and latency per solved task</strong>, not per token.</p>

<p>Chain-of-thought prompting in 2022 showed that asking a model to work through a problem improved the answer. Process reward models in 2023 showed you could grade the reasoning steps rather than only the final answer. In 2024 this stopped being a prompting trick and became trained behaviour, and in January 2025 an open recipe using group-relative policy optimisation made the method reproducible outside the largest labs.</p>

<p>This is where <a href="/2026/08/31/how-an-ai-model-turns-input-into-output.html">the reasoning tokens</a> you pay for but never see come from. The metric shift here is the real story: the meaningful unit stopped being cost per token and became cost per solved task.</p>

<p><strong>Distance to the bound: bounded by the verifier, and unquantified.</strong> Reinforcement learning on reasoning works where an answer can be checked, and achievable accuracy can’t exceed what the checker can distinguish. Where verification is cheap and exact (mathematics, code, formal proof) the bound is high and the method works spectacularly. Outside those domains, which is most work of economic value, there’s no bound because there’s barely a method.</p>

<p>This lane didn’t meaningfully exist before 2024. Its arrival is the clearest case in the whole chart of a <strong>handoff</strong>: a new axis opening precisely as an old one stiffened.</p>

<h2 id="6-the-delivery">6. The delivery</h2>

<h3 id="serving">Serving</h3>

<p><strong>Metrics:</strong> dollars per million tokens · time to first token · time between tokens · concurrent sessions per card · cache hit rate.</p>

<p>Four years of work here, and nearly all of it attacks the same quantity. Continuous batching let new requests join a running batch. FlashAttention cut memory traffic without changing the mathematics. PagedAttention managed the KV cache the way an operating system manages virtual memory, ending the fragmentation that wasted most of a card. Speculative decoding drafted several tokens cheaply and verified them in one pass. Quantization to four bits put capable models on laptops. Prompt caching reused an identical prefix rather than recomputing it.</p>

<p>Why they all converge is visible in one calculation.</p>

<details class="technical-detail">
  <summary>For the extra curious: the roofline that explains the whole lane</summary>

  <p>Generating one token requires reading the weights that token uses. So for a single stream:</p>

  <div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>tokens per second  ≤  memory bandwidth ÷ bytes read per token
</code></pre></div>  </div>

  <p>For a dense 70B model at 8-bit precision on one H100:</p>

  <div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>3.35 TB/s ÷ 70 GB ≈ 48 tokens/second
</code></pre></div>  </div>

  <p>That’s a ceiling, not an estimate — real systems land below it. Every technique above is an attack on the same denominator. Batching amortises one read across many users. Mixture-of-experts reads only the active parameters. Quantization shrinks the bytes. Speculative decoding gets several tokens per read.</p>

  <p>Decoding is bounded by <strong>reads, not by arithmetic</strong>. This is also why the <a href="/2026/08/31/how-an-ai-model-turns-input-into-output.html">KV cache</a> occupies such expensive memory: it exists so that earlier positions don’t have to be read and recomputed at every step.</p>
</details>

<p><strong>Distance to the bound: pressed against two of them.</strong> The bandwidth roofline above, and a second floor that no engineering touches at all: a round trip between continents can’t beat light in fibre, which comes to roughly 200 ms between Mumbai and Virginia. That latency floor is the entire argument for running models on the device.</p>

<h3 id="agents">Agents</h3>

<p><strong>Metrics:</strong> end-to-end task success · <strong>task horizon</strong> · per-step reliability · tool-call accuracy · cost per completed task · human interventions per task.</p>

<p>Retrieval, then interleaved reasoning and action, then tools with typed arguments, then a shared protocol for those tools, then the screen itself as an interface. The through-line is that a single answer became a long-running task — exactly the shift <a href="/2026/08/30/ai-system-design-for-product-managers.html">AI System Design</a> is about.</p>

<p>The metric that matters most here is task horizon: the length of task a model completes reliably. METR’s January 2026 update puts the 50% time horizon at <strong>320 minutes</strong>, and finds the doubling time has been accelerating: 89 days for models released since 2024, against roughly seven months measured across the longer run.</p>

<p><strong>Distance to the bound: no ceiling has been drawn, but the arithmetic bites now.</strong> Per-step reliability <em>r</em> across <em>n</em> steps compounds as <em>r<sup>n</sup></em>. Ninety percent success over a hundred steps demands 99.9% reliability per step. That isn’t a theorem, just multiplication, which is why it’s the practical constraint today rather than eventually.</p>

<p>The open question is whether verification and error correction can break the compounding the way fault-tolerance broke it for unreliable hardware: whether agency has a <strong>threshold theorem</strong>. Nobody has one. Everything being built right now is a bet on the answer.</p>

<h3 id="distribution">Distribution</h3>

<p><strong>Metrics:</strong> <strong>open-weight lag in months behind the closed frontier</strong> · capability density · price at fixed capability · what class of device can run it.</p>

<p>BERT, then GPT-2’s staged release, then open replications of GPT-3, then a weights release in early 2023 that seeded most of what followed. Then permissive licensing, then frontier-scale open weights, then open reasoning models. By 2026 the lag is measured in months rather than years, and the models are cheap enough that an open model on commodity infrastructure is the default for most production work rather than the compromise.</p>

<p><strong>Distance to the bound: real headroom.</strong> Allen-Zhu and Li found that language models store <strong>2 bits of knowledge per parameter</strong>, and only 2, a figure that holds even under 8-bit quantization. That places a genuine floor under capability density: a 7B model tops out near 14 billion bits, which by their estimate exceeds English Wikipedia and textbooks combined. We’re not near it.</p>

<p>The unanswered question in this lane isn’t whether the open gap closes but <strong>what it converges to</strong>, and no theory predicts that number.</p>

<h2 id="7-the-instrument">7. The instrument</h2>

<h3 id="evaluation">Evaluation</h3>

<p><strong>Metrics:</strong> benchmark saturation half-life · contamination rate · construct validity · label-noise floor.</p>

<p>Every benchmark follows the same lifecycle: introduced, chased, saturated, contaminated, replaced. GLUE gave way to SuperGLUE within a year. MMLU held for longer. Human preference voting sidestepped static test sets. Then came benchmarks built from real work: actual repository bugs, questions hard enough to stump specialists, mathematics that stops research mathematicians.</p>

<p><strong>Distance to the bound: past it, for the most-cited benchmark in the field.</strong> A benchmark’s ceiling isn’t 100%. It’s 100% minus the share of items that are mislabelled or ambiguous. MMLU carries documented label errors on the order of 6–9% of items, capping usable headroom near 95%. Frontier systems now report above 92%.</p>

<p>The remaining gap is smaller than the measurement error, so movement there is noise. SWE-bench Verified exists precisely because of this problem: ninety engineers hand-screened the task pool to remove broken tests, underspecified issues and unreliable environments.</p>

<p>This is why evaluation belongs beside the other eight areas rather than in a footnote. <strong>It’s the instrument every other claim in this article depends on</strong>, and when the instrument saturates, confident statements about progress quietly stop meaning anything.</p>

<h2 id="8-what-the-export-controls-actually-produced">8. What the export controls actually produced</h2>

<p>The China thread isn’t a separate history. It’s the clearest available case of <strong>constraint-shift</strong>, and the nine-lane frame explains it without needing geopolitics.</p>

<p>Blocked at the substrate, the one area where a cap could be imposed from outside, Chinese labs directed effort into the areas where a cap couldn’t reach. The result is a body of work concentrated almost entirely in architecture, serving and distribution: multi-head latent attention compressing the KV cache, fine-grained expert routing so that a small fraction of a very large model runs per token, 8-bit training end to end, careful overlap of communication and computation, and a cheap reinforcement-learning recipe for reasoning. Then the weights were released openly, which made that efficiency work everybody’s baseline rather than one company’s advantage.</p>

<p>The honest accounting matters. Headline training-cost figures typically describe a final run and exclude the research, failed runs and infrastructure around it. A good deal of this work is excellent execution of ideas already published rather than invention from nothing. And the ecosystem still trails at frontier-scale pretraining and in evaluation depth.</p>

<p>But the structural lesson generalises well beyond one country, and it’s the most useful thing in this section: <strong>abundance optimises capability, scarcity optimises efficiency, and efficiency work compounds for everyone.</strong> A lab with unlimited chips has little reason to halve its KV cache. A lab without them has no other option. Once the technique is published, everyone’s serving costs fall.</p>

<h2 id="9-now-do-the-same-for-things-that-touch-the-world">9. Now do the same for things that touch the world</h2>

<p>Everything above is necessary for a robot or a self-driving car and nowhere near sufficient. Five further areas gate anything embodied, and they behave completely differently, because <strong>their limits are physical and close, rather than theoretical and distant.</strong></p>

<p>That contrast is the strongest evidence for this whole way of looking at the field.</p>

<ol>
  <li><strong>Sensing</strong> — measuring the world: lidar, cameras, radar, touch. Everything downstream is limited by what was perceived.</li>
  <li><strong>Onboard energy</strong> — carrying the power. Every runtime, range and payload figure in robotics is downstream of one number.</li>
  <li><strong>Actuation</strong> — moving against the world: motors, gearboxes, tendons. Turning decisions into force.</li>
  <li><strong>Embodied policy</strong> — learning to act. The model that turns what was sensed into what the body does next.</li>
  <li><strong>Safety validation</strong> — proving it is safe enough to deploy.</li>
</ol>

<div class="lanes">
  <div class="lanes-bar">
    <button class="lanes-btn" id="lb2" aria-pressed="true">Binding constraint</button>
    <span class="lanes-hint">Note how many gauges are hatched &rarr;</span>
  </div>
  <div class="lanes-shell"><div class="lanes-scroll"><div class="lanes-grid emb band" id="lg2"></div></div></div>
</div>

<p>Set the two charts side by side and the divergence explains itself.</p>

<p>Compute sits roughly <strong>ten million times</strong> above its thermodynamic floor. Battery energy density sits within about <strong>1.5×</strong> of what lithium-ion chemistry permits: 250 to 300 watt-hours per kilogram today against a theoretical ceiling near 400 to 460, while petrol carries around 12,500. Electric actuators are close to what the magnetic saturation of iron allows and have improved by only a small multiple in fifty years.</p>

<p>Language models could ride lanes with enormous headroom. Robots couldn’t. <strong>That difference in available headroom, more than any difference in algorithmic difficulty, is why one sprinted and the other crawled.</strong></p>

<p>The exception proves the rule. One embodied lane genuinely did collapse: lidar fell from $75,000 a unit to under $200, a reduction of more than 99%. It moved because its binding limit was <strong>economic, not physical</strong>, and manufacturing scale defeats a price. It doesn’t defeat eye-safety limits on laser power, or the number of watt-hours in a kilogram. The same pattern holds in actuation, where cost and integration improved enormously while power density barely moved at all.</p>

<p>Two further bounds in this group deserve naming, because both get mistaken for engineering problems.</p>

<p><strong>Embodied policy faces the data-stock problem in its worst form.</strong> Pretraining had an internet to read. Robotics doesn’t, and can’t: there’s no web-scale corpus of robot trajectories waiting to be found, because every hour of manipulation data has to be manufactured at the speed of physical reality, one robot-hour at a time. The escapes under way — simulation, pooling data across robot types, learning priors from human video — are all attempts to substitute a stock that can’t simply be collected.</p>

<p><strong>Safety validation faces a statistical wall that better models don’t move.</strong> RAND’s result is that demonstrating with 95% confidence that a driverless fleet beats the human fatality rate would require roughly <strong>275 million miles driven without a fatality</strong>, and under some assumptions, hundreds of billions. This is a property of how rare fatal crashes are, not of how good the driving is. No improvement in perception or planning shortens the distance needed to <em>prove</em> the improvement. It’s why the field leans on simulation, staged rollout and disengagement proxies, and why deployment has advanced one metropolitan area at a time.</p>

<h2 id="10-what-this-predicts">10. What this predicts</h2>

<p>The distance column turns the usual futurology into something closer to measurement. Three groups, three different expectations.</p>

<p><strong>Where the gap is large, progress continues.</strong> Substrate has seven orders of magnitude of thermodynamic headroom; what binds it is capital, power and policy, all of which can change. Distribution has real room under the capability-density limit. Neither of these areas is going to stall for reasons of physics.</p>

<p><strong>Where the gap is closed, progress stalls regardless of investment.</strong> Exact attention is at its complexity-theoretic threshold. Serving is pressed against the bandwidth roofline and the speed of light. The most-cited benchmark in the field has passed its own label-noise floor. In the embodied group, energy density and actuation are at the wall, and the safety-validation arithmetic doesn’t care how good the model gets. Money doesn’t buy past any of these; only a change of problem does.</p>

<p><strong>Where no bound has been drawn, we genuinely don’t know.</strong> This is the uncomfortable part, because it’s precisely the set of areas everything now depends on. Post-training has no impossibility theory. Run-time reasoning is bounded by verifier quality in checkable domains and by nothing legible outside them. Agents have compounding arithmetic and no threshold theorem. These three carry most of the current expectations about the next few years, and they’re the three where no one can say what the ceiling is.</p>

<p>If you want one question to track, it’s the threshold question for agents: can verification break the compounding of per-step error the way error correction broke it for unreliable hardware? Every other open question in the delivery plane follows from that one.</p>

<h2 id="what-to-keep-in-your-head">What to keep in your head</h2>

<ol>
  <li>Progress was never one curve. It was nine areas moving in parallel on different schedules.</li>
  <li>At any moment one area was the binding constraint, and the history is that constraint migrating: architecture, then the training recipe, then alignment, then serving cost, then run-time reasoning, then reliability.</li>
  <li>The visible jump in late 2022 was mostly post-training and interface. The underlying capability had existed for two years.</li>
  <li>Scaling was predictable, which is what made it financeable. Chinchilla’s correction, not raw size, is why small models got good.</li>
  <li>Constraints determine which area a lab attacks. Abundance optimises capability; scarcity optimises efficiency; the efficiency work then compounds for everyone.</li>
  <li>Three areas are at a hard bound, two have vast headroom, and three have no bound anyone has drawn.</li>
  <li>Embodied AI is slow not because it’s intellectually harder but because its lanes sit close to physical limits that software lanes don’t.</li>
  <li>When a benchmark passes its label-noise floor, claims of progress on it stop carrying information.</li>
</ol>

<p>The last nine years felt abrupt because several curves compounded while only one of them was visible from outside. The reason to watch metrics rather than announcements: the next jump will show up as a number in one of these lanes well before it shows up in a product.</p>

<hr />

<p><strong>A note on the figures.</strong> The Landauer comparison, the irreducible-loss term, the 2-bits-per-parameter result, the attention hardness threshold, the ~300 trillion token text stock, the MMLU label-noise ceiling, RAND’s 275 million miles, lithium-ion energy densities and the lidar price history are drawn from published work rather than recalled. The bandwidth roofline and the step-compounding figure are derived here from stated specifications and hold only for the configuration named. Assigning a milestone to a single lane is an editorial judgement; several belong in two at once, which is rather the point of drawing them in parallel.</p>]]></content><author><name>Yash Tambawala</name></author><summary type="html"><![CDATA[Nine areas of AI progress, the metrics that define each, and how much room is left in every one before physics or information theory says stop]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://yashtambawala.com/assets/og/what-actually-improved.png" /><media:content medium="image" url="https://yashtambawala.com/assets/og/what-actually-improved.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">How an AI Model Turns Input Into Output</title><link href="https://yashtambawala.com/2026/08/31/how-an-ai-model-turns-input-into-output.html" rel="alternate" type="text/html" title="How an AI Model Turns Input Into Output" /><published>2026-08-31T11:15:00+05:30</published><updated>2026-08-31T11:15:00+05:30</updated><id>https://yashtambawala.com/2026/08/31/how-an-ai-model-turns-input-into-output</id><content type="html" xml:base="https://yashtambawala.com/2026/08/31/how-an-ai-model-turns-input-into-output.html"><![CDATA[<p>AI models are often explained from the middle. Someone introduces attention, embeddings or GPUs before explaining what calculation the model is trying to perform. Each definition may be correct, but the definitions do not join into a picture.</p>

<p>This article follows one incomplete sentence through a language model:</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>The bank approved the
</code></pre></div></div>

<p>We will introduce a term only when the model encounters the problem that term solves. The main explanation does not require mathematics. The exact calculations are available in expandable sections for readers who want them.</p>

<p>This is the technical companion to <a href="/2026/08/30/ai-system-design-for-product-managers.html">AI System Design</a>, which applies these mechanics to product architecture, context design, latency and cost.</p>

<h2 id="1-the-model-begins-with-one-job">1. The model begins with one job</h2>

<p>Given the incomplete sentence, the model calculates a score for every possible next token in its vocabulary. Those scores are converted into probabilities.</p>

<p>An illustrative result might be:</p>

<table>
  <thead>
    <tr>
      <th>Possible next token</th>
      <th style="text-align: right">Probability</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">loan</code></td>
      <td style="text-align: right">31%</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">request</code></td>
      <td style="text-align: right">14%</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">payment</code></td>
      <td style="text-align: right">8%</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">application</code></td>
      <td style="text-align: right">6%</td>
    </tr>
  </tbody>
</table>

<p>The model produces the distribution. It does not choose from it. Something outside the model applies a selection rule, which <a href="#7-a-selection-rule-turns-probabilities-into-one-token">section 7</a> describes.</p>

<p>If the selection rule picks <code class="language-plaintext highlighter-rouge">loan</code>, the sequence becomes:</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>The bank approved the loan
</code></pre></div></div>

<p>The model then calculates a new distribution for the token after <code class="language-plaintext highlighter-rouge">loan</code>. Producing an answer means repeating this operation.</p>

<p>This creates four requirements:</p>

<ol>
  <li>Text must become numbers because the model performs arithmetic.</li>
  <li>The numbers must preserve order because word order changes meaning.</li>
  <li>Each position must be able to use information from other positions.</li>
  <li>The final result must become scores for possible next tokens.</li>
</ol>

<p>Tokens, embeddings, position information and attention exist to satisfy these requirements.</p>

<h2 id="2-text-is-divided-into-tokens">2. Text is divided into tokens</h2>

<p>A model does not normally process one English word at a time. A <strong>tokenizer</strong> divides text into entries from a fixed vocabulary.</p>

<p>Our sentence might become:</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>["The", " bank", " approved", " the"]
</code></pre></div></div>

<p>Tokens are not always complete words. An uncommon word may become several tokens. Punctuation and spaces can also affect the division.</p>

<p>Every vocabulary entry has an integer identifier:</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>"The"       → 791
" bank"     → 3821
" approved" → 9342
" the"      → 279
</code></pre></div></div>

<p>The model now sees the sequence of IDs:</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>[791, 3821, 9342, 279]
</code></pre></div></div>

<p>These numbers are labels, not measurements. Token 9342 is not greater or more meaningful than token 791. The IDs merely identify vocabulary entries.</p>

<p>That makes them useful for lookup, but not yet useful for the model’s learned arithmetic.</p>

<h2 id="3-each-token-id-selects-a-learned-list-of-numbers">3. Each token ID selects a learned list of numbers</h2>

<p>The model contains a large table with one row for every vocabulary token. Our second token, ` bank`, had ID 3821, so it selects row 3821.</p>

<p>A shortened version of that row might look like:</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>[0.18, -0.42, 0.07, 0.91, ...]
</code></pre></div></div>

<p>This ordered list of numbers is a <strong>vector</strong>. The learned starting vector for a token is its <strong>embedding</strong>.</p>

<p>The numbers were learned during training. The model repeatedly predicted tokens, measured its errors and adjusted its internal values to reduce those errors. The embedding table was adjusted along with the rest of the model.</p>

<p>Every occurrence of ` bank` starts from this identical row — the financial one and the river one alike. Nothing here distinguishes them yet; that only happens once the surrounding tokens are allowed to act on it, which is section 5.</p>

<p>An embedding dimension does not usually have a clean label such as “financial meaning.” Information is distributed across many numbers and interpreted by later model operations.</p>

<h3 id="embedding-and-activation-are-not-the-same-term">Embedding and activation are not the same term</h3>

<p>The embedding is the token’s starting vector. As the input passes through the model, that vector is repeatedly updated. The vector at a particular point in the calculation is called an <strong>activation</strong>.</p>

<table>
  <thead>
    <tr>
      <th>Term</th>
      <th>Meaning</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Token</td>
      <td>A discrete unit of input or output</td>
    </tr>
    <tr>
      <td>Token ID</td>
      <td>The vocabulary identifier for that token</td>
    </tr>
    <tr>
      <td>Embedding</td>
      <td>The token’s learned starting vector</td>
    </tr>
    <tr>
      <td>Activation</td>
      <td>The token’s current vector at one stage of processing</td>
    </tr>
  </tbody>
</table>

<p>The term “activation” therefore does not refer to another stored meaning. It refers to the numerical values currently moving through the model for this particular input.</p>

<details class="technical-detail">
  <summary>For the extra curious: the embedding table's shape</summary>

  <p>If a vocabulary contains 100,000 tokens and the model represents each token using 4,096 numbers, the embedding table has:</p>

  <div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>100,000 rows × 4,096 columns
</code></pre></div>  </div>

  <p>If <code class="language-plaintext highlighter-rouge">E</code> is that table, the starting vector for token ID <code class="language-plaintext highlighter-rouge">i</code> is the row <code class="language-plaintext highlighter-rouge">E[i]</code>.</p>
</details>

<h2 id="4-the-model-must-preserve-order-and-build-context">4. The model must preserve order and build context</h2>

<p>These sentences contain many of the same tokens:</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>The dog chased the man.
The man chased the dog.
</code></pre></div></div>

<p>They do not mean the same thing. The model therefore incorporates information about each token’s position. Different model families implement position differently, but the requirement is constant: the calculation must preserve sequence order.</p>

<p>Order is not enough. Consider <code class="language-plaintext highlighter-rouge">bank</code> in these sentences:</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>The bank approved the loan.
They sat beside the river bank.
</code></pre></div></div>

<p>The token <code class="language-plaintext highlighter-rouge">bank</code> begins with the same embedding in both. Its later activation must become different because the surrounding tokens provide different information.</p>

<p>The model therefore needs an operation that lets each token position update itself using other permitted positions. That operation is <strong>attention</strong>.</p>

<h2 id="5-attention-decides-how-positions-contribute-to-one-another">5. Attention decides how positions contribute to one another</h2>

<p>For each position, attention performs three steps:</p>

<ol>
  <li>Calculate how strongly that position should use each available earlier position.</li>
  <li>Convert those strengths into weights.</li>
  <li>Combine information from the earlier positions using those weights.</li>
</ol>

<p>The model first creates three numerical projections for every position:</p>

<ul>
  <li>A <strong>query</strong> participates in deciding what other positions should contribute to the position being updated.</li>
  <li>A <strong>key</strong> participates in the matching calculation for each available position.</li>
  <li>A <strong>value</strong> carries the information that can be combined into the result.</li>
</ul>

<p>Queries are not literal questions. Keys are not database keys. Values are not human-readable facts. They are different learned transformations of the current activations.</p>

<p>Queries and keys calculate the weights. The corresponding values are combined using those weights. This separation lets the model use one set of numerical features to determine relevance and another set to carry information.</p>

<h3 id="the-same-operation-on-our-four-tokens">The same operation, on our four tokens</h3>

<p>Abstractly that is hard to hold. Run it on the sentence instead.</p>

<p>Our sequence is four tokens. Every position produces a query, a key and a value. To update position 4, <code class="language-plaintext highlighter-rouge">the</code>, the model takes position 4’s <strong>query</strong> and compares it against the <strong>key</strong> of every position it is allowed to see — positions 1 through 4. Four comparisons give four numbers; softmax turns them into weights that sum to 1.</p>

<p>An illustrative result for this sentence:</p>

<figure class="post-figure" aria-labelledby="attn-fig-title">
<svg viewBox="0 0 580 330" role="img" aria-labelledby="attn-svg-title attn-svg-desc" class="attn-grid">
  <title id="attn-svg-title">Attention weights between four tokens</title>
  <desc id="attn-svg-desc">A four by four grid. Rows are the position being updated, columns the position being read. The upper right half is blocked by the causal mask. In the bottom row, the position "the" puts most of its weight on "approved" and "bank".</desc>
  <defs>
    <pattern id="attn-hatch" width="7" height="7" patternUnits="userSpaceOnUse" patternTransform="rotate(45)">
      <rect class="attn-hatch-bg" width="7" height="7" />
      <line class="attn-hatch-line" x1="0" y1="0" x2="0" y2="7" />
    </pattern>
  </defs>
  <text x="150" y="26" class="attn-cap">reading from →</text>
  <text x="140" y="60" class="attn-col">The</text>
  <text x="228" y="60" class="attn-col">bank</text>
  <text x="316" y="60" class="attn-col">approved</text>
  <text x="404" y="60" class="attn-col">the</text>

  <text x="96" y="98" class="attn-row">The</text>
  <rect x="108" y="76" width="84" height="40" class="attn-cell" style="--w:1.00" /><text x="150" y="101" class="attn-num">1.00</text>
  <rect x="196" y="76" width="84" height="40" class="attn-mask" />
  <rect x="284" y="76" width="84" height="40" class="attn-mask" />
  <rect x="372" y="76" width="84" height="40" class="attn-mask" />

  <text x="96" y="146" class="attn-row">bank</text>
  <rect x="108" y="124" width="84" height="40" class="attn-cell" style="--w:0.35" /><text x="150" y="149" class="attn-num">0.35</text>
  <rect x="196" y="124" width="84" height="40" class="attn-cell" style="--w:0.65" /><text x="238" y="149" class="attn-num">0.65</text>
  <rect x="284" y="124" width="84" height="40" class="attn-mask" />
  <rect x="372" y="124" width="84" height="40" class="attn-mask" />

  <text x="96" y="194" class="attn-row">approved</text>
  <rect x="108" y="172" width="84" height="40" class="attn-cell" style="--w:0.12" /><text x="150" y="197" class="attn-num">0.12</text>
  <rect x="196" y="172" width="84" height="40" class="attn-cell" style="--w:0.48" /><text x="238" y="197" class="attn-num">0.48</text>
  <rect x="284" y="172" width="84" height="40" class="attn-cell" style="--w:0.40" /><text x="326" y="197" class="attn-num">0.40</text>
  <rect x="372" y="172" width="84" height="40" class="attn-mask" />

  <text x="96" y="242" class="attn-row">the</text>
  <rect x="108" y="220" width="84" height="40" class="attn-cell" style="--w:0.08" /><text x="150" y="245" class="attn-num">0.08</text>
  <rect x="196" y="220" width="84" height="40" class="attn-cell" style="--w:0.31" /><text x="238" y="245" class="attn-num">0.31</text>
  <rect x="284" y="220" width="84" height="40" class="attn-cell" style="--w:0.44" /><text x="326" y="245" class="attn-num">0.44</text>
  <rect x="372" y="220" width="84" height="40" class="attn-cell" style="--w:0.17" /><text x="414" y="245" class="attn-num">0.17</text>

  <text x="472" y="146" class="attn-side">rows sum</text>
  <text x="472" y="164" class="attn-side">to 1.00</text>
  <rect x="108" y="278" width="18" height="14" class="attn-mask" />
  <text x="134" y="290" class="attn-key">blocked by the causal mask — these tokens do not exist yet</text>
  <text x="20" y="180" class="attn-cap" transform="rotate(-90 20 180)">updating ↓</text>
</svg>
<figcaption id="attn-fig-title">Illustrative attention weights for one head in one layer. Each row is a position being updated; each cell is how much of that column's value it takes.</figcaption>
</figure>

<p>Read the bottom row. The position that will predict the next token puts 0.44 on <code class="language-plaintext highlighter-rouge">approved</code> and 0.31 on <code class="language-plaintext highlighter-rouge">bank</code>, and very little on <code class="language-plaintext highlighter-rouge">The</code>. Its updated activation is the values of those four positions, mixed in exactly those proportions. That mixture is why the distribution at the top of this article favours <code class="language-plaintext highlighter-rouge">loan</code>: the position carries <code class="language-plaintext highlighter-rouge">bank</code> and <code class="language-plaintext highlighter-rouge">approved</code> together, and in the training data that combination is followed by a small set of tokens.</p>

<p>Now the second sentence. In <code class="language-plaintext highlighter-rouge">They sat beside the river bank</code>, the position holding <code class="language-plaintext highlighter-rouge">bank</code> has <code class="language-plaintext highlighter-rouge">river</code> available to attend to, and puts weight there instead. Same token, same starting embedding, same weights in the model — a different mixture, and therefore a different activation. <strong>Nothing looked up a definition of <code class="language-plaintext highlighter-rouge">bank</code>.</strong> The disambiguation is a consequence of which positions were available and how strongly each was weighted.</p>

<p>This is also the honest answer to what the intermediate numbers are. They are not a hidden sentence, and they are not meaningless. Each position’s vector is a location in a space the model learned during training, where the operations in later layers can act on the distinctions that mattered for prediction. <code class="language-plaintext highlighter-rouge">bank</code>-after-<code class="language-plaintext highlighter-rouge">approved</code> and <code class="language-plaintext highlighter-rouge">bank</code>-after-<code class="language-plaintext highlighter-rouge">river</code> end up in different regions, because putting them in different regions is what made the model’s predictions less wrong.</p>

<h3 id="attention-runs-many-times-in-parallel-and-that-is-where-heads-come-from">Attention runs many times in parallel, and that is where heads come from</h3>

<p>One round of the calculation above finds one kind of relationship. A model runs many in parallel, each with its own learned <code class="language-plaintext highlighter-rouge">Wq</code>, <code class="language-plaintext highlighter-rouge">Wk</code> and <code class="language-plaintext highlighter-rouge">Wv</code>. Each parallel copy is an <strong>attention head</strong>.</p>

<p>The model does not give every head the full vector. A representation of 4,096 numbers split across 32 heads gives each head 128 numbers to work with — that slice size is the <strong>head dimension</strong>. The heads run independently and their outputs are joined back together before the layer finishes. Different heads end up sensitive to different things; none of them was told what to specialise in.</p>

<p>Now the distinction that decides how much memory a conversation costs.</p>

<p>Queries are used once, in the step that produces the current token, and then discarded. Keys and values have to be kept, because every future token will attend back to them. So model designers noticed they could keep <strong>fewer key/value heads than query heads</strong> and let several query heads share one set. This is <strong>grouped-query attention</strong>, and it is the reason the two counts differ:</p>

<table>
  <thead>
    <tr>
      <th>Llama 3.1 8B</th>
      <th style="text-align: right">Count</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Query heads</td>
      <td style="text-align: right">32</td>
    </tr>
    <tr>
      <td>Key/value heads</td>
      <td style="text-align: right">8</td>
    </tr>
    <tr>
      <td>Head dimension</td>
      <td style="text-align: right">128</td>
    </tr>
    <tr>
      <td>Layers</td>
      <td style="text-align: right">32</td>
    </tr>
  </tbody>
</table>

<p>Four query heads share each key/value head. The saving is not a rounding detail: the cache stores keys and values, so <strong>the key/value head count is what multiplies into memory</strong>, and dropping it from 32 to 8 cuts the cost of every cached token by four.</p>

<p>This is the number that appears in the memory formula later in this article, and in the <a href="/2026/08/30/ai-system-design-for-product-managers.html#derivation">cost arithmetic</a> in the companion guide. When a table says a token costs 128 KB, this is where three of its four factors come from.</p>

<p>A language model also applies a <strong>causal mask</strong>: a position may use earlier tokens, but it cannot use tokens that have not occurred yet. When predicting after <code class="language-plaintext highlighter-rouge">approved</code>, the model may use <code class="language-plaintext highlighter-rouge">The bank approved</code>; it cannot use <code class="language-plaintext highlighter-rouge">the loan</code> before those tokens exist.</p>

<details class="technical-detail">
  <summary>For the extra curious: the attention calculation</summary>

  <p>Starting with the current activation matrix <code class="language-plaintext highlighter-rouge">X</code>, the model produces:</p>

  <div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Q = XWq
K = XWk
V = XWv
</code></pre></div>  </div>

  <p><code class="language-plaintext highlighter-rouge">Wq</code>, <code class="language-plaintext highlighter-rouge">Wk</code> and <code class="language-plaintext highlighter-rouge">Wv</code> are learned parameter matrices. A common compact form of attention is:</p>

  <div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Attention(Q, K, V) = softmax((QKᵀ / √d) + M)V
</code></pre></div>  </div>

  <p><code class="language-plaintext highlighter-rouge">QKᵀ</code> produces query-key compatibility scores. <code class="language-plaintext highlighter-rouge">M</code> applies the causal mask. Softmax converts the permitted scores into weights. Multiplication by <code class="language-plaintext highlighter-rouge">V</code> produces weighted combinations of the value vectors.</p>
</details>

<h2 id="6-many-layers-produce-the-next-token-scores">6. Many layers produce the next-token scores</h2>

<p>Attention is only one part of a transformer layer. A layer also contains operations that transform each position’s values, normalise them and carry earlier information forward.</p>

<p>The model repeats these operations across many layers — 32 of them in the model tabulated above. At each layer, the activation for <code class="language-plaintext highlighter-rouge">bank</code>, <code class="language-plaintext highlighter-rouge">approved</code> and every other position changes as information is combined and transformed.</p>

<p>Each layer has its own attention, and therefore its own keys and values. Nothing is shared between layers. That is why the layer count multiplies into the memory a conversation occupies: the per-token cost is paid once per layer, every layer, for as long as the conversation is held.</p>

<p>After the final layer, the model takes the activation at the <strong>last position</strong> — position 4, <code class="language-plaintext highlighter-rouge">the</code>, the one whose row we read in the grid above — and converts it into one score for every token in the vocabulary. The other three positions were computed too, and during prefill their keys and values are kept, but only the last position’s activation is asked what comes next. Those raw scores are called <strong>logits</strong>. Softmax turns the logits into probabilities, returning us to the table at the beginning of the article.</p>

<p>The complete path is now visible:</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>text
→ token IDs
→ embedding vectors
→ position-aware activations
→ repeated transformer layers
→ vocabulary scores
→ next-token probabilities
→ selection rule
→ selected token
</code></pre></div></div>

<p>Everything up to the probabilities is the model. The step that follows is not, and it is the next section.</p>

<details class="technical-detail">
  <summary>For the extra curious: what a learned matrix transformation does</summary>

  <p>A common neural-network operation is:</p>

  <div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>y = xW
</code></pre></div>  </div>

  <p><code class="language-plaintext highlighter-rouge">x</code> is an input vector, <code class="language-plaintext highlighter-rouge">W</code> is a matrix of learned parameters and <code class="language-plaintext highlighter-rouge">y</code> is the output vector. Each number in <code class="language-plaintext highlighter-rouge">y</code> is a weighted combination of numbers in <code class="language-plaintext highlighter-rouge">x</code>.</p>

  <p>For an entire sequence, <code class="language-plaintext highlighter-rouge">XW</code> applies the transformation to the matrix of token activations. Modern accelerators are designed to perform enormous numbers of these operations efficiently.</p>
</details>

<h2 id="7-a-selection-rule-turns-probabilities-into-one-token">7. A selection rule turns probabilities into one token</h2>

<p>The model’s output is a probability distribution over the whole vocabulary. A distribution is not a token. Some rule has to reduce thousands of candidates to the single token that gets appended, and that rule lives in the serving layer, not in the model weights.</p>

<p>The simplest rule is <strong>greedy decoding</strong>: always take the highest-probability token. Given the distribution from section 1, greedy decoding selects <code class="language-plaintext highlighter-rouge">loan</code> every time.</p>

<p>Greedy decoding is not always what a product wants. It makes the model repeat itself on open-ended tasks and produces one fixed answer for one fixed prompt. The alternative is <strong>sampling</strong>: draw a token at random, in proportion to the probabilities. <code class="language-plaintext highlighter-rouge">loan</code> is then selected about 31% of the time and <code class="language-plaintext highlighter-rouge">request</code> about 14%.</p>

<p>Three controls shape that draw.</p>

<p><strong>Temperature</strong> reshapes the distribution before the draw. The logits are divided by a number <code class="language-plaintext highlighter-rouge">T</code> before softmax. Below 1, the gaps between candidates widen and the leading token dominates. Above 1, the distribution flattens and unlikely tokens gain a real chance. At <code class="language-plaintext highlighter-rouge">T = 0</code> the rule collapses back to greedy decoding.</p>

<table>
  <thead>
    <tr>
      <th>Temperature</th>
      <th>Effect on our example</th>
      <th>Typical use</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>0</td>
      <td><code class="language-plaintext highlighter-rouge">loan</code> every time</td>
      <td>Extraction, classification, structured output</td>
    </tr>
    <tr>
      <td>~0.7</td>
      <td><code class="language-plaintext highlighter-rouge">loan</code> usually, other plausible tokens sometimes</td>
      <td>Assistants, general chat</td>
    </tr>
    <tr>
      <td>~1.3</td>
      <td>The tail becomes reachable</td>
      <td>Brainstorming, variation</td>
    </tr>
  </tbody>
</table>

<p><strong>Top-k</strong> limits the draw to the <code class="language-plaintext highlighter-rouge">k</code> highest-probability tokens. <strong>Top-p</strong>, or nucleus sampling, limits it to the smallest set of tokens whose probabilities sum to <code class="language-plaintext highlighter-rouge">p</code>. Both exist to cut off the long tail of near-zero candidates that temperature alone can make reachable — the tokens responsible for an answer that starts sensibly and then goes strange.</p>

<h3 id="temperature-zero-is-not-a-determinism-guarantee">Temperature zero is not a determinism guarantee</h3>

<p>Product teams routinely assume <code class="language-plaintext highlighter-rouge">temperature = 0</code> means identical output for identical input. It removes the deliberate randomness, which is the largest source of variation, but it does not make the arithmetic reproducible.</p>

<p>Floating-point addition is not associative, so the order in which an accelerator sums values changes the last bits of a result. That order depends on how requests were grouped into a batch, which depends on the traffic arriving alongside yours. Mixture-of-experts routing and changes to the serving stack add further variation. Occasionally two candidate tokens sit close enough that the last-bit difference flips which one leads, and the outputs diverge from that token onward.</p>

<p>Design for it. If a downstream system requires an exact match, validate the output’s structure rather than comparing it byte for byte, and pin the behaviour you need in tests as a property, not a fixture.</p>

<details class="technical-detail">
  <summary>For the extra curious: where temperature enters the calculation</summary>

  <p>Section 6 ended with softmax applied to the logits. Temperature is a division applied first:</p>

  <div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>P = softmax(logits / T)
</code></pre></div>  </div>

  <p>Because softmax is exponential, dividing by a small <code class="language-plaintext highlighter-rouge">T</code> multiplies the ratio between any two candidates. If one logit exceeds another by 2.0, then at <code class="language-plaintext highlighter-rouge">T = 1</code> their probability ratio is <code class="language-plaintext highlighter-rouge">e²</code> ≈ 7.4; at <code class="language-plaintext highlighter-rouge">T = 0.5</code> it is <code class="language-plaintext highlighter-rouge">e⁴</code> ≈ 54.6. <code class="language-plaintext highlighter-rouge">T = 0</code> is not evaluated as a division — implementations special-case it to selecting the maximum.</p>
</details>

<h2 id="8-reading-the-input-and-writing-the-output-behave-differently">8. Reading the input and writing the output behave differently</h2>

<p>Everything to this point described producing <strong>one</strong> token. Our four input tokens were all known in advance, so the model can push all four through its layers together — position 4 does not have to wait for position 2, because position 2’s token was already there. Serving systems call this <strong>prefill</strong>.</p>

<p>Generating is different. Having selected <code class="language-plaintext highlighter-rouge">loan</code>, the model appends it and runs the calculation again on five tokens to get the sixth. It cannot start the sixth before the fifth exists, because the sixth attends to it. This phase is <strong>decode</strong>, and it is strictly one token at a time:</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>prefill   The bank approved the          ← 4 tokens, together
decode    ... loan                       ← token 5
decode    ... loan application           ← token 6, only after 5
decode    ... loan application yesterday ← token 7, only after 6
</code></pre></div></div>

<p>That asymmetry is the whole reason input and output are priced and timed differently. Input is a wide, parallel pass. Output is a queue.</p>

<p>Compare:</p>

<table>
  <thead>
    <tr>
      <th>Request</th>
      <th style="text-align: right">Input</th>
      <th style="text-align: right">Output</th>
      <th>Dominant work</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Contract extraction</td>
      <td style="text-align: right">50,000 tokens</td>
      <td style="text-align: right">200 tokens</td>
      <td>Processing known input</td>
    </tr>
    <tr>
      <td>Long-form writing</td>
      <td style="text-align: right">2,000 tokens</td>
      <td style="text-align: right">4,000 tokens</td>
      <td>Sequential generation</td>
    </tr>
  </tbody>
</table>

<p>The same total token count does not imply the same latency or hardware behaviour — which is the observation the <a href="/2026/08/30/ai-system-design-for-product-managers.html">companion guide</a> is built on.</p>

<h3 id="some-of-the-generated-tokens-are-not-the-answer">Some of the generated tokens are not the answer</h3>

<p>Nothing in the loop so far distinguishes a token meant for the reader from a token the model produces to work something out. Both are generated the same way, one at a time, each conditioned on everything before it.</p>

<p>Modern models exploit that. Given a hard problem, a model can generate a stretch of tokens that reason toward the answer — considering an approach, rejecting it, trying another — and only then generate the response. These are usually called <strong>reasoning tokens</strong> or <strong>thinking tokens</strong>. They are ordinary output tokens: same sequential generation, same per-token cost, same contribution to the KV cache. The only thing that makes them different is that the product usually does not show them.</p>

<p>This has three consequences a product team feels directly:</p>

<ol>
  <li><strong>They are on the bill as output.</strong> Output tokens are the expensive ones. A request whose visible answer is 200 tokens may have generated 4,000 to get there, and it is billed accordingly.</li>
  <li><strong>They are slow in the same way any output is slow.</strong> Decode is sequential, so reasoning is time the user spends waiting before the answer begins.</li>
  <li><strong>They are not always worth it.</strong> On a hard problem, more reasoning measurably improves the answer. On classification or extraction it is close to pure cost.</li>
</ol>

<p>Because the tradeoff depends on the task rather than the prompt, providers expose it as a request parameter — a thinking budget, or an effort level — rather than something you phrase your way into. It is one of the few dials where the same prompt, unchanged, can differ several-fold in both cost and latency. The <a href="/2026/08/30/ai-system-design-for-product-managers.html#config">companion guide</a> shows where that dial sits in a real request.</p>

<h3 id="the-kv-cache-prevents-repeated-work">The KV cache prevents repeated work</h3>

<p>Look again at the grid in section 5. To produce token 5, the model needed the keys and values of positions 1 to 4. To produce token 6, it needs positions 1 to 5 — the same four, plus one. Positions 1 to 4 have not changed: <code class="language-plaintext highlighter-rouge">The</code> cannot see anything after it, so its key and value are the same as they were.</p>

<p>Recomputing them anyway, for every generated token, in every layer, would be enormous waste. So the serving system keeps them. That retained state is the <strong>KV cache</strong>. Each new token adds one column to the grid and reuses every column already there.</p>

<p>This is also why the cache stores keys and values but not queries: a query is used in the row being computed and never referred to again, which is exactly the asymmetry that makes grouped-query attention worth doing.</p>

<p>The KV cache is not the conversation transcript. The transcript is text that an application can store in a database. The cache is model-specific numerical state created from a particular sequence.</p>

<p>It grows as the sequence grows and occupies fast accelerator memory. If it is discarded, the transcript still exists, but the model has to process that text again to rebuild usable state.</p>

<h3 id="where-the-cache-can-live">Where the cache can live</h3>

<p>An active model uses several kinds of memory. They are not interchangeable.</p>

<table>
  <thead>
    <tr>
      <th>Location</th>
      <th>What it can hold</th>
      <th>What must happen before inference continues</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Accelerator memory, usually HBM</td>
      <td>Model weights and live KV-cache state</td>
      <td>Nothing; the accelerator can use it directly</td>
    </tr>
    <tr>
      <td>Host memory, usually DDR</td>
      <td>Parked numerical cache state</td>
      <td>Transfer the state back to accelerator memory</td>
    </tr>
    <tr>
      <td>Application database or storage</td>
      <td>Transcript as text</td>
      <td>Process the text again to recreate model state</td>
    </tr>
  </tbody>
</table>

<p><strong>HBM</strong>, or high-bandwidth memory, sits close to the accelerator’s compute units. It is fast and scarce. The model weights, current calculations and live KV caches compete for this capacity.</p>

<p><strong>DDR</strong>, the system memory attached to the host machine, is larger and cheaper but slower. A serving system can move idle cache state there, but it must transfer that state back before the accelerator can use it again.</p>

<p>A <strong>database</strong> stores a different object. It can preserve the conversation text, but not the numerical state created inside the model. Returning from text requires computation, not merely a memory copy.</p>

<p>The exact capacity and transfer time depend on the hardware, interconnect, model and cache size. The product consequence is stable: keeping a conversation ready consumes scarce memory; parking it adds transfer delay; discarding it requires repeated input processing.</p>

<h3 id="why-cached-state-disappears">Why cached state disappears</h3>

<p>There are three distinct events that are often described loosely as a cache miss.</p>

<ul>
  <li><strong>TTL expiry:</strong> Retained state is removed after it has gone unused for a configured period. The next request starts more slowly because state must be restored or rebuilt.</li>
  <li><strong>LRU eviction:</strong> Memory fills before the timer expires, so the system removes the least recently used state to make room. Quick follow-ups may become slower during busy periods.</li>
  <li><strong>Preemption:</strong> Active generation is paused or rescheduled so its resources can be used elsewhere. The user may see an answer stall after it has already begun.</li>
</ul>

<p>These mechanisms have different symptoms and should be measured separately. Time to first token helps reveal cold or rebuilt state. Time between output tokens helps reveal interruptions during generation.</p>

<h3 id="why-an-unchanged-beginning-can-be-reused">Why an unchanged beginning can be reused</h3>

<p>The state at each position depends on the tokens before it — and, because of the causal mask, on nothing after it. That is the property prompt caching exploits. Consider two requests:</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>The bank approved the loan
The bank approved the mortgage
</code></pre></div></div>

<p>The keys and values for <code class="language-plaintext highlighter-rouge">The bank approved the</code> are identical in both. Not similar — identical, because none of those four positions was allowed to see the fifth token when its state was computed. A serving system can reuse all four and begin work at position 5.</p>

<p>Reverse it and the property disappears:</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>The bank approved the loan
On Tuesday the bank approved the loan
</code></pre></div></div>

<p>Every token has shifted position and each one now has different tokens before it. Nothing is reusable, despite the two sequences sharing almost every word.</p>

<p>If a request changes near the beginning, later cached state may no longer be valid because later positions were calculated using the changed earlier token. This is why prompt caching generally reuses an exact shared prefix rather than matching identical fragments anywhere in a request.</p>

<details class="technical-detail">
  <summary>For the extra curious: why cache memory grows</summary>

  <p>A simplified KV-cache estimate is:</p>

  <div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>2 × layers × cached tokens × KV heads × head dimension × bytes per number
</code></pre></div>  </div>

  <p>The factor of two represents keys and values. Exact requirements vary by architecture, precision, quantisation and serving technique. For a fixed setup, the important relationship is that KV-cache memory grows approximately linearly with sequence length.</p>
</details>

<h2 id="9-images-and-audio-enter-through-different-front-ends">9. Images and audio enter through different front ends</h2>

<p>Text starts with a tokenizer because it contains discrete text symbols. Images and audio need different encoders.</p>

<p>A typical image path is:</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>pixels
→ image patches or vision-encoder inputs
→ visual feature vectors
→ model-compatible activations
</code></pre></div></div>

<p>A typical audio path is:</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>waveform
→ time segments or spectral features
→ audio feature vectors
→ model-compatible activations
</code></pre></div></div>

<p>Architectures differ, but the shared destination is an ordered collection of numerical representations the model can process.</p>

<h2 id="what-to-keep-in-your-head">What to keep in your head</h2>

<p>You do not need to memorise every matrix to understand the path:</p>

<ol>
  <li>The model calculates probabilities for possible next tokens.</li>
  <li>A tokenizer turns text into vocabulary IDs.</li>
  <li>Each ID selects a learned starting vector called an embedding.</li>
  <li>Activations are those numerical representations as they change through the model.</li>
  <li>Position information preserves order.</li>
  <li>Attention lets positions combine information from permitted earlier positions, running as many parallel heads — of which the key/value heads are the ones that cost memory.</li>
  <li>Repeated layers produce scores for the next token.</li>
  <li>A selection rule — greedy, or sampling shaped by temperature, top-k and top-p — reduces that distribution to one token.</li>
  <li>Some generated tokens are reasoning rather than answer; they cost and take time like any other output token.</li>
  <li>Generation repeats the calculation one selected token at a time.</li>
  <li>The KV cache retains earlier attention state so it does not have to be recreated at every step.</li>
</ol>

<p>Run the sentence through all of it once. <code class="language-plaintext highlighter-rouge">The bank approved the</code> becomes four IDs; each ID pulls a learned row; position information keeps them in order; attention lets position 4 mix in 0.44 of <code class="language-plaintext highlighter-rouge">approved</code> and 0.31 of <code class="language-plaintext highlighter-rouge">bank</code>; 32 layers repeat that; the last position’s vector becomes 100,000-odd scores; softmax makes them probabilities; a selection rule picks <code class="language-plaintext highlighter-rouge">loan</code>; the four positions’ keys and values are kept so the next token costs one column instead of five. That is the whole machine. Everything else is scale.</p>

<p>These mechanics matter because they connect visible product behaviour to computation: long inputs affect the initial wait, long outputs take sequential time, growing conversations occupy more working memory, a changed prompt beginning can prevent reuse, and the selection rule — not the model — decides how much the same prompt varies.</p>]]></content><author><name>Yash Tambawala</name></author><summary type="html"><![CDATA[A step-by-step explanation of tokens, embeddings, attention, generation and model memory—without assuming a machine-learning background]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://yashtambawala.com/assets/og/how-an-ai-model-turns-input-into-output.png" /><media:content medium="image" url="https://yashtambawala.com/assets/og/how-an-ai-model-turns-input-into-output.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">AI System Design</title><link href="https://yashtambawala.com/2026/08/30/ai-system-design-for-product-managers.html" rel="alternate" type="text/html" title="AI System Design" /><published>2026-08-30T09:00:00+05:30</published><updated>2026-08-30T09:00:00+05:30</updated><id>https://yashtambawala.com/2026/08/30/ai-system-design-for-product-managers</id><content type="html" xml:base="https://yashtambawala.com/2026/08/30/ai-system-design-for-product-managers.html"><![CDATA[<section class="think-orientation" aria-labelledby="orientation-title">
  <p class="think-kicker">The premise</p>
  <h2 id="orientation-title">A model call is not a product.</h2>
  <p>A document pipeline and a coding agent might each process one million tokens. The document pipeline reads independent files overnight and returns short results. The coding agent carries one task for hours, repeatedly pausing for tools while its working information grows. The token count is the same. Almost every important system decision is different.</p>
  <p>The difference is measurable. On one NVIDIA H100, the document pipeline runs ten jobs at a time through a queue that drains before morning. The coding agent gets two sessions on the same card, all afternoon. <a href="#arithmetic">The arithmetic is worked out below</a>, and it is the reason a token budget is not a capacity plan.</p>
  <p class="sys-orientation-note">Two forces have made this question urgent rather than academic. Products stopped making single model calls and started running agents that loop through tools for hours. And the same products now have to run partly on a phone, a robot or a factory gateway, where the memory budget is not elastic and there is no larger machine to move to.</p>
  <p class="sys-orientation-note">This guide follows one real request, decides where its information should live, then uses five questions to compare the systems behind chat, agent fleets, coding, on-device assistants, industrial edge, monitoring, documents and voice.</p>
</section>

<section class="think-act think-act--mechanics" id="mechanics" data-act="mechanics" aria-labelledby="mechanics-title">
  <div class="think-act-head">
    <p class="think-act-number">Act 01 <span>·</span> Draw the system</p>
    <h2 id="mechanics-title">Draw the whole system.</h2>
    <p>Begin with the path from a user’s need to a useful result. The model sits in that path; it does not replace it.</p>
  </div>

  <div class="sys-request-path" aria-label="The path of one support request">
    <div><span>01</span><strong>User need</strong><p>“Where is order 4821, and what happens if it is late?”</p></div>
    <i aria-hidden="true">→</i>
    <div><span>02</span><strong>Context builder</strong><p>Fetch the order, select the policy and carry forward relevant conversation.</p></div>
    <i aria-hidden="true">→</i>
    <div><span>03</span><strong>Model and tools</strong><p>Interpret the request, decide whether another lookup is needed and form an answer.</p></div>
    <i aria-hidden="true">→</i>
    <div><span>04</span><strong>Product response</strong><p>Show the status, explain the policy and offer the next action.</p></div>
  </div>

  <div class="think-lens-row">
    <article class="think-lens">
      <p class="think-thinker">One support request</p>
      <h3>The application decides what the model gets to know.</h3>
      <p>The model does not automatically know the current location of order 4821 or the company’s delayed-delivery policy. The application has to retrieve those facts and decide which parts belong in this request.</p>
      <p>A weak system sends the entire customer profile, every order, all previous tickets, the complete policy handbook and the full conversation. A better system sends the current order, the relevant policy section, a small amount of active context and clear instructions.</p>
      <p>The second request is not merely cheaper. It is easier to keep current, easier to audit, less likely to expose unrelated data and less likely to distract the model with conflicting evidence.</p>
      <aside class="think-example"><span>First decision</span>Do not ask, “How much can the model fit?” Ask, “What is the smallest complete set of information required for this decision?”</aside>
      <blockquote class="think-question">Who decides what enters the request—and by what rule?</blockquote>
    </article>

    <aside class="think-diagram sys-diagram sys-diagram--pass" data-pass-demo="" aria-label="Animation showing supplied information being processed before a response is produced">
      <span class="diagram-label">What the user waits for</span>
      <div class="pass-phase" data-pass-phase="">Idle</div>
      <div class="pass-block" aria-hidden="true"><span class="pass-block-label">Information supplied</span><div class="pass-grid" data-pass-grid=""></div></div>
      <div class="pass-arrow" aria-hidden="true">↓</div>
      <div class="pass-block" aria-hidden="true"><span class="pass-block-label">Response produced</span><div class="pass-out" data-pass-out=""></div></div>
      <div class="pass-meters" aria-hidden="true"><div><span>Processing input</span><i data-pass-meter="compute"></i></div><div><span>Producing output</span><i data-pass-meter="memory"></i></div></div>
      <button class="think-print sys-button" type="button" data-pass-play="">Run the request</button>
      <p>The initial wait and the time spent producing the answer are different parts of the experience.<br /><strong>Measure them separately.</strong></p>
    </aside>
  </div>

  <div class="think-memory sys-levers">
    <div class="think-memory-head">
      <p class="think-kicker">Five questions</p>
      <h3>Describe the work before choosing the model.</h3>
    </div>
    <ol>
      <li><strong>How much information is needed?</strong> A single order lookup is small. A repository-wide coding task may carry files, logs, errors and decisions for hours.</li>
      <li><strong>How long does the work continue?</strong> A one-shot extraction ends cleanly. An agent that loops through tools for two hours must control growth and decide what it can forget.</li>
      <li><strong>How long are the gaps?</strong> A two-second tool call and a thirty-minute human pause should not keep expensive working state in the same way. An agent working autonomously has almost no gaps at all, which is why it is expensive to hold.</li>
      <li><strong>How quickly must it respond?</strong> A batch result can wait. A voice assistant cannot leave a person in silence. A robot arm has a hard deadline.</li>
      <li><strong>Where does it run?</strong> A datacentre lets you add memory. A phone, a robot or a plant-floor gateway does not, and its budget is shared with everything else the device is doing.</li>
    </ol>
    <p class="sys-levers-note">The fifth question used to be assumed away. It no longer can be: the same assistant increasingly answers on the device when it can and escalates to a server when it must, and those two paths are different systems wearing one product name.</p>
  </div>
</section>

<section class="think-act think-act--shape" id="shape" data-act="shape" aria-labelledby="shape-title">
  <div class="think-act-head">
    <p class="think-act-number">Act 02 <span>·</span> Place the information</p>
    <h2 id="shape-title">Put information in the right place.</h2>
    <p>Not everything the product knows should become conversation history.</p>
  </div>

  <div class="sys-decide" id="decide">
    <div class="sys-decide-head">
      <p class="think-kicker">The same request, decided line by line</p>
      <h3>“Where is order 4821, and what happens if it is late?”</h3>
      <p>Here is everything the product knows that could plausibly go into this one request, and the decision for each. The column that does the work is the last one — most of what is available should not be there.</p>
    </div>

    <div class="sys-decide-table" role="table" aria-label="What enters the request for order 4821">
      <div class="sys-decide-head-row" role="row"><span role="columnheader">What it is</span><span role="columnheader">How often it changes</span><span role="columnheader">Where it lives</span><span role="columnheader">In this request?</span></div>
      <div role="row"><strong data-label="What it is" role="cell">Support instructions and refund policy</strong><span data-label="How often it changes" role="cell">Rarely — a versioned document</span><span data-label="Where it lives" role="cell">The prompt prefix</span><b data-label="In this request?" class="is-in" role="cell">Yes — first, and byte-identical every time</b></div>
      <div role="row"><strong data-label="What it is" role="cell">Status of order 4821</strong><span data-label="How often it changes" role="cell">Constantly</span><span data-label="Where it lives" role="cell">Order service</span><b data-label="In this request?" class="is-in" role="cell">Yes — fetched now, never carried</b></div>
      <div role="row"><strong data-label="What it is" role="cell">The late-delivery clause</strong><span data-label="How often it changes" role="cell">Quarterly</span><span data-label="Where it lives" role="cell">Policy store, searched</span><b data-label="In this request?" class="is-in" role="cell">Yes — the clause, not the handbook</b></div>
      <div role="row"><strong data-label="What it is" role="cell">What has been tried this session</strong><span data-label="How often it changes" role="cell">Every turn</span><span data-label="Where it lives" role="cell">The conversation</span><b data-label="In this request?" class="is-in" role="cell">Yes — while the task is live</b></div>
      <div role="row"><strong data-label="What it is" role="cell">The customer’s other eleven orders</strong><span data-label="How often it changes" role="cell">Constantly</span><span data-label="Where it lives" role="cell">Order service</span><b data-label="In this request?" role="cell">No — not this decision</b></div>
      <div role="row"><strong data-label="What it is" role="cell">The full policy handbook</strong><span data-label="How often it changes" role="cell">Quarterly</span><span data-label="Where it lives" role="cell">Policy store</span><b data-label="In this request?" role="cell">No — retrieve, don’t carry</b></div>
      <div role="row"><strong data-label="What it is" role="cell">Complete profile and ticket history</strong><span data-label="How often it changes" role="cell">Constantly</span><span data-label="Where it lives" role="cell">CRM and ticketing</span><b data-label="In this request?" role="cell">No — last relevant turn only</b></div>
    </div>

    <div class="sys-decide-verdict">
      <article><span>What the lazy version sends</span><strong>Everything above</strong><p>Large, expensive on every turn, stale the moment the order moves, impossible to audit after a bad refund, and full of material that competes for the model’s attention.</p></article>
      <article><span>What this version sends</span><strong>Four of the seven</strong><p>Smaller, current at the moment of the decision, and explainable line by line when someone asks why the refund was approved.</p></article>
    </div>

    <p class="sys-decide-note">Notice that “where it lives” and “how often it changes” decide the answer between them, and neither is a question about the model. A fact with an owner and an update rule belongs with its owner. The conversation carries the active problem — not a copy of the product’s database.</p>
  </div>

  <div class="sys-dual-copy">
    <article class="think-lens">
      <p class="think-thinker">Conversation is not memory</p>
      <h3>Text, model state and durable facts are different objects.</h3>
      <p>The application can store a transcript as text. While the model processes that text, it also creates temporary numerical working state—commonly called the KV cache—that makes an immediate continuation efficient. Keeping that state ready consumes fast memory; losing it means processing the text again.</p>
      <p>Neither object should be confused with durable product state. An order status belongs in the order system. A camera event belongs in an event store. A code change belongs in a file. The model can retrieve these facts when they matter instead of dragging them through every turn.</p>
      <aside class="think-example"><span>Practical rule</span>If information has a clear owner, structure and update rule, store it there. Conversation should carry the active problem, not become the database for the entire product.</aside>
      <blockquote class="think-question">If this conversation disappeared, which facts would the product lose?</blockquote>
    </article>

    <article class="think-lens">
      <p class="think-thinker">Long tasks</p>
      <h3>Compaction is a decision about forgetting.</h3>
      <p>A coding agent may accumulate file contents, test output, errors, diffs and failed approaches until the conversation becomes unwieldy. Shortening it into a summary creates room, but anything omitted is no longer available unless it was written somewhere durable.</p>
      <p>The asymmetry is what matters: source files and test results can always be read again, so losing them costs a tool call. The reasoning — why this approach was abandoned, what the error actually meant — existed only in the conversation. Lose that and the agent repeats the work it already did.</p>
      <blockquote class="think-question">What must survive when this task is shortened or restarted?</blockquote>
    </article>
  </div>

  <div class="sys-approaches" id="approaches">
    <div class="sys-approaches-head">
      <p class="think-kicker">Approaches</p>
      <h3>Seven ways to keep the request small.</h3>
      <p>“Externalise state and compact strategically” is the right instinct and a useless instruction on its own. These are the mechanisms underneath it, each with the price it charges.</p>
    </div>
    <ol class="sys-approach-list">
      <li>
        <span>01</span>
        <strong>Give every fact a system of record</strong>
        <p>Order status lives in the order service, entitlements in billing, events in the event store. The conversation carries an identifier, not a copy. Nothing in the transcript can go stale, because the transcript no longer claims to know.</p>
        <em>Cost: a lookup on the critical path. Budget it in your first-response time.</em>
      </li>
      <li>
        <span>02</span>
        <strong>Retrieve at decision time, not at session start</strong>
        <p>Preloading the customer’s profile “in case it is needed” pays for it on every turn and is wrong the moment it changes. Fetch the two fields this decision requires, when it requires them.</p>
        <em>Cost: the model must be able to ask. This is what tool definitions are for.</em>
      </li>
      <li>
        <span>03</span>
        <strong>Keep tool output out of the transcript by default</strong>
        <p>A test run, a query result or a page of logs can be tens of thousands of tokens, most of it irrelevant. Write the full result somewhere addressable, put a short structured summary and a handle in the conversation, and let the model request the slice it needs.</p>
        <em>Cost: an extra round trip when the model does need the detail. Usually far cheaper than carrying it every turn thereafter.</em>
      </li>
      <li>
        <span>04</span>
        <strong>Isolate work in sub-agents</strong>
        <p>A narrow task — search this codebase, check these three invoices — can run in its own context with its own brief and return only its conclusion. The exploration never enters the parent conversation, so the parent does not carry it for the next two hours.</p>
        <em>Cost: the parent cannot see the reasoning it did not receive. Make the sub-agent return its evidence, not just its verdict.</em>
      </li>
      <li>
        <span>05</span>
        <strong>Compact to a schema, not to a summary</strong>
        <p>Free-text summarisation loses whatever the summariser judged unimportant, which is exactly the thing you will need. Compact into fixed fields instead: goal, constraints, decisions made and why, files changed, tests passing, tests failing, open problem, next action.</p>
        <em>Cost: schema design. The upside is that a fixed shape can be tested, and a missing field is visible.</em>
      </li>
      <li>
        <span>06</span>
        <strong>Write progress into durable artifacts as you go</strong>
        <p>If the milestone exists in a file, a commit, a task record or a test result, compaction cannot destroy it. If it exists only in the conversation, compaction is a data-loss event.</p>
        <em>Cost: the discipline to checkpoint before the context is nearly full rather than after.</em>
      </li>
      <li>
        <span>07</span>
        <strong>Order the request by rate of change</strong>
        <p>Stable instructions first, then slow-moving reference material, then session state, then the current turn. The shared opening stays byte-identical across requests, which is the only condition under which it can be reused.</p>
        <em>Cost: none, and it is the most commonly skipped step in the list.</em>
      </li>
    </ol>
  </div>

  <article class="think-lens sys-tension">
    <p class="think-thinker">The trade-off nobody mentions</p>
    <h3>Dynamic retrieval and prefix reuse pull in opposite directions.</h3>
    <p>Fetching authoritative rules at request time keeps them current. Reusing a cached prefix requires the beginning of the request to be identical to last time. Do both carelessly and you get neither: a freshly retrieved policy paragraph, injected above the instructions, changes byte one of the request and invalidates every cached position after it. The system is now current and slow, and the metric that moved is first-response time.</p>
    <p>The resolution is placement, not choice. Anything stable goes first and stays byte-identical. Anything freshly fetched goes after it, as late in the request as the task allows. A timestamp, a session ID or a personalised greeting at the top of a prompt is one of the most expensive lines a product can ship, and it is almost always there by accident.</p>
    <aside class="think-example"><span>How to find it</span>Diff two consecutive real requests from the same session. The position of the first differing byte is your reuse ceiling. Most teams have never looked.</aside>
    <blockquote class="think-question">In your prompt, what is the first thing that changes—and does it need to be there?</blockquote>
  </article>

  <div class="sys-serving-bridge">
    <div class="sys-serving-bridge-head">
      <p class="think-kicker">What happens between turns</p>
      <h3>Keeping a conversation ready is a memory decision.</h3>
      <p>The product sees one conversation. The serving system sees model state competing for a limited amount of fast memory.</p>
    </div>

    <div class="think-lens-row">
      <article class="think-lens">
        <p class="think-thinker">Three places</p>
        <h3>Ready, parked or rebuilt.</h3>
        <p>Between two turns, a conversation’s working state is sitting in one of three places, and the choice is a pricing decision the product team rarely knows it is making. Keeping it ready costs scarce accelerator memory. Parking it in cheaper host memory costs a transfer before the next turn can start. Discarding it costs nothing until the user returns, then costs a full reprocessing of the transcript.</p>
        <p>The serving team usually owns this setting. The product team owns the fact it optimises for: how long a user actually takes to reply. Those two groups often never speak.</p>
        <aside class="think-example"><span>What to bring</span>Not an average. Bring the distribution of gaps between turns, and the share of sessions that resume after the current retention window expires.</aside>
      </article>

      <aside class="think-diagram sys-diagram sys-diagram--tiers" data-tier-demo="" aria-label="Interactive showing where model working state can live between turns">
        <span class="diagram-label">Where the working state is</span>
        <div class="tier-list">
          <button type="button" class="tier is-active" data-tier="hbm"><em>Accelerator memory · HBM</em><strong>Working state, ready</strong><span class="tier-bar"><i style="--fill: 3%"></i></span><span class="tier-cost">Continue directly</span></button>
          <button type="button" class="tier" data-tier="ddr"><em>Host memory · DDR</em><strong>Working state, parked</strong><span class="tier-bar"><i style="--fill: 30%"></i></span><span class="tier-cost">Transfer it back</span></button>
          <button type="button" class="tier" data-tier="db"><em>Application storage</em><strong>Transcript as text</strong><span class="tier-bar"><i style="--fill: 100%"></i></span><span class="tier-cost">Process it again</span></button>
        </div>
        <div class="tier-detail" data-tier-detail="">The working state is already where the model can use it, so the conversation can continue without a recovery step. This is also the scarcest place to keep it.</div>
      </aside>
    </div>

    <p class="sys-companion-link">Working state disappears for three different reasons—an idle timer expiring, a busy machine evicting the least recently used state, and generation being preempted mid-answer. They produce three different user complaints, and <a href="#levers">the symptom table at the end of this guide</a> separates them. For why they exist, along with the attention, KV-cache and exact-prefix mechanics underneath, read <a href="/2026/08/31/how-an-ai-model-turns-input-into-output.html">How an AI Model Turns Input Into Output</a>.</p>
  </div>
</section>

<section class="think-act think-act--workloads" id="workloads" data-act="workloads" aria-labelledby="workloads-title">
  <div class="think-act-head">
    <p class="think-act-number">Act 03 <span>·</span> Compare the workloads</p>
    <h2 id="workloads-title">The same model can sit inside very different products.</h2>
    <p>Change the information, duration, gaps, latency requirement or location and the architecture changes with it.</p>
  </div>

  <div class="sys-classifier" data-classifier="" aria-label="Interactive comparing the workload shape of six AI products">
    <div class="sys-classifier-head">
      <p class="think-kicker">Compare the workload</p>
      <h3>Pick a product. Watch the five constraints move.</h3>
      <p>The model may be similar. The system around it is not.</p>
    </div>
    <div class="frame-tabs classifier-choices" role="tablist" aria-label="Workloads">
      <button type="button" role="tab" data-workload="chat" class="is-active" aria-selected="true">Chat</button>
      <button type="button" role="tab" data-workload="agent" aria-selected="false">Agent fleet</button>
      <button type="button" role="tab" data-workload="coding" aria-selected="false">Coding</button>
      <button type="button" role="tab" data-workload="device" aria-selected="false">On-device</button>
      <button type="button" role="tab" data-workload="industrial" aria-selected="false">Industrial edge</button>
      <button type="button" role="tab" data-workload="cctv" aria-selected="false">Monitoring</button>
      <button type="button" role="tab" data-workload="batch" aria-selected="false">Documents</button>
      <button type="button" role="tab" data-workload="voice" aria-selected="false">Voice</button>
    </div>
    <div class="classifier-panel">
      <div class="classifier-dials">
        <div class="dial"><span>Information</span><div class="bayes-track"><i data-dial="context"></i></div><em data-dial-value="context">Small, grows slowly</em></div>
        <div class="dial"><span>Duration</span><div class="bayes-track"><i data-dial="duration"></i></div><em data-dial-value="duration">Minutes</em></div>
        <div class="dial"><span>Gap between steps</span><div class="bayes-track"><i data-dial="gap"></i></div><em data-dial-value="gap">Tens of seconds to minutes</em></div>
        <div class="dial"><span>Latency pressure</span><div class="bayes-track"><i data-dial="latency"></i></div><em data-dial-value="latency">Fast first word</em></div>
        <div class="dial"><span>Memory elasticity</span><div class="bayes-track"><i data-dial="place"></i></div><em data-dial-value="place">Server-side, can add capacity</em></div>
      </div>
      <div class="classifier-verdict">
        <span class="diagram-label">Design consequence</span>
        <p data-classifier-verdict="">Short sessions, long gaps and modest context. This is the ordinary chat workload most serving systems expect.</p>
        <div class="frame-facts" data-classifier-tags=""></div>
      </div>
    </div>

    <div class="classifier-static" aria-label="Every workload compared, as a list">
      <article data-workload-data="chat">
        <h4>Chat</h4>
        <dl>
          <dt>Information</dt><dd data-dial-src="context" data-level=".22">Small, grows slowly</dd>
          <dt>Duration</dt><dd data-dial-src="duration" data-level=".3">Minutes</dd>
          <dt>Gap between steps</dt><dd data-dial-src="gap" data-level=".65">Tens of seconds to minutes</dd>
          <dt>Latency pressure</dt><dd data-dial-src="latency" data-level=".6">Fast first word</dd>
          <dt>Memory elasticity</dt><dd data-dial-src="place" data-level=".15">Server-side, can add capacity</dd>
        </dl>
        <p data-verdict-src="">Chat begins small, grows gradually and pauses unpredictably while people read and think. Design for quick follow-ups without assuming every user will return soon.</p>
        <ul><li>Stable shared opening</li><li>Relevant history only</li><li>Measure return gaps</li></ul>
      </article>
      <article data-workload-data="agent">
        <h4>Agent fleet</h4>
        <dl>
          <dt>Information</dt><dd data-dial-src="context" data-level="1">Very large, grows every step</dd>
          <dt>Duration</dt><dd data-dial-src="duration" data-level=".9">Hours, unattended</dd>
          <dt>Gap between steps</dt><dd data-dial-src="gap" data-level=".04">None — it never waits for a human</dd>
          <dt>Latency pressure</dt><dd data-dial-src="latency" data-level=".2">Cost per task, not first word</dd>
          <dt>Memory elasticity</dt><dd data-dial-src="place" data-level=".15">Server-side, can add capacity</dd>
        </dl>
        <p data-verdict-src="">An autonomous agent holds a large context for hours with no idle gaps, and each step resends everything before it. Cache reuse and context curation decide whether the product is economically viable, long before model quality does.</p>
        <ul><li>Protect the prefix</li><li>Isolate in sub-agents</li><li>Cost per completed task</li></ul>
      </article>
      <article data-workload-data="coding">
        <h4>Coding</h4>
        <dl>
          <dt>Information</dt><dd data-dial-src="context" data-level="1">Very large, grows fast</dd>
          <dt>Duration</dt><dd data-dial-src="duration" data-level=".95">Hours</dd>
          <dt>Gap between steps</dt><dd data-dial-src="gap" data-level=".12">Seconds</dd>
          <dt>Latency pressure</dt><dd data-dial-src="latency" data-level=".25">Throughput over speed</dd>
          <dt>Memory elasticity</dt><dd data-dial-src="place" data-level=".15">Server-side, can add capacity</dd>
        </dl>
        <p data-verdict-src="">A coding task carries files, tool results, errors and decisions for hours. Control what enters the conversation and preserve milestones somewhere durable.</p>
        <ul><li>Filter tool output</li><li>Checkpoint progress</li><li>Compact at milestones</li></ul>
      </article>
      <article data-workload-data="device">
        <h4>On-device</h4>
        <dl>
          <dt>Information</dt><dd data-dial-src="context" data-level=".1">Small, and hard-capped by hardware</dd>
          <dt>Duration</dt><dd data-dial-src="duration" data-level=".25">While the app is in front</dd>
          <dt>Gap between steps</dt><dd data-dial-src="gap" data-level=".6">Tens of seconds to minutes</dd>
          <dt>Latency pressure</dt><dd data-dial-src="latency" data-level=".75">Immediate, or it feels broken</dd>
          <dt>Memory elasticity</dt><dd data-dial-src="place" data-level="1">On the device — no larger machine exists</dd>
        </dl>
        <p data-verdict-src="">A phone shares a few gigabytes of unified memory with everything else running and hits a thermal limit before a memory limit. The engineering is in curating what gets sent and deciding when to escalate to a server.</p>
        <ul><li>Hard context ceiling</li><li>Define the offline set</li><li>Design the escalation path</li></ul>
      </article>
      <article data-workload-data="industrial">
        <h4>Industrial edge</h4>
        <dl>
          <dt>Information</dt><dd data-dial-src="context" data-level=".12">Bounded, nothing accumulates</dd>
          <dt>Duration</dt><dd data-dial-src="duration" data-level="1">Runs continuously</dd>
          <dt>Gap between steps</dt><dd data-dial-src="gap" data-level=".02">Milliseconds</dd>
          <dt>Latency pressure</dt><dd data-dial-src="latency" data-level="1">Hard deadline</dd>
          <dt>Memory elasticity</dt><dd data-dial-src="place" data-level="1">On the machine, often without a network</dd>
        </dl>
        <p data-verdict-src="">A robot cell, vehicle or line inspector must finish every decision inside a deadline, on hardware it cannot expand, sometimes with no connectivity. Worst-case latency matters more than average throughput, and a cloud fallback that occasionally takes four seconds is not a fallback.</p>
        <ul><li>Bounded rolling state</li><li>Worst-case deadline</li><li>Degrade safely offline</li></ul>
      </article>
      <article data-workload-data="cctv">
        <h4>Monitoring</h4>
        <dl>
          <dt>Information</dt><dd data-dial-src="context" data-level=".3">Fixed by design</dd>
          <dt>Duration</dt><dd data-dial-src="duration" data-level="1">Runs forever</dd>
          <dt>Gap between steps</dt><dd data-dial-src="gap" data-level=".5">Fixed interval</dd>
          <dt>Latency pressure</dt><dd data-dial-src="latency" data-level=".4">Seconds, and soft</dd>
          <dt>Memory elasticity</dt><dd data-dial-src="place" data-level=".5">Split: edge detectors, server judgement</dd>
        </dl>
        <p data-verdict-src="">Continuous video should not become a continuous conversation. Let inexpensive detectors create event records and call an expensive model only when a judgement is needed.</p>
        <ul><li>Detect first</li><li>Store events</li><li>Escalate selectively</li></ul>
      </article>
      <article data-workload-data="batch">
        <h4>Documents</h4>
        <dl>
          <dt>Information</dt><dd data-dial-src="context" data-level=".8">Large, all at once</dd>
          <dt>Duration</dt><dd data-dial-src="duration" data-level=".08">One call</dd>
          <dt>Gap between steps</dt><dd data-dial-src="gap" data-level="0">No turns at all</dd>
          <dt>Latency pressure</dt><dd data-dial-src="latency" data-level=".03">Nobody waiting</dd>
          <dt>Memory elasticity</dt><dd data-dial-src="place" data-level=".05">Server-side, and time-shiftable</dd>
        </dl>
        <p data-verdict-src="">Independent documents with no person waiting can be queued and processed in batches. There is no reason to buy an interactive experience the product does not need.</p>
        <ul><li>Independent jobs</li><li>Queue the work</li><li>Return structured results</li></ul>
      </article>
      <article data-workload-data="voice">
        <h4>Voice</h4>
        <dl>
          <dt>Information</dt><dd data-dial-src="context" data-level=".18">Small</dd>
          <dt>Duration</dt><dd data-dial-src="duration" data-level=".3">Minutes</dd>
          <dt>Gap between steps</dt><dd data-dial-src="gap" data-level=".06">Under a second</dd>
          <dt>Latency pressure</dt><dd data-dial-src="latency" data-level=".95">Hard, conversational</dd>
          <dt>Memory elasticity</dt><dd data-dial-src="place" data-level=".4">Often split between device and server</dd>
        </dl>
        <p data-verdict-src="">Voice looks like chat but is governed by silence. The full path from speech detection through retrieval and speech generation must fit inside one conversational latency budget.</p>
        <ul><li>Budget every stage</li><li>Start speaking early</li><li>Optimise the full chain</li></ul>
      </article>
    </div>
  </div>

  <div class="sys-arithmetic" id="arithmetic">
    <div class="sys-arithmetic-head">
      <p class="think-kicker">The arithmetic</p>
      <h3>One million tokens, three different capacity plans.</h3>
      <p>The figures below are for a specific, checkable setup: <strong>Llama 3.1 8B</strong> — 32 layers, eight key/value heads, 128 dimensions per head — running at two bytes per number on <strong>one NVIDIA H100</strong>, the 80 GB data-centre GPU that most current inference runs on. Your model and card will give different numbers. The ratios are the point, and the method is the part worth copying.</p>
    </div>

    <div class="sys-arith-constants">
      <div><span>Working state per token</span><strong>128 KB</strong><em>2 × 32 layers × 8 KV heads × 128 dims × 2 bytes</em></div>
      <div><span>Memory left after weights</span><strong>64 GB</strong><em>H100 has 80 GB; the 8B model’s weights take about 16</em></div>
      <div><span>Tokens held at once</span><strong>~524,000</strong><em>64 GB ÷ 128 KB. This, not the context window, is the budget</em></div>
    </div>

    <div class="sys-ladder-block" id="derivation">
      <div class="sys-ladder-head">
        <p class="think-kicker">Where 128 KB comes from</p>
        <h3>One token to one invoice, in seven multiplications.</h3>
        <p>Every step below is one multiplication and the reason it is there. Nothing is rounded until the last line, so you can redo it with your own model and your own rate card. If <em>head</em>, <em>key/value head</em> and <em>layer</em> are not yet solid, the companion explains <a href="/2026/08/31/how-an-ai-model-turns-input-into-output.html#attention-runs-many-times-in-parallel-and-that-is-where-heads-come-from">where they come from</a> — the three of them supply three of the four factors here.</p>
      </div>

      <ol class="sys-ladder">
        <li><span>01</span><div><strong>One key vector, one attention head</strong><p>Attention runs many times in parallel; each parallel copy is a <a href="/2026/08/31/how-an-ai-model-turns-input-into-output.html#attention-runs-many-times-in-parallel-and-that-is-where-heads-come-from">head</a>, and each works on a slice of 128 numbers. At two bytes per number — the usual 16-bit precision — that slice is a fixed cost per token, per head.</p></div><b>128 × 2 B<i>256 B</i></b></li>
        <li><span>02</span><div><strong>Keys <em>and</em> values</strong><p>This is the factor people miss. Attention keeps two vectors per position, not one: the key it matches against and the value it carries forward. Hence the leading 2 in the formula.</p></div><b>256 B × 2<i>512 B</i></b></li>
        <li><span>03</span><div><strong>Every key/value head in the layer</strong><p>Llama 3.1 8B has 32 query heads but only eight key/value heads. That gap is deliberate: queries are used once and thrown away, keys and values are kept forever, so <a href="/2026/08/31/how-an-ai-model-turns-input-into-output.html#attention-runs-many-times-in-parallel-and-that-is-where-heads-come-from">sharing them across query heads</a> cuts this bill fourfold. This is the number that multiplies.</p></div><b>512 B × 8<i>4 KB</i></b></li>
        <li><span>04</span><div><strong>Every layer</strong><p>Each of the 32 layers runs its own attention and keeps its own keys and values; nothing is shared between them. This is the line that produces the number the table above uses.</p></div><b>4 KB × 32<i>128 KB per token</i></b></li>
        <li><span>05</span><div><strong>A whole coding session</strong><p>A long agent run carrying 180,000 tokens of files, tool output and history. The relationship is linear: double the context, double the memory.</p></div><b>180,000 × 128 KB<i>22 GB</i></b></li>
        <li><span>06</span><div><strong>What one H100 has left</strong><p>80 GB of high-bandwidth memory, about 16 GB of it holding the model’s weights. The remainder is what every concurrent session competes for.</p></div><b>80 − 16 GB<i>64 GB for context</i></b></li>
        <li><span>07</span><div><strong>How many fit</strong><p>Two coding sessions. Or, at 4,000 tokens a turn, about 130 chat conversations held ready at once on the same card.</p></div><b>64 GB ÷ 22 GB<i>2 sessions</i></b></li>
      </ol>

      <div class="sys-cost-paths">
        <div class="sys-cost-paths-head">
          <p class="think-kicker">And then the invoice</p>
          <h3>There are two ways to pay for that memory.</h3>
          <p>You either rent the hardware and do this arithmetic yourself, or you rent the tokens and let someone else do it. The price list in the second case <em>is</em> the arithmetic in the first, with a margin on top.</p>
        </div>

        <article>
          <p class="think-thinker">Path A · Rent the GPU</p>
          <h4>The memory budget is the bill.</h4>
          <p>An H100 runs roughly <a href="https://intuitionlabs.ai/articles/h100-rental-prices-cloud-comparison">two to four dollars an hour</a>. Take three. That hour costs the same whatever you run on it, so the cost per user is set entirely by how many users fit.</p>
          <table class="sys-cost-table">
            <thead><tr><th>Workload</th><th>Fits</th><th>Cost per session-hour</th></tr></thead>
            <tbody>
              <tr><td>Chat, 4,000 tokens</td><td>~130</td><td><b>2.3¢</b></td></tr>
              <tr><td>Coding agent, 180,000 tokens</td><td>2</td><td><b>$1.50</b></td></tr>
            </tbody>
          </table>
          <p><strong>Same GPU, same hour, a 65× difference per user.</strong> Not because one workload is harder, but because context length decides how many people share the card. This is the number a pricing model has to survive.</p>
        </article>

        <article>
          <p class="think-thinker">Path B · Rent the tokens</p>
          <h4>The quadratic shows up as a line item.</h4>
          <p>Take the 50-step agent from above: each step adds about 2,000 tokens and resends everything before it, so the model processes 2,550,000 input tokens to produce a 100,000-token transcript.</p>
          <table class="sys-cost-table">
            <thead><tr><th>Per agent run</th><th>Input processed</th><th>Cost</th></tr></thead>
            <tbody>
              <tr><td>No prefix reuse</td><td>2,550,000</td><td><b>$12.75</b></td></tr>
              <tr><td>With prefix reuse</td><td>100,000 fresh + 2,450,000 cached</td><td><b>$1.85</b></td></tr>
            </tbody>
          </table>
          <p>At <span class="sys-rate">$5 per million input tokens</span>, cache reads billed near a tenth of that and writes at a small premium. <strong>Roughly 7× on the same completed task</strong>, from one decision about the order of the request.</p>
        </article>
      </div>

      <p class="sys-ladder-note">Figures are illustrative and the rate card moves; the shape does not. Two things are worth carrying away: memory scales linearly with context, so a context limit is a concurrency decision wearing different clothes. And an agent’s bill scales with the <em>square</em> of its steps unless the prefix is reused, which is why that one setting is worth more than most prompt engineering.</p>
    </div>

    <div class="sys-arith-table" role="table" aria-label="Four workloads compared by working state and concurrency on one H100">
      <div class="sys-arith-head" role="row"><span role="columnheader">Workload</span><span role="columnheader">Working context</span><span role="columnheader">State held</span><span role="columnheader">Held for</span><span role="columnheader">Fits on one H100</span></div>
      <div role="row"><strong data-label="Workload" role="cell">Chat turn</strong><span data-label="Working context" role="cell">4,000 tokens</span><span data-label="State held" role="cell">0.5 GB</span><span data-label="Held for" role="cell">Seconds, then idle</span><b data-label="Fits on one H100" role="cell">~80 conversations</b></div>
      <div role="row"><strong data-label="Workload" role="cell">Document extraction</strong><span data-label="Working context" role="cell">50,000 in, 200 out</span><span data-label="State held" role="cell">6.1 GB</span><span data-label="Held for" role="cell">~30 seconds, then released</span><b data-label="Fits on one H100" role="cell">6 jobs in parallel</b></div>
      <div role="row"><strong data-label="Workload" role="cell">Coding agent</strong><span data-label="Working context" role="cell">180,000 tokens, growing</span><span data-label="State held" role="cell">22 GB</span><span data-label="Held for" role="cell">Hours, continuously</span><b data-label="Fits on one H100" role="cell">1, with nothing to spare</b></div>
      <div role="row"><strong data-label="Workload" role="cell">Phone assistant</strong><span data-label="Working context" role="cell">8,000 tokens</span><span data-label="State held" role="cell">1 GB</span><span data-label="Held for" role="cell">While the app is foregrounded</span><b data-label="Fits on one H100" role="cell">Not the question — see below</b></div>
    </div>

    <p class="sys-arith-punchline">Every row above can be described as “about a million tokens a day.” One is about a hundred and thirty users, one is a queue that drains overnight, one is two engineers, and one never reaches your H100 at all. Token count told you the volume. It did not tell you how many machines to buy, and it is the number most roadmaps are still written in.</p>

    <div class="sys-arith-split">
      <article>
        <p class="think-thinker">On the device</p>
        <h3>The budget stops being elastic.</h3>
        <p>A phone running a small model locally might have two to three gigabytes of unified memory available to it, shared with the operating system and every other app, and it is thermally limited before it is memory limited. At the 128 KB per token calculated above, working state alone would consume the entire budget within a few thousand tokens — which is why on-device models use aggressive quantisation, far smaller key/value dimensions and hard context ceilings rather than the generous windows their server counterparts advertise.</p>
        <p>The design consequence is not “use a smaller model.” It is that context has to be curated before it is sent, that the device cannot hold a long conversation in working state between sessions, and that anything requiring a large context is a routing decision: answer here, or escalate to a server and accept the round trip, the privacy question and the offline failure mode.</p>
        <aside class="think-example"><span>The decision</span>Define which requests must be answerable with no network, and treat that set as a hard requirement. Everything else can escalate. A product that degrades unpredictably in a tunnel or a basement was designed without this line.</aside>
      </article>
      <article>
        <p class="think-thinker">In the loop</p>
        <h3>An agent’s cost grows with the square of its steps.</h3>
        <p>An agent step is not an incremental request. It is a new request whose input is the entire conversation so far. If each step adds about 2,000 tokens of tool result and reasoning, then step 50 sends 100,000 tokens of input to produce a few hundred.</p>
        <p>Across the whole run, the model processes the sum of every intermediate length:</p>
        <p class="sys-arith-formula">2,000 × (1 + 2 + … + 50) = <strong>2,550,000 tokens</strong> processed<br />for a transcript that ends at only 100,000</p>
        <p>Twenty-five times the transcript. That multiple is what prefix caching exists to remove: when the shared beginning is reused, a step reprocesses only what is new, and the run costs closer to the transcript length than to its square. This is why cache-hit rate is not an infrastructure statistic on an agentic product. It is the difference between a feature that ships and one that gets cancelled at the second invoice.</p>
        <aside class="think-example"><span>Failure mode</span>An agent that appends a timestamp, a step counter or a re-fetched policy near the top of each step is invalidating its own prefix every turn, and paying the full quadratic. It will look correct in every test and ruinous in production.</aside>
      </article>
    </div>
  </div>

  <div class="sys-opposed-cases">
    <article>
      <p class="think-thinker">Agent fleet, in the datacentre</p>
      <h3>Memory is elastic. Attention is not.</h3>
      <p>An autonomous agent holds a large, growing context for hours with almost no idle gaps, because it is never waiting for a human to type. It is the most expensive thing to hold in memory and the least likely to be evicted safely. You can always add machines; what you cannot do is make a 200,000-token context produce a good decision when 180,000 of those tokens are stale tool output.</p>
      <strong>Primary constraint: context grows faster than its usefulness. Budget for curation, isolation and compaction as first-class work.</strong>
    </article>
    <div aria-hidden="true">↔</div>
    <article>
      <p class="think-thinker">Assistant, on the device</p>
      <h3>Attention is manageable. Memory is fixed.</h3>
      <p>A phone or a handheld industrial terminal has a memory ceiling set by hardware you do not control, shared with everything else running, and a thermal limit that arrives before the memory limit does. There is no larger instance. The context is small enough to reason about entirely, and the engineering goes into deciding what deserves to be in it and when to give up and call a server.</p>
      <strong>Primary constraint: a hard ceiling. Design the escalation path and the offline behaviour before the prompt.</strong>
    </article>
  </div>

</section>

<section class="think-act think-act--config" id="config" data-act="config" aria-labelledby="config-title">
  <div class="think-act-head">
    <p class="think-act-number">Act 04 <span>·</span> Configure it</p>
    <h2 id="config-title">Now configure it.</h2>
    <p>Everything above is a decision. Each one has a parameter attached to it, and most teams never touch them because nobody told them the parameters existed.</p>
  </div>

  <div class="sys-config-intro">
    <p>A product manager’s exposure to a model is usually one endpoint, a prompt and a model name. That is a fair description of a single call and a poor description of an agent: the controls that decide whether an autonomous agent is affordable are request parameters, not prompt wording. What follows builds one, in the Anthropic Messages API because its parameter names are public and stable enough to print — the same controls exist elsewhere under other names.</p>
  </div>

  <div class="sys-usecase">
    <p class="think-kicker">The use case</p>
    <h3>Overnight refund resolution.</h3>
    <p><a href="#mechanics">Order 4821</a>, the late delivery this guide opened with, was one of them. So were nine hundred others this week. Each claim needs the carrier record checked, the delayed-delivery policy applied, the customer’s history weighed, and a decision written: refund, partial credit, or escalate to a human. Nobody is waiting at 2am. The agent runs unattended until the queue is empty.</p>
    <div class="sys-usecase-shape">
      <div><span>Information</span><strong>Grows per claim, discarded between claims</strong></div>
      <div><span>Duration</span><strong>Hours, unattended</strong></div>
      <div><span>Gaps</span><strong>None — it never waits for a human</strong></div>
      <div><span>Latency</span><strong>Irrelevant. Cost per resolved claim is the metric</strong></div>
      <div><span>Where</span><strong>Server-side</strong></div>
    </div>
  </div>

  <div class="sys-param-map">
    <div class="sys-param-map-head">
      <p class="think-kicker">The mapping</p>
      <h3>Every decision in this guide is a parameter.</h3>
    </div>
    <div class="sys-param-table" role="table" aria-label="Design decisions mapped to API parameters">
      <div class="sys-param-head" role="row"><span role="columnheader">The decision</span><span role="columnheader">The parameter</span><span role="columnheader">What it does</span></div>
      <div role="row"><strong data-label="The decision" role="cell">Put stable content first<em><a href="#approaches">Order by rate of change</a></em></strong><code data-label="The parameter" role="cell">cache_control</code><span data-label="What it does" role="cell">Marks the end of the stable prefix so it is reused instead of reprocessed each step.</span></div>
      <div role="row"><strong data-label="The decision" role="cell">Fetch facts when needed<em><a href="#approaches">Retrieve at decision time</a></em></strong><code data-label="The parameter" role="cell">tools</code><span data-label="What it does" role="cell">The agent asks for the carrier record when it needs it, rather than carrying every order.</span></div>
      <div role="row"><strong data-label="The decision" role="cell">Keep tool output out of the transcript<em><a href="#approaches">Store it, reference it</a></em></strong><code data-label="The parameter" role="cell">context_management</code><span data-label="What it does" role="cell">Clears old tool results from the context automatically, so claim 400 is not still carrying claim 1.</span></div>
      <div role="row"><strong data-label="The decision" role="cell">How long the work continues<em><a href="#mechanics">One of the five questions</a></em></strong><code data-label="The parameter" role="cell">task_budget</code><span data-label="What it does" role="cell">Tells the agent its token ceiling so it paces itself and finishes, instead of being cut off mid-claim.</span></div>
      <div role="row"><strong data-label="The decision" role="cell">How hard to work each claim<em>Reasoning is billed, so it is bought</em></strong><code data-label="The parameter" role="cell">effort</code><span data-label="What it does" role="cell">Before answering, the model can generate <a href="/2026/08/31/how-an-ai-model-turns-input-into-output.html#some-of-the-generated-tokens-are-not-the-answer">reasoning tokens</a> — working out that the reader never sees but pays for as output. This sets how much of it to buy.</span></div>
      <div role="row"><strong data-label="The decision" role="cell">Isolate exploration<em><a href="#approaches">Work in sub-agents</a></em></strong><code data-label="The parameter" role="cell">model</code><span data-label="What it does" role="cell">A cheaper model reads the long carrier logs; the expensive one only sees the conclusion.</span></div>
    </div>
  </div>

  <div class="sys-config-block">
    <p class="think-kicker">The configuration</p>
    <h3>The same agent, written down.</h3>
    <pre class="sys-config-code"><code>from anthropic import Anthropic

client = Anthropic()

with client.beta.messages.stream(
    model="claude-opus-5",
    max_tokens=64000,

    # Stable prefix. Refund rules, tone, escalation policy — the things that
    # do not change between claim 1 and claim 900. The cache breakpoint goes
    # at the END of this block, so everything above it is reused every step.
    system=[{
        "type": "text",
        "text": REFUND_POLICY_AND_INSTRUCTIONS,
        "cache_control": {"type": "ephemeral"},
    }],

    # How long the work continues, and how hard to think about each claim.
    # "effort" buys reasoning tokens: output the reader never sees, billed
    # and waited on like any other output. task_budget is advisory - the
    # agent sees the countdown and wraps up gracefully. max_tokens is a
    # hard cut, mid-sentence, with no warning.
    output_config={
        "effort": "medium",
        "task_budget": {"type": "tokens", "total": 400000},
    },
    thinking={"type": "adaptive"},

    # Keeping tool output out of the transcript, enforced by the server.
    # Old tool results are cleared as the run proceeds, so a 900-claim night
    # does not end with claim 1's carrier dump still resent on every request.
    context_management={"edits": [{"type": "clear_tool_uses_20250919"}]},

    # Fetching facts when needed. Nothing about order 4821 is preloaded —
    # the agent asks for the authoritative record when it decides it needs it.
    tools=[get_carrier_record, get_customer_history, issue_refund, escalate],

    betas=["task-budgets-2026-03-13", "context-management-2025-06-27"],
    messages=[{"role": "user", "content": "Work the delayed-delivery queue."}],
) as stream:
    result = stream.get_final_message()

# The number that decides whether this is affordable.
print(result.usage.cache_read_input_tokens)</code></pre>
    <p class="sys-config-note">Two lines deserve a note. <code>thinking</code> lets the model generate <a href="/2026/08/31/how-an-ai-model-turns-input-into-output.html#some-of-the-generated-tokens-are-not-the-answer">reasoning tokens</a> before it answers — real output tokens, billed and waited on, that the reader never sees; <code>effort</code> is how much of that to buy, and it is the difference between a careful refund decision and an expensive one on a claim that needed no thought. And streaming is not stylistic: a run with <code>max_tokens</code> this large will hit an HTTP timeout without it.</p>
  </div>

  <div class="sys-config-notes">
    <article>
      <p class="think-thinker">The line that matters most</p>
      <h3>Watch <code>cache_read_input_tokens</code>, not the token total.</h3>
      <p>This is <a href="#arithmetic">the quadratic cost calculated earlier</a>, made observable. If that number is zero across repeated steps, the agent is reprocessing its entire history every time and you are paying the full square. It is the single most valuable number in the response object, and almost nobody reads it.</p>
      <p>The usual cause is something small and volatile sitting above the cache breakpoint — a timestamp, a claim ID, a re-fetched policy paragraph, a greeting with the customer’s name. Any byte change in the prefix invalidates everything after it. The fix is placement: stable content first, volatile content after the breakpoint.</p>
      <aside class="think-example"><span>The 2am check</span>Log <code>cache_read_input_tokens</code> for every step of one real overnight run. If it climbs and holds, the design is working. If it sits at zero, stop tuning prompts and go find your first differing byte.</aside>
    </article>
    <article>
      <p class="think-thinker">The escape hatch</p>
      <h3>Give the reading job to a cheaper model.</h3>
      <p>Carrier logs are long, mostly irrelevant and read once. Putting them through the expensive model is the most common avoidable cost in an agent like this. Run that sub-task separately on a smaller model and return only its conclusion to the main loop — the exploration never enters the expensive context, and never gets carried for the rest of the night.</p>
      <pre class="sys-config-code sys-config-code--small"><code>summary = client.messages.create(
    model="claude-haiku-4-5",
    max_tokens=1024,
    messages=[{"role": "user", "content": f"{CARRIER_LOG}\n\nDid it arrive late, and by how long?"}],
)
# Only `summary` goes back to the Opus loop. The log never does.</code></pre>
      <p>One caution: caches are scoped to a model. Switching models <em>inside</em> a single conversation invalidates the cached prefix, which is why this runs as a separate call rather than a mid-conversation swap.</p>
    </article>
  </div>

  <div class="sys-providers" id="providers">
    <div class="sys-providers-head">
      <p class="think-kicker">The same controls elsewhere</p>
      <h3>Every serving stack has these knobs. Only the spelling changes.</h3>
      <p>The names above are Anthropic’s. The reason to learn the underlying idea rather than the parameter is that the idea ports and the parameter does not. Here are the same four controls in three other places a team might actually be working.</p>
    </div>

    <div class="sys-provider-table" role="table" aria-label="Equivalent controls across inference providers">
      <div class="sys-provider-head" role="row"><span role="columnheader">The control</span><span role="columnheader">Anthropic</span><span role="columnheader">Google Gemini</span><span role="columnheader">vLLM, self-hosted</span></div>
      <div role="row">
        <strong data-label="The control" role="cell">Reuse the stable prefix</strong>
        <span data-label="Anthropic" role="cell"><code>cache_control</code> on the last stable block</span>
        <span data-label="Google Gemini" role="cell"><a href="https://ai.google.dev/gemini-api/docs/caching">Implicit caching</a>, on by default above 2,048 tokens; explicit caches for guaranteed reuse</span>
        <span data-label="vLLM" role="cell"><a href="https://docs.vllm.ai/en/stable/design/prefix_caching/"><code>--enable-prefix-caching</code></a>, hashing the KV cache in fixed blocks</span>
      </div>
      <div role="row">
        <strong data-label="The control" role="cell">Confirm it actually happened</strong>
        <span data-label="Anthropic" role="cell"><code>usage.cache_read_input_tokens</code></span>
        <span data-label="Google Gemini" role="cell"><code>cachedContentTokenCount</code> in the response metadata</span>
        <span data-label="vLLM" role="cell">Prefix cache hit rate on the metrics endpoint</span>
      </div>
      <div role="row">
        <strong data-label="The control" role="cell">Buy more reasoning before the answer</strong>
        <span data-label="Anthropic" role="cell"><code>effort</code>, inside <code>output_config</code></span>
        <span data-label="Google Gemini" role="cell">A thinking budget on the request</span>
        <span data-label="vLLM" role="cell">A property of the model you chose, not a dial the server gives you</span>
      </div>
      <div role="row">
        <strong data-label="The control" role="cell">Set the memory budget</strong>
        <span data-label="Anthropic" role="cell">Not yours to set — it is priced in</span>
        <span data-label="Google Gemini" role="cell">Not yours to set — it is priced in</span>
        <span data-label="vLLM" role="cell"><a href="https://docs.vllm.ai/en/v0.6.5/usage/engine_args.html"><code>--gpu-memory-utilization</code></a> and <code>--max-model-len</code>: the arithmetic above, as a config flag</span>
      </div>
    </div>

    <div class="sys-provider-notes">
      <article>
        <p class="think-thinker">The discount is real</p>
        <h3>Prefix reuse is the largest single lever, on every platform.</h3>
        <p>Google documents a <a href="https://ai.google.dev/gemini-api/docs/caching">90% discount on cached input tokens</a> for its 2.5-generation models and later. Anthropic prices cache reads far below fresh input. vLLM’s saving is not a discount at all — it is compute you simply never perform.</p>
        <p>Three different commercial models, one shared conclusion: the stable-prefix decision is worth more than almost any prompt change you could make, and it costs nothing but paying attention to the order of your request.</p>
      </article>
      <article>
        <p class="think-thinker">If you rent the GPU yourself</p>
        <h3>The arithmetic stops being theoretical.</h3>
        <p>An H100 rents for roughly <a href="https://intuitionlabs.ai/articles/h100-rental-prices-cloud-comparison">two to four dollars an hour</a> depending on provider and commitment. At that point the calculation earlier is not an analogy for your cost — it <em>is</em> your cost. The tokens you choose to hold in memory are the difference between two concurrent sessions and a hundred and twenty-eight on hardware you are paying for by the hour either way.</p>
        <p>This is also where the abstraction stops protecting you. A managed API absorbs a bad context design into a slightly larger bill. A GPU you rented absorbs it by falling over.</p>
      </article>
    </div>
  </div>

  <p class="sys-config-caveat">Parameter names, beta identifiers and model availability move faster than any article. Treat the shapes above as the current form of durable ideas — a stable prefix, a bounded task, a filtered context and a cheap reader — and check the provider’s reference before shipping. The ideas outlive the spellings.</p>
</section>

<section class="think-act think-act--levers" id="levers" data-act="levers" aria-labelledby="levers-title">
  <div class="think-act-head">
    <p class="think-act-number">Act 05 <span>·</span> Measure it</p>
    <h2 id="levers-title">Measure the experience, then find the system cause.</h2>
    <p>Infrastructure metrics become useful only when they explain something a user or product team can observe.</p>
  </div>

  <div class="sys-diagnostic-table" role="table" aria-label="User symptoms, evidence and product actions">
    <div class="sys-diagnostic-head" role="row"><span role="columnheader">What happens</span><span role="columnheader">Inspect</span><span role="columnheader">Likely action</span></div>
    <div role="row"><strong role="cell">The response always starts slowly</strong><span role="cell">Context length, retrieval time and time to first token</span><span role="cell">Remove irrelevant input and shorten the retrieval path.</span></div>
    <div role="row"><strong role="cell">Returning after a pause is slow</strong><span role="cell">Gap between turns and cache reuse</span><span role="cell">Match retention to real return behaviour; do not optimise for an average pause.</span></div>
    <div role="row"><strong role="cell">The answer freezes halfway</strong><span role="cell">Time between output tokens and serving interruptions</span><span role="cell">Add capacity headroom or change scheduling for latency-sensitive work.</span></div>
    <div role="row"><strong role="cell">Long sessions become costly or confused</strong><span role="cell">Context p50, p95 and p99; compactions per session</span><span role="cell">Find the largest inputs, filter tool output and improve handoffs.</span></div>
    <div role="row"><strong role="cell">Shared instructions are repeatedly processed</strong><span role="cell">Cache-hit rate and the first changing prompt block</span><span role="cell">Place stable material before timestamps, user fields and request-specific data.</span></div>
    <div role="row"><strong role="cell">An agent’s cost scales worse than its usefulness</strong><span role="cell">Tokens processed per completed task, and cache-hit rate per step</span><span role="cell">Find what invalidates the prefix each step; move tool output out of the transcript.</span></div>
    <div role="row"><strong role="cell">The agent loops without converging</strong><span role="cell">Steps per task, repeated tool calls, context at the point of failure</span><span role="cell">Cap the loop, isolate exploration in sub-agents and compact to a schema at each milestone.</span></div>
    <div role="row"><strong role="cell">The on-device path silently stops being used</strong><span role="cell">Share of requests escalated to a server, and why each escalated</span><span role="cell">Check whether the local context ceiling or a thermal limit is causing it, not the model’s ability.</span></div>
  </div>

  <p class="sys-coda">None of the rows above start with a model. They start with something a user noticed, and end with a decision about where information lives. That is the whole argument: the model is one component in a system you already know how to reason about, and the questions worth asking — what does this decision need, how long must we hold it, where does it run, what does that cost — are questions you were qualified to answer before any of this arrived.</p>

</section>]]></content><author><name>Yash Tambawala</name></author><summary type="html"><![CDATA[How workload shape—not just token count—determines the cost, speed and architecture of an AI product]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://yashtambawala.com/assets/og/ai-system-design-for-product-managers.png" /><media:content medium="image" url="https://yashtambawala.com/assets/og/ai-system-design-for-product-managers.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">How to Think</title><link href="https://yashtambawala.com/2026/08/26/how-to-think.html" rel="alternate" type="text/html" title="How to Think" /><published>2026-08-26T09:00:00+05:30</published><updated>2026-08-26T09:00:00+05:30</updated><id>https://yashtambawala.com/2026/08/26/how-to-think</id><content type="html" xml:base="https://yashtambawala.com/2026/08/26/how-to-think.html"><![CDATA[<section class="think-orientation" aria-labelledby="orientation-title">
  <p class="think-kicker">Before you begin</p>
  <h2 id="orientation-title">There is no single best framework.</h2>
  <p>I use these thinkers as different ways into a problem. They often disagree, which is useful: each one helps me notice something the others miss.</p>
  <div class="think-lens-index" aria-label="The field guide at a glance">
    <div><strong>Deleuze</strong><span>What generated this?</span></div>
    <div><strong>Korzybski</strong><span>What did the map omit?</span></div>
    <div><strong>Meadows</strong><span>What system produced this?</span></div>
    <div><strong>Bateson</strong><span>What pattern connects it?</span></div>
    <div><strong>Kuhn</strong><span>What does the frame hide?</span></div>
    <div><strong>Simon</strong><span>What can I know before acting?</span></div>
    <div><strong>Gigerenzer</strong><span>What simple rule is enough?</span></div>
    <div><strong>Bayes / Tetlock</strong><span>What moves my probability?</span></div>
  </div>
</section>

<section class="think-act think-act--model" id="model" data-act="model" aria-labelledby="model-title">
  <div class="think-act-head">
    <p class="think-act-number">Act 01</p>
    <h2 id="model-title">A label is not an explanation.</h2>
    <p>A label tells you what something resembles. An explanation tells you how it works and how it got there.</p>
  </div>

  <div class="think-lens-row">
    <article class="think-lens">
      <p class="think-thinker">Gilles Deleuze</p>
      <h3>How did this become what it is?</h3>
      <p>Start with the thing you can see, then ask what made it. Calling a retailer “high quality” is not an explanation. Ask what concrete conditions—customer loyalty, buying power, store density, low costs, and capable operators—produced the result.</p>
      <p class="think-formula">thing → category<br /><strong>difference → relation → process → emergence</strong></p>

      <h4>Representation</h4>
      <p>Labels such as “emerging market,” “monopoly,” “luxury,” or “high ROCE” save time. But naming something can create a false sense that we understand it.</p>
      <aside class="think-example"><span>Example</span>A retailer looks like “the Costco of Country X.” That analogy is a starting map, not an explanation. Ask where it breaks: membership economics, supplier power, land costs, customer density, or culture may be completely different.</aside>

      <h4>Same result, different causes</h4>
      <p>When two companies look similar, ask how each one got there.</p>
      <aside class="think-example"><span>Example</span>Two firms both report 25% margins. One charges a premium; the other keeps costs low and its factories busy. The number is the same. The businesses are not.</aside>

      <h4>One number hides many realities</h4>
      <p>A company or country is made up of people, incentives, contracts, technology, institutions, geography, and history. “Country X grew 7%” can be true even while agriculture is weak, youth unemployment is high, and some regions are falling behind.</p>

      <h4>Thresholds matter</h4>
      <p>A factory at 60% utilisation and the same factory at 90% can have very different economics. Look for bottlenecks and tipping points, not just gradual change.</p>

      <h4>Problems before solutions</h4>
      <p>“How do we increase app engagement?” invites notifications and gamification. “Why is there no recurring reason to open the app?” may reveal that the product solves an occasional problem and should not be optimised for daily use.</p>

      <blockquote class="think-question">What produced this result—and what does the label hide?</blockquote>
    </article>

    <aside class="think-diagram think-diagram--genesis" aria-label="Six generative forces combining to produce a visible outcome">
      <span class="diagram-label">Visible outcome</span>
      <div class="genesis-output">A RETAILER WITH<br />LOYAL CUSTOMERS</div>
      <div class="genesis-arrow" aria-hidden="true">↑</div>
      <span class="diagram-label diagram-label--forces">Generative forces</span>
      <div class="genesis-forces">
        <span>low prices</span><span>member renewals</span><span>dense stores</span><span>supplier terms</span><span>low staff turnover</span><span>enough sales volume</span>
      </div>
      <p>A category names the outcome.<br /><strong>Explanation traces the forces that made it.</strong></p>
    </aside>
  </div>

  <div class="think-lens-row think-lens-row--reverse">
    <article class="think-lens">
      <p class="think-thinker">Alfred Korzybski</p>
      <h3>The map is not the territory</h3>
      <p>Every useful map leaves things out. A road map ignores trees and building interiors; a DCF leaves out much of the day-to-day business. Simplifying is necessary. Forgetting what you removed is the danger.</p>
      <dl class="think-definitions">
        <div><dt>Territory</dt><dd>What is actually happening.</dd></div>
        <div><dt>Map</dt><dd>How you represented it.</dd></div>
        <div><dt>Omission</dt><dd>What the representation cannot contain.</dd></div>
      </dl>
      <aside class="think-example"><span>Example</span>A P/E multiple may be useful shorthand for valuation while omitting balance-sheet fragility, reinvestment needs, accounting quality, and the durability of current earnings.</aside>
      <aside class="think-example"><span>Example</span>An org chart maps formal authority. The real flow of information may run through friendships, old colleagues, project teams, and informal experts.</aside>
      <blockquote class="think-question">What did this representation have to omit in order to become useful?</blockquote>
    </article>
    <aside class="think-diagram think-diagram--omission" aria-label="A valuation multiple is a visible map while debt, cycle, reinvestment, and earnings quality remain outside it">
      <span class="diagram-label">Territory: the whole business</span>
      <div class="omission-territory">
        <i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i>
        <span class="omission-item omission-item--debt">Debt</span><span class="omission-item omission-item--cycle">Cycle</span><span class="omission-item omission-item--quality">Earnings quality</span><span class="omission-item omission-item--reinvestment">Reinvestment</span>
        <div class="omission-map"><span>Map: one visible metric</span><strong>12×</strong><em>P/E multiple</em></div>
      </div>
      <p>The map is useful because it compresses.<br /><strong>Its omissions are the next questions.</strong></p>
    </aside>
  </div>
</section>

<section class="think-act think-act--system" id="system" data-act="system" aria-labelledby="system-title">
  <div class="think-act-head">
    <p class="think-act-number">Act 02</p>
    <h2 id="system-title">Look past the event. Find the system.</h2>
    <p>When a problem recurs, the structure is often a better explanation than the latest incident.</p>
  </div>

  <div class="think-lens-row">
    <article class="think-lens">
      <p class="think-thinker">Donella Meadows</p>
      <h3>Structure creates behaviour</h3>
      <p>Look for what accumulates, what flows in and out, what feeds back, and where delays sit. A system can turn sensible individual choices into a bad collective result.</p>

      <h4>Reinforcing loops</h4>
      <p>A change produces more change in the same direction: more users → more content → a more useful platform → more users.</p>

      <h4>Balancing loops</h4>
      <p>A change activates forces that push back: high prices → lower demand → inventory builds → discounting → prices fall. Extrapolating the first leg misses the system’s response.</p>

      <h4>Delays</h4>
      <p>Cut marketing today; revenue stays strong for three months; management concludes marketing was wasteful; six months later the pipeline collapses. Delayed feedback makes the original decision look smarter than it was.</p>

      <h4>Stocks and flows</h4>
      <p>Customer base is a stock. Acquisition is an inflow; churn is an outflow. Celebrating record acquisition while churn quietly rises is like praising a stronger faucet while the bathtub drain widens.</p>

      <h4>Leverage points</h4>
      <p>Small parameter changes often do little. Information, incentives, rules, and goals usually matter more. If service agents rush customers because promotions depend on calls per hour, another training manual will not fix it. Change the metric.</p>
      <blockquote class="think-question">What system would make this behaviour predictable rather than surprising?</blockquote>
    </article>

    <aside class="think-diagram think-diagram--loop" aria-label="A reinforcing platform feedback loop">
      <div class="loop-node loop-node--one">More users</div>
      <div class="loop-node loop-node--two">More content</div>
      <div class="loop-node loop-node--three">More useful</div>
      <svg viewBox="0 0 400 400" aria-hidden="true">
        <defs><marker id="loop-arrow" markerWidth="7" markerHeight="7" refX="6" refY="3" orient="auto"><path d="M0,0 L0,6 L7,3 z" /></marker></defs>
        <path d="M236 58 C292 78 326 128 322 186" />
        <path d="M304 256 C263 316 139 316 96 256" />
        <path d="M78 198 C70 132 108 80 164 58" />
      </svg>
      <div class="loop-core" aria-label="Reinforcing feedback loop">R<span>reinforcing<br />loop</span></div>
      <p><strong>R = reinforcing feedback:</strong> each turn makes the next turn stronger.</p>
    </aside>
  </div>

  <div class="think-lens-row think-lens-row--reverse">
    <article class="think-lens">
      <p class="think-thinker">Gregory Bateson</p>
      <h3>Look at the relationship, not just the thing</h3>
      <p>We tend to blame isolated things: the manager, the customer, the price. Bateson’s useful point is that behaviour often comes from the relationship between them.</p>

      <h4>A difference that makes a difference</h4>
      <p>A ₹100 cup of coffee is not “high” or “low” by itself. ₹100 relative to its ₹50 cost, a nearby café’s ₹120 price, a customer’s ₹140 willingness-to-pay, or yesterday’s ₹90 price are four different pieces of information.</p>

      <h4>Patterns over traits</h4>
      <p>A “bad manager” may improve dramatically under a different boss, incentive system, and team. The old environment helped produce the old behaviour.</p>

      <h4>Double binds</h4>
      <p>Leadership says “take initiative,” but punishes decisions not pre-approved. Employees learn that the safest form of initiative is to ask permission for everything.</p>
      <blockquote class="think-question">What pattern of relationships is producing the behaviour I am attributing to an isolated object?</blockquote>
    </article>
    <aside class="think-diagram think-diagram--relation" aria-label="The meaning of a one hundred rupee price depends on the comparison">
      <span class="diagram-label">A café charges ₹100 for one coffee</span>
      <div class="relation-price">₹100<span>coffee</span></div>
      <div class="relation-comparisons">
        <span style="--angle:-64deg">₹50 to make it</span>
        <span style="--angle:-18deg">₹120 nearby café</span>
        <span style="--angle:28deg">₹140 customer limit</span>
        <span style="--angle:72deg">₹90 yesterday</span>
      </div>
      <p>Compare it with cost, a competitor, willingness-to-pay, or yesterday.<br /><strong>The comparison creates the information.</strong></p>
    </aside>
  </div>
</section>

<section class="think-act think-act--frame" id="frame" data-act="frame" aria-labelledby="frame-title">
  <div class="think-act-head">
    <p class="think-act-number">Act 03</p>
    <h2 id="frame-title">Your starting assumption shapes what you notice.</h2>
    <p>Two people can study the same facts and focus on completely different things.</p>
  </div>

  <div class="think-lens-row">
    <article class="think-lens">
      <p class="think-thinker">Thomas Kuhn</p>
      <h3>Frames and inconvenient facts</h3>
      <p>A frame tells you which questions to ask and which evidence to take seriously. That is why two analysts can read the same filing and reach very different conclusions.</p>
      <aside class="think-example"><span>Frame A</span>When deciding whether to invest in an airline, start with “airlines are cyclical businesses.” You notice fixed costs and demand swings.</aside>
      <aside class="think-example"><span>Frame B</span>Or start with “scarce airport slots protect this airline’s routes.” You notice entry barriers, route concentration, and pricing discipline.</aside>

      <h4>Pay attention to what does not fit</h4>
      <p>You believe commodity producers cannot sustain high returns, yet one firm does so across several cycles. It may be temporary—or your original view may be missing a real advantage.</p>
      <p>The useful habit is simple: when a fact does not fit your view, do not explain it away too quickly.</p>
      <blockquote class="think-question">What is my current view making hard to notice?</blockquote>
    </article>

    <aside class="think-diagram think-diagram--frames" data-frame-demo="" aria-label="One airline seen through three competing frames">
      <div class="frame-tabs" role="group" aria-label="Choose an analytical frame">
        <button type="button" class="is-active" data-frame="bad">Cyclical business</button>
        <button type="button" data-frame="slots">Slot scarcity</button>
        <button type="button" data-frame="network">Network scale</button>
      </div>
      <div class="frame-object" aria-hidden="true"><span>Investment question: is this airline attractive?</span><strong>ONE AIRLINE</strong></div>
      <div class="frame-facts" aria-live="polite">
        <span>Fixed costs</span><span>Fuel exposure</span><span>Demand cycles</span>
      </div>
      <p data-frame-copy="">Frame: “Airlines are cyclical.” So cost volatility and demand cycles dominate.<br /><strong>The frame decides what gets highlighted.</strong></p>
    </aside>
  </div>
</section>

<section class="think-act think-act--limits" id="limits" data-act="limits" aria-labelledby="limits-title">
  <div class="think-act-head">
    <p class="think-act-number">Act 04</p>
    <h2 id="limits-title">You cannot know everything. Decide what is enough.</h2>
    <p>The world may be complicated. Your decision rule does not always need to be.</p>
  </div>

  <div class="think-lens-row">
    <article class="think-lens">
      <p class="think-thinker">Herbert Simon</p>
      <h3>Bounded rationality</h3>
      <p>Perfect optimisation assumes that you can see every option and consequence. In practice, time, information, and attention are limited. A good process must include a point at which you stop searching and decide.</p>

      <h4>Satisficing</h4>
      <p>Set a clear bar and choose an option that clears it. In hiring, define the non-negotiables—competence, reliability, learning ability, and compensation fit—then hire when a candidate meets them well enough.</p>

      <h4>Search cost</h4>
      <p>If another week of research has only a 5% chance of changing a small purchase decision, continued analysis may be less rational than acting and learning.</p>

      <h4>Adaptive strategy</h4>
      <p>Under uncertainty, a choice that is good enough, reversible, and informative can beat a polished five-year plan. Pilot five stores, learn, and update before planning a national rollout.</p>
      <blockquote class="think-question">How much uncertainty is actually reducible before I need to act?</blockquote>
    </article>

    <aside class="think-diagram think-diagram--search" aria-label="The diminishing benefit of more research crosses the rising cost of delay, showing when to act">
      <span class="diagram-label">Example: deciding whether to research onboarding drop-off for another week</span>
      <div class="search-chart">
        <svg viewBox="0 0 420 310" role="img" aria-labelledby="search-title search-desc">
          <title id="search-title">When to stop researching and act</title>
          <desc id="search-desc">The expected gain from another week of research declines while the cost of delaying the launch rises. Their crossing point is the decision point.</desc>
          <line class="search-axis" x1="56" y1="258" x2="390" y2="258" /><line class="search-axis" x1="56" y1="258" x2="56" y2="25" />
          <path class="search-benefit" d="M68 52 C144 58, 208 88, 276 156 S348 218, 380 231" />
          <path class="search-cost-line" d="M68 235 L380 68" />
          <line class="search-marker" x1="250" y1="258" x2="250" y2="142" />
          <circle class="search-dot" cx="250" cy="142" r="7" />
          <text class="search-benefit-label" x="82" y="45">Expected gain from another week</text>
          <text class="search-cost-label" x="270" y="84">Cost of delaying the launch</text>
          <text class="search-y-label" x="17" y="36">Expected value</text>
          <text class="search-x-label" x="188" y="292">Weeks spent researching</text>
        </svg>
        <div class="search-stop-label"><strong>Week 3: stop → act</strong><span>After this point, the likely insight from another week is worth less than the cost of delaying the experiment.</span></div>
      </div>
      <p>Thinking has a marginal cost.<br /><strong>A stopping rule is part of reason.</strong></p>
    </aside>
  </div>

  <div class="think-lens-row think-lens-row--reverse">
    <article class="think-lens">
      <p class="think-thinker">Gerd Gigerenzer</p>
      <h3>When a simple rule is enough</h3>
      <p>A decision rule does not need to copy all the complexity of the world. In the right setting, a few reliable signals can be clearer and more robust than a large model.</p>

      <h4>Less can be more</h4>
      <p>A 40-variable credit model may look sophisticated. If repayment history, income stability, and debt burden capture most default risk, the simpler rule may generalise better and fail more transparently.</p>

      <h4>The task matters</h4>
      <p>Understanding, prediction, and decision are different jobs. You may need a detailed model to understand churn. To decide which support tickets need immediate attention, three clear rules may work better.</p>

      <p>First understand enough of the situation to avoid a crude answer. Then ask whether more detail will actually change the choice. If it will not, stop.</p>
      <blockquote class="think-question">What is the simplest rule that captures enough of what matters for this decision?</blockquote>
    </article>
    <aside class="think-diagram think-diagram--simplify" aria-label="A small-business loan decision using three clear cues instead of twenty possible signals">
      <div class="variable-cloud" aria-hidden="true"><span>20 possible signals<br />for a small-business loan</span></div>
      <div class="simplify-arrow">→</div>
      <ol><li>Repayment history</li><li>Income stability</li><li>Debt burden</li></ol>
      <p>For this loan decision, three cues may beat twenty weak ones.<br /><strong>Match the rule to the environment.</strong></p>
    </aside>
  </div>
</section>

<section class="think-act think-act--probability" id="probability" data-act="probability" aria-labelledby="probability-title">
  <div class="think-act-head">
    <p class="think-act-number">Act 05</p>
    <h2 id="probability-title">Use probabilities, not declarations.</h2>
    <p>A belief should change when the evidence changes.</p>
  </div>

  <div class="think-lens-row">
    <article class="think-lens">
      <p class="think-thinker">Bayes</p>
      <h3>Put a number on the belief</h3>
      <p>Start with an estimate, look at the new evidence, and revise the estimate. You do not need the equation to use the habit.</p>
      <p>A company has a 30% chance of losing its largest customer within two years. The customer begins testing a rival: 30% → 45%. A contract extension arrives: 45% → 35%. Updating is not indecision; it is the point.</p>

      <h4>Diagnostic evidence</h4>
      <p>“The company raised price 15%” is weak evidence of pricing power if the entire industry raised price. “It raised price 15%, competitors did not, and volume still grew” is much harder to explain without genuine pricing power.</p>
      <blockquote class="think-question">How much more likely is this evidence under my hypothesis than under the alternatives?</blockquote>
    </article>
    <aside class="think-diagram think-diagram--bayes" data-bayes-demo="" aria-label="Interactive probability update">
      <div class="bayes-readout"><strong data-probability="">30%</strong><span>chance customer leaves</span></div>
      <div class="bayes-track"><i data-probability-bar=""></i></div>
      <div class="bayes-events" role="group" aria-label="Update with evidence">
        <button type="button" class="is-active" data-probability-value="30">Prior</button>
        <button type="button" data-probability-value="45">Tests rival</button>
        <button type="button" data-probability-value="35">Extends contract</button>
      </div>
      <p>Write down the number.<br /><strong>Change it when the facts change.</strong></p>
    </aside>
  </div>

  <div class="think-probability-grid">
    <article>
      <p class="think-thinker">Base rates</p>
      <h3>Check what usually happens</h3>
      <p>The story explains why this case feels special. The base rate tells you what usually happens in similar cases. A brilliant restaurant concept or charismatic turnaround CEO can improve the odds, but should not erase them.</p>
      <blockquote class="think-question">What happens in the relevant reference class before I tell myself why this case is different?</blockquote>
    </article>
    <article>
      <p class="think-thinker">Kahneman &amp; Tversky</p>
      <h3>Four common probability mistakes</h3>
      <ul class="think-bug-list">
        <li><strong>Representativeness</strong><span>Resemblance is mistaken for probability.</span></li>
        <li><strong>Availability</strong><span>Vivid examples feel statistically common.</span></li>
        <li><strong>Anchoring</strong><span>A starting number contaminates later estimates.</span></li>
        <li><strong>Base-rate neglect</strong><span>The specific story overwhelms the statistical pattern.</span></li>
      </ul>
      <blockquote class="think-question">Is this shortcut a bug here—or a useful heuristic?</blockquote>
    </article>
    <article>
      <p class="think-thinker">Philip Tetlock</p>
      <h3>Make forecasts you can check</h3>
      <p>Replace “confident” with a number. Define the event, record the forecast, update it, and later see whether you were right.</p>
      <p>“AI will transform banking” cannot be scored. A dated forecast about the five largest private banks, a measurable service-interaction threshold, and a definition of human escalation can.</p>
      <p>Before the evidence arrives, write down what would change your mind.</p>
      <blockquote class="think-question">What probability am I assigning, and what evidence would move it by 10–15 points?</blockquote>
    </article>
  </div>
</section>

<section class="think-specimen" id="specimen" aria-labelledby="specimen-title">
  <div class="specimen-switch" role="tablist" aria-label="Choose a worked example">
    <button type="button" class="is-active" role="tab" aria-selected="true" data-specimen-choice="investing">Investing</button>
    <button type="button" role="tab" aria-selected="false" data-specimen-choice="product">Product management</button>
  </div>
  <div data-specimen-case="investing">
    <div class="think-specimen-intro"><p class="think-kicker">Worked example · investing</p><h2 id="specimen-title">A “cheap” company, seen six ways.</h2><p>Company X trades at 12× earnings while peers trade at 25×. That makes it look cheap. It does not tell us whether it is a good investment.</p></div>
    <div class="think-specimen-stage">
      <div class="think-specimen-steps">
        <article data-specimen-step="0"><span>Start with the claim</span><h3>It trades at 12×.</h3><p>That is a useful fact, but not yet an investment case.</p></article><article data-specimen-step="1"><span>Check what is missing</span><h3>Look beyond the multiple.</h3><p>The P/E says nothing about debt, cyclicality, reinvestment, accounting quality, or whether today’s earnings will last.</p></article><article data-specimen-step="2"><span>Trace the business</span><h3>Find out why it is cheap.</h3><p>Lower quality can lead to churn, lower volume, poor utilisation, weaker margins, and less money to reinvest. Cheapness may be a symptom, not an opportunity.</p></article><article data-specimen-step="3"><span>Try other explanations</span><h3>What story fits the facts?</h3><p>It could be undervalued, in structural decline, or simply near the bottom of a cycle. Each explanation requires different evidence.</p></article><article data-specimen-step="4"><span>Set a rule</span><h3>Know what would make you pass.</h3><p>For example: avoid highly leveraged cyclical companies with weak interest coverage, however cheap they look.</p></article><article data-specimen-step="5"><span>Write the forecast</span><h3>Make the thesis testable.</h3><p>Start with the success rate of similar turnarounds, estimate the odds, and write down what evidence would raise or lower them.</p></article>
      </div>
      <aside class="specimen-visual" data-specimen-visual="" data-step="0" aria-live="polite"><p class="specimen-label">COMPANY X</p><div class="specimen-multiple"><strong>12×</strong><span>earnings</span></div><div class="specimen-peer">Peers <strong>25×</strong></div><div class="specimen-omissions"><span>debt</span><span>cycle</span><span>quality</span><span>reinvestment</span></div><div class="specimen-loop"><span>quality</span><i>→</i><span>churn</span><i>→</i><span>volume</span><i>→</i><span>margin</span></div><div class="specimen-frames"><span>undervalued</span><span>decline</span><span>cyclical trough</span></div><div class="specimen-rule">LEVERAGE + WEAK COVERAGE <strong>→ PASS</strong></div><div class="specimen-probability"><span>Turnaround probability</span><strong>38%</strong><i></i></div><p class="specimen-caption" data-specimen-caption="">A low multiple is only a starting point.</p></aside>
    </div>
  </div>
  <div data-specimen-case="product" hidden="">
    <div class="think-specimen-intro"><p class="think-kicker">Worked example · product management</p><h2>Daily active users have stopped growing.</h2><p>Before launching notifications or campaigns, work out what has actually changed.</p></div>
    <div class="think-specimen-stage">
      <div class="think-specimen-steps">
        <article data-specimen-step="0"><span>Start with the claim</span><h3>DAU has stopped growing.</h3><p>That tells us what happened, not why.</p></article><article data-specimen-step="1"><span>Break down the metric</span><h3>What changed inside DAU?</h3><p>Check new-user activation, retained users, usage frequency, churn, and seasonality before choosing a remedy.</p></article><article data-specimen-step="2"><span>Trace the product loop</span><h3>Find the weak link.</h3><p>A poor first session can lower retention. That means less content and fewer referrals, which makes the product weaker for the next group of users.</p></article><article data-specimen-step="3"><span>Try other explanations</span><h3>What kind of problem is this?</h3><p>It may be a marketing problem, an activation problem, or simply a product people do not need every day. Each calls for a different test.</p></article><article data-specimen-step="4"><span>Choose the decision metrics</span><h3>Focus on two numbers.</h3><p>Use first-session completion and seven-day retention. Leave the rest aside unless it can change the roadmap.</p></article><article data-specimen-step="5"><span>Run the test</span><h3>Try one reversible change.</h3><p>Pilot a clearer onboarding path for one cohort. Decide in advance how much retention must improve before you roll it out.</p></article>
      </div>
      <aside class="specimen-visual" data-specimen-visual="" data-step="0" aria-live="polite"><p class="specimen-label">PRODUCT APP</p><div class="specimen-multiple"><strong>DAU</strong><span>flat growth</span></div><div class="specimen-peer">7-day retention <strong>18%</strong></div><div class="specimen-omissions"><span>activation</span><span>frequency</span><span>churn</span><span>seasonality</span></div><div class="specimen-loop"><span>first use</span><i>→</i><span>retention</span><i>→</i><span>content</span><i>→</i><span>referrals</span></div><div class="specimen-frames"><span>marketing problem</span><span>activation problem</span><span>occasional use</span></div><div class="specimen-rule">FIRST SESSION + 7-DAY RETENTION <strong>→ DECIDE</strong></div><div class="specimen-probability"><span>Chance onboarding improves retention</span><strong>45%</strong><i></i></div><p class="specimen-caption" data-specimen-caption="">DAU tells us what happened, not why.</p></aside>
    </div>
  </div>
</section>

<section class="think-act think-act--stack" id="stack" data-act="stack" aria-labelledby="stack-title">
  <div class="think-act-head">
    <p class="think-act-number">Act 06</p>
    <h2 id="stack-title">Put the ideas to work.</h2>
    <p>No single approach is enough. The value comes from using them together on the same decision.</p>
  </div>

  <div class="think-tensions">
    <article><p><strong>Understand first</strong> Do not simplify the situation too early.</p><span>↔</span><p><strong>Decide eventually</strong> Do not demand a complete theory before acting.</p></article>
    <article><p><strong>Change the frame</strong> Sometimes the model itself is wrong.</p><span>↔</span><p><strong>Update the odds</strong> Sometimes the model is fine and the probability changed.</p></article>
    <article><p><strong>Use the map</strong> A useful model must leave things out.</p><span>↔</span><p><strong>Check the gaps</strong> Know what was removed and whether it matters.</p></article>
    <article><p><strong>Look one layer deeper</strong> There is usually more to understand.</p><span>↔</span><p><strong>Know when to stop</strong> If more detail will not change the choice, act.</p></article>
  </div>

  <section class="think-workflow" aria-labelledby="workflow-title">
    <div class="think-workflow-head">
      <p class="think-kicker">A practical workflow</p>
      <h3 id="workflow-title">Twelve steps for a real decision</h3>
    </div>
    <ol>
      <li><span>01</span><div><strong>Define the problem</strong><p>Write the question clearly. Expose assumptions built into its wording.</p></div></li>
      <li><span>02</span><div><strong>Separate map from territory</strong><p>List the categories, metrics, and models in use; note what each omits.</p></div></li>
      <li><span>03</span><div><strong>Trace the system</strong><p>Map what builds up, what flows in and out, where feedback appears, and where delays sit.</p></div></li>
      <li><span>04</span><div><strong>Map relationships</strong><p>Find variables whose meaning depends on each other rather than standing alone.</p></div></li>
      <li><span>05</span><div><strong>State your starting view</strong><p>Write down your frame and try at least one competing explanation.</p></div></li>
      <li><span>06</span><div><strong>Find the cause</strong><p>Trace the process that produced the visible result.</p></div></li>
      <li><span>07</span><div><strong>Set a search limit</strong><p>Identify which unknowns can actually change the decision.</p></div></li>
      <li><span>08</span><div><strong>Test a simple rule</strong><p>Ask whether a few robust cues can make the decision reliably.</p></div></li>
      <li><span>09</span><div><strong>Check the base rate</strong><p>Find out what usually happens in similar cases before getting lost in this story.</p></div></li>
      <li><span>10</span><div><strong>Assign a probability</strong><p>Use a number rather than “likely,” “confident,” or “bullish.”</p></div></li>
      <li><span>11</span><div><strong>Search for anomalies</strong><p>Find the strongest fact your model struggles to explain.</p></div></li>
      <li><span>12</span><div><strong>Update, decide, review</strong><p>Move with the evidence; act when analysis has low value; later score the forecast.</p></div></li>
    </ol>
  </section>

  <section class="think-transfer" aria-labelledby="transfer-title">
    <div>
      <p class="think-kicker">Transfer test</p>
      <h3 id="transfer-title">“Country X will become a manufacturing superpower.”</h3>
      <p>Turn the slogan into questions you can answer.</p>
    </div>
    <ul>
      <li><strong>Define success.</strong> Export share? Manufacturing value added? Employment? Technological sophistication?</li>
      <li><strong>State the story.</strong> China replacement, demographic dividend, friend-shoring, or domestic-market scale?</li>
      <li><strong>Trace what is required.</strong> Power, logistics, labour, capital, education, regulation, suppliers, currency, and scale.</li>
      <li><strong>Compare costs properly.</strong> “Cheap wages” mean little without productivity, energy, logistics, quality, and training costs.</li>
      <li><strong>Look within the country.</strong> Regions and institutions may differ enough to produce very different outcomes.</li>
      <li><strong>Say what would prove you wrong.</strong> Identify the few variables that could kill the thesis and write a dated forecast.</li>
    </ul>
  </section>

  <section class="think-memory" aria-labelledby="memory-title">
    <div class="think-memory-head">
      <p class="think-kicker">Keep these</p>
      <h3 id="memory-title">The 12 questions worth memorising</h3>
    </div>
    <ol>
      <li>What generated this?</li>
      <li>Do I understand it, or have I only named it?</li>
      <li>What did my model leave out?</li>
      <li>What system would naturally produce this behaviour?</li>
      <li>Where are the feedback loops and delays?</li>
      <li>What relationships create the pattern?</li>
      <li>What does my framework make hard to see?</li>
      <li>What can I realistically know before I have to act?</li>
      <li>What is the simplest rule that captures enough?</li>
      <li>What is the base rate?</li>
      <li>What probability am I actually assigning?</li>
      <li>What evidence would materially change that number?</li>
    </ol>
  </section>

  <section class="think-integration" aria-labelledby="integration-title">
    <p class="think-kicker">The practical takeaway</p>
    <h3 id="integration-title">You do not need to memorise the thinkers. Use these five checks.</h3>
    <ol>
      <li><strong>Name the claim.</strong> What are you calling this, and what does that label leave out?</li>
      <li><strong>Explain the mechanism.</strong> What process, relationships, incentives, and feedback loops produced it?</li>
      <li><strong>Challenge the frame.</strong> What competing explanation would direct your attention somewhere else?</li>
      <li><strong>Set the bar for action.</strong> Which unknowns can actually change the decision, and when will you stop researching?</li>
      <li><strong>Put a number on it.</strong> What probability are you assigning, and what evidence would change it?</li>
    </ol>
  </section>
</section>]]></content><author><name>Yash Tambawala</name></author><summary type="html"><![CDATA[A practical field guide to models, systems, uncertainty, and changing your mind]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://yashtambawala.com/assets/og/how-to-think.png" /><media:content medium="image" url="https://yashtambawala.com/assets/og/how-to-think.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">India in a World of Open Markets and Closed Chokepoints</title><link href="https://yashtambawala.com/2026/07/26/india-in-a-world-of-open-markets-and-closed-chokepoints.html" rel="alternate" type="text/html" title="India in a World of Open Markets and Closed Chokepoints" /><published>2026-07-26T19:49:00+05:30</published><updated>2026-07-26T19:49:00+05:30</updated><id>https://yashtambawala.com/2026/07/26/india-in-a-world-of-open-markets-and-closed-chokepoints</id><content type="html" xml:base="https://yashtambawala.com/2026/07/26/india-in-a-world-of-open-markets-and-closed-chokepoints.html"><![CDATA[<p>Two recent technology debates appear unrelated.</p>

<p>The first began when <a href="https://www.tomshardware.com/tech-industry/artificial-intelligence/nvidia-and-24-other-companies-sign-open-weights-letter-as-washington-weighs-chinese-ai-model-ban">NVIDIA and other technology companies</a> asked the US government to avoid premature restrictions on open-weight AI models. Critics pointed out that CUDA, NVIDIA’s most valuable software ecosystem, remains proprietary. The second concerns the Cockroach Janta Party, or CJP, which built an enormous following on Instagram before becoming part of a wider protest movement in India.</p>

<p>One debate concerns artificial intelligence and the other political attention. Both raise the same question: who controls the infrastructure through which a society communicates, builds technology, and participates in the modern economy?</p>

<p>Edward Fishman’s <em><a href="https://www.penguinrandomhouse.com/books/726149/chokepoints-by-edward-fishman/">Chokepoints</a></em> offers a useful framework. Fishman applies it to economic warfare, but the logic extends to digital networks. Modern power often comes from controlling a few indispensable nodes that others must use and cannot easily replace. These chokepoints usually begin as successful commercial products. Dependence builds quietly. Eventually, access can be priced, prioritised, restricted, or withdrawn.</p>

<p>A system can look open on the surface while remaining concentrated underneath.</p>

<h2 id="the-two-gateways-into-the-digital-economy">The Two Gateways into the Digital Economy</h2>

<p>Attention and artificial intelligence are two gateways into the same digital economy.</p>

<p>The attention layer determines what people see, which businesses find customers, and which political movements gain visibility. The intelligence layer determines what developers can build, what enterprises can automate, and who captures the gains from AI. Beneath both lie the harder layers: semiconductor equipment, fabs, accelerators, cloud data centres, electricity, software ecosystems, advertising systems, and global distribution.</p>

<p>India is deeply involved at the upper layers. It supplies users, creators, developers, advertisers, and demand. Control of much of the underlying machinery sits elsewhere. That is the connection between Instagram, open weights, CUDA, TSMC, ASML, and the large cloud providers.</p>

<h2 id="attention-is-infrastructure">Attention Is Infrastructure</h2>

<p>The CJP example should be handled carefully. Its follower count does not reveal where its audience came from, and it does not prove foreign funding or coordination. What we do know is that a movement founded by a Boston University student <a href="https://apnews.com/article/india-cockroach-party-youth-movement-protest-modi-59d045fbd485635def01b01f3898307b">grew rapidly on Instagram and later acquired a street presence</a> after its founder returned to India.</p>

<p>The grievances may be domestic, the supporters genuine, and the movement organic. The structural point remains. An organisation can build legitimacy and mobilise people in India through distribution infrastructure controlled by a foreign company.</p>

<p>Instagram determines how content is recommended. Meta controls moderation, appeals, verification, advertising rules, and most of the data available to researchers. A recommendation change, an automated moderation mistake, or a policy revision can materially damage an organisation that depends on the platform. Malicious intent is not necessary. Dependency itself creates power.</p>

<p>This extends far beyond politics. Indian businesses find customers through Google and Meta. Professionals build reputations through LinkedIn. Creators depend on YouTube and Instagram. App companies rely on Android and iOS. India bears the social and political consequences while foreign companies operate much of the distribution machinery.</p>

<p>Attention should therefore be treated as strategic infrastructure. This does not justify banning foreign platforms or placing political speech under direct government control. Both responses could cause more harm than the dependence they seek to address.</p>

<p>India instead needs stronger institutions for the digital public sphere. Meta already maintains a <a href="https://about.fb.com/news/2018/12/ad-transparency-in-india/">searchable archive for political advertising in India</a>, but formal political ads are only part of the influence system. Transparency should also cover issue-based campaigns, proxy advertisers, paid creators, and cross-border sponsorship. Users need meaningful appeals against automated moderation, while independent researchers need better access to platform data.</p>

<p>Regulation is the defensive response. The offensive response is to build important Indian networks of our own.</p>

<p>UPI shows how public digital infrastructure can widen access and competition. Yet open rails do not guarantee that Indian companies will own the customer relationship, behavioural data, intelligence layer, or distribution built above them. An open protocol at the bottom can coexist with concentration at the application layer.</p>

<h2 id="the-same-structure-appears-in-ai">The Same Structure Appears in AI</h2>

<p>Open models are genuinely valuable. They reduce dependence on a few foreign APIs, lower experimentation costs, allow local deployment, and give enterprises more control over data and inference. NVIDIA’s commercial interest does not make those benefits any less real.</p>

<p>But the open-weights debate is also a negotiation over which part of the AI stack becomes cheaper and which parts retain pricing power.</p>

<p>The commercial logic is simple. Companies favour openness in layers that create demand for their bottlenecks and protect the layers that make them difficult to replace. NVIDIA can contribute heavily to open-source AI while keeping CUDA under its control. Open models encourage more training and inference, which increases demand for accelerators, networking, servers, and NVIDIA’s software ecosystem. Opening model weights makes models easier to use. Making CUDA hardware-neutral would make NVIDIA hardware easier to replace.</p>

<p>The pattern appears elsewhere. Meta can release model weights because its deepest advantages lie in attention, advertising, data, and distribution. Microsoft can support open models while protecting cloud contracts, enterprise identity, and customer relationships. Companies want the inputs they buy to become cheaper while the products they sell remain differentiated. That is ordinary platform economics.</p>

<p>AI adoption adds further pressure. Chip companies want more workloads, cloud companies want higher utilisation, and application companies want intelligence to become an inexpensive input. Open weights, distillation, routing, smaller models, and self-hosting all push model prices down.</p>

<p>The harder bottlenecks remain below: advanced lithography, foundry capacity, high-bandwidth memory, packaging, accelerator design, interconnects, CUDA, electricity, and data centres. NVIDIA does not need one model company to win. It needs total AI computation to grow. Cloud providers do not need model companies to preserve large margins. They need AI workloads consuming cloud capacity.</p>

<p>The economic pressure therefore runs in one direction. Model intelligence becomes cheaper and more abundant while the owners of scarce infrastructure continue to collect the toll.</p>

<h2 id="open-weights-are-useful-they-are-not-sovereignty">Open Weights Are Useful. They Are Not Sovereignty.</h2>

<p>A company can release final model weights while retaining its training data, data-selection methods, training code, post-training systems, evaluations, inference optimisations, and production infrastructure. Indian developers can still run the model locally, adapt it, and reduce their dependence on a single API. But receiving the weights is not the same as receiving the productive system that created them.</p>

<p>Open weights are like receiving an industrial machine rather than the factory that built it. You can operate and modify the machine without possessing the knowledge, tooling, capital, supply chain, or organisation required to reproduce it or build its successor.</p>

<p>India’s progression has five stages: API access, weight access, adaptation capability, reproduction capability, and original technological leadership. Most discussions stop at stage two and call it democratisation. A country running foreign-designed models on foreign accelerators through foreign software and cloud infrastructure has gained useful access. It has not gained sovereignty. It has diversified its dependence.</p>

<h2 id="open-markets-and-physical-chokepoints">Open Markets and Physical Chokepoints</h2>

<p>The attention and AI stories converge at the infrastructure layer.</p>

<p>ASML holds a critical position in advanced lithography. TSMC combines process knowledge, yield learning, packaging, and customer trust. NVIDIA combines accelerators with CUDA, networking, libraries, and developer familiarity. The large cloud providers combine capital, electricity, data centres, software, and enterprise distribution.</p>

<p>These positions were built over decades. Open markets do not distribute power evenly because participants enter them with different levels of capital, knowledge, scale, and control over standards. Free trade and open source still benefit latecomers by lowering costs and accelerating learning. But access is not productive capacity. Model access is not model-building capability. GPU rental is not control of the compute stack. Access to a foreign fab is not accumulated process knowledge.</p>

<p>Chokepoints are not permanent. Once they are used for coercion, customers and countries have stronger incentives to find substitutes. Their power comes from the fact that replacement is slow, expensive, and uncertain. That delay is where leverage lives.</p>

<h2 id="from-access-to-capability">From Access to Capability</h2>

<p>Indian technology policy often focuses on access to GPUs, foreign fabs, open models, cloud capacity, capital, and export markets. All of it is useful. But access still leaves someone else in control of the terms. The sharper question is what the rest of the world will eventually need from India.</p>

<p>Fishman’s framework gives India two jobs.</p>

<p>The first is <strong>resilience</strong>. India must reduce the leverage others hold over it through supplier diversification, interoperability, domestic maintenance, alternative vendors, and credible substitutes for critical foreign systems.</p>

<p>The second is <strong>indispensability</strong>. India must build capabilities that others cannot easily replace through patient capital, domestic procurement, technical learning, supplier coordination, standards, exports, and developer ecosystems.</p>

<p>Resilience reduces the leverage others hold over India. Indispensability creates leverage for India. A serious strategy needs both.</p>

<p>India cannot build everything, so it must choose. The strongest candidates will combine large domestic demand, an existing base of skills and suppliers, learning that spreads into adjacent industries, export potential, and a realistic path to global relevance. Semiconductor equipment subsystems, advanced packaging, power electronics, grid technology, multilingual computing, and software for the physical economy deserve consideration by those standards.</p>

<p>In AI, the capability ladder runs from consuming APIs to adapting models, building training systems, evaluations, compilers, and inference platforms, and eventually creating original model factories. In attention and distribution, the equivalent task is to create accountable rules for foreign platforms while building Indian networks to meaningful scale.</p>

<h2 id="the-real-question">The Real Question</h2>

<p>This was never mainly about whether NVIDIA is hypocritical or whether one Instagram account is suspicious. It is about the architecture of power in an economy where communication, commerce, and innovation move through a small number of privately controlled networks and industrial systems.</p>

<p>The CJP episode shows how a political movement can grow through infrastructure controlled outside India. The open-weights debate shows how one layer of a technology stack can open while the layers beneath remain concentrated. The semiconductor industry shows how open trade can coexist with extraordinary control over a handful of machines and production systems.</p>

<p>India should continue using global platforms, foreign capital, open-source software, open models, and international supply chains. The goal is not autarky. It is selective indispensability and meaningful control over the dependencies that matter.</p>

<p>Twenty years from now, India should not simply supply users, labour, data, and demand for systems someone else controls. There should be parts of the global economy where other countries need India’s technology, manufacturing systems, networks, and knowledge, with no easy substitute.</p>

<p>That is what a moat looks like at the level of a nation.</p>]]></content><author><name>Yash Tambawala</name></author><summary type="html"><![CDATA[Why access to technology and platforms is not the same as control over them]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://yashtambawala.com/assets/og/india-in-a-world-of-open-markets-and-closed-chokepoints.png" /><media:content medium="image" url="https://yashtambawala.com/assets/og/india-in-a-world-of-open-markets-and-closed-chokepoints.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">You May Not Care About Politics, But Politics Cares About You</title><link href="https://yashtambawala.com/2026/05/30/you-may-not-care-about-politics-but-politics-cares-about-you.html" rel="alternate" type="text/html" title="You May Not Care About Politics, But Politics Cares About You" /><published>2026-05-30T00:00:00+05:30</published><updated>2026-05-30T00:00:00+05:30</updated><id>https://yashtambawala.com/2026/05/30/you-may-not-care-about-politics-but-politics-cares-about-you</id><content type="html" xml:base="https://yashtambawala.com/2026/05/30/you-may-not-care-about-politics-but-politics-cares-about-you.html"><![CDATA[<p>People often say they do not care about politics. Usually they mean something narrower: they do not care for party arguments, television debates, ideological branding, campaign gossip, or the daily theatre of public life.</p>

<p>That reaction is understandable. A great deal of political conversation is low quality. It is noisy, tribal, repetitive, and often designed to exhaust rather than enlighten. Serious people respond by withdrawing from it. They focus on work, investing, family, health, business, travel, education, and self-improvement. They assume that if they avoid politics, politics will avoid them.</p>

<p>That is one of the great illusions of modern life.</p>

<p>Politics is not just what politicians say. Politics is the machinery that decides what can be owned, what can be taxed, what can be licensed, what can be exported, what counts as money, what counts as a valid contract, which industries are legitimate, which profits are acceptable, which people are trusted, and which exits remain open.</p>

<p>You may opt out of political discussion. You do not get to opt out of political consequences.</p>

<h2 id="markets-are-political-settlements">Markets Are Political Settlements</h2>

<p>The error begins when we imagine markets as if they existed in a clean room, separate from power. We picture the economy as a contractual domain where people freely exchange goods, sign agreements, raise capital, build businesses, and allocate resources. Politics appears only as interference.</p>

<p>This view is emotionally attractive because it lets us believe private competence can insulate us from public disorder. Work hard, save money, invest well, follow the law, get the right passport, choose the right country. That is the fantasy.</p>

<p>The harsher truth is that markets are political settlements that have become stable enough to look natural.</p>

<p>A bank account is a political settlement. A currency note is a political settlement. A mining lease, telecom license, airport concession, payment license, drug approval, education permit, tax rule, zoning approval, or power tariff is a political settlement. The market is what remains after politics has decided what may be owned, exchanged, taxed, subsidized, banned, inspected, capped, protected, punished, or nationalized.</p>

<p>When that settlement is stable, we call it the rule of law. When it shifts suddenly, we call it political risk. But it was political in both cases.</p>

<h2 id="the-state-can-reopen-yesterday">The State Can Reopen Yesterday</h2>

<p>The clearest example is retrospective taxation.</p>

<p>Most people assume the past is closed. If a transaction was completed under one tax understanding, and the state later changes the law, the new law should apply only going forward. That is how ordinary fairness works.</p>

<p>Sovereign power has a subtler instrument available to it. It can say the rule is not new at all. It can call the amendment clarificatory or explanatory. It can argue that taxpayers misunderstood what the law always meant. With that move, a new burden becomes an old obligation. A policy change becomes an interpretation. The past is reopened without the state admitting that it changed the past.</p>

<p>India has used this logic in ways that should permanently alter how investors think about legal certainty. The Vodafone retrospective tax episode was the warning shot. The state lost in court, then changed the law retrospectively. The online gaming GST dispute revealed the same instinct in another form: what the industry saw as a new tax treatment, the state framed as a clarification of the old law.</p>

<p>The lesson is not merely that India can be aggressive on tax. The lesson is larger. If the fiscal or political stakes are high enough, the state may not simply tax the future. It may reinterpret yesterday.</p>

<h2 id="money-is-a-legal-object-before-it-is-an-economic-one">Money Is a Legal Object Before It Is an Economic One</h2>

<p>Once you understand that move, you begin to see the same structure elsewhere. Money itself is political.</p>

<p>We use it every day, so it feels natural. But a currency note is valuable because the state maintains the legal and institutional system in which others must accept it. India’s demonetization made this visible. One evening, widely used currency notes were money. Then the state announced that they would cease to be legal tender. The paper had not changed. Its legal status had changed. Wealth held in one form became a claim that had to pass through an administrative process.</p>

<p>The United States offers an even deeper example. During the Great Depression, many contracts contained gold clauses that protected creditors by linking payment obligations to gold. When those clauses obstructed monetary policy, the government invalidated them. Contract sanctity yielded to monetary sovereignty.</p>

<p>This did not happen in some openly lawless regime. It happened inside a constitutional democracy with courts, lawyers, property rights, and a deep commercial tradition. The point is simple: money is not merely an economic object. It is a legal and political object whose character can change when the state decides a higher priority is at stake.</p>

<p>Bank deposits carry the same lesson. People treat deposits as safe private money, but deposits sit inside a hierarchy of banking law, resolution rules, deposit insurance limits, central bank policy, and crisis politics. Cyprus in 2013 showed what happens when that hierarchy is rearranged. Large uninsured depositors discovered that their deposits could become part of a bank rescue. What looked like cash became loss-absorbing capital.</p>

<p>In normal times, a deposit is money. In crisis, it can become an instrument of policy. The difference is political necessity.</p>

<h2 id="licenses-are-assets-built-on-permission">Licenses Are Assets Built on Permission</h2>

<p>Much of modern capitalism runs on permission.</p>

<p>Telecom spectrum, mining blocks, airport concessions, banking licenses, NBFC registrations, insurance approvals, environmental clearances, land-use permissions, hospital approvals, and power purchase agreements all fall into this category. Investors treat these permissions as assets. They lend against them, value them, securitize them, and build projections around them.</p>

<p>But a license is not property in the ancient sense. It is a permission structure granted by the state, and its value depends not only on the text of the document but on the legitimacy of the political process that produced it.</p>

<p>India’s 2G spectrum cancellations made this painfully clear. Companies had licenses, business plans, investors, debt, employees, and operating assumptions. Then the allocation process itself was judged illegitimate. The asset did not merely lose value because market conditions changed. Its legal origin was attacked.</p>

<p>The same principle appeared in coal block cancellations. If a business depends on a state-granted scarce resource, it owns more than the resource. It owns the political history of how that resource was allocated. If that history becomes unacceptable, the asset becomes fragile.</p>

<h2 id="contracts-exist-inside-a-sovereign-system">Contracts Exist Inside a Sovereign System</h2>

<p>Contracts are supposed to create private order. But contracts do not float above the state. They exist inside a sovereign system, and when that system faces monetary stress, fiscal pressure, regulatory conflict, or mass public anger, contract language can bend.</p>

<p>India’s telecom AGR dispute is an instructive example. For years, telecom companies and the government disputed what counted as adjusted gross revenue for license fee calculations. The issue appeared in legal disclosures, investor presentations, and contingent liability notes. Then it crystallized into enormous liabilities. What had looked like an interpretive dispute became a balance-sheet event.</p>

<p>This is one of the most important lessons for investors: regulatory ambiguity is not a footnote. It is often a hidden option owned by the state.</p>

<p>The company may think the issue is probabilistic, manageable, or remote. The state may later convert that ambiguity into revenue, penalties, back dues, or operating restrictions. A spreadsheet may treat the law as an input. In reality, the law itself may be one of the variables.</p>

<h2 id="trade-platforms-and-profits-depend-on-political-legitimacy">Trade, Platforms, and Profits Depend on Political Legitimacy</h2>

<p>Trade is political in the same way. Free trade exists only until a country decides strategic autonomy, domestic employment, food security, national security, or foreign policy matter more than efficiency. Exports can be banned. Imports can be restricted. Tariffs can rise. Entire industries can be repriced by political decision.</p>

<p>The same is true of digital platforms and fast-growing industries. Many founders act as though scale alone makes a business legitimate. It does not. Legitimacy is political.</p>

<p>A payments business exists because regulators permit private actors near money. A gaming company exists because the state permits monetized chance, monetized attention, or both. A healthcare chain exists because the state permits profit inside vulnerability. An education business exists because the state permits parental anxiety to become revenue. An AI company exists inside unresolved political questions about data, labor substitution, liability, and truth.</p>

<p>Every large profit pool rests on a political settlement. Some settlements are durable. Some are fragile. Some are just waiting for a scandal.</p>

<h2 id="citizenship-is-also-a-political-asset">Citizenship Is Also a Political Asset</h2>

<p>People usually think of politics as something that affects regulation, tax, or business. But politics also governs categories of belonging.</p>

<p>Citizen, foreigner, enemy alien, refugee, illegal migrant, minority, infiltrator, dissident, dual national, non-resident, beneficial owner, politically exposed person, strategic threat: these labels can alter the legal treatment of the same human being.</p>

<p>A person’s bank account, passport, property, speech, movement, business, inheritance, and physical safety can depend on which category the state places him into. In calm times, these categories feel bureaucratic. In crisis, they become destiny.</p>

<p>The internment of Japanese Americans during World War II is a lasting warning. It happened inside a constitutional democracy with courts, elections, lawyers, newspapers, and constitutional language. That is what makes it instructive rather than exotic. A person may believe citizenship, property, and rights are secure because the legal text says so. In a crisis, another question appears: does the political community still recognize him as part of the group that deserves protection?</p>

<p>If that answer becomes uncertain, formal rights may remain on paper while lived protection collapses.</p>

<h2 id="exit-is-harder-than-it-looks">Exit Is Harder Than It Looks</h2>

<p>Albert Hirschman’s framework of exit, voice, and loyalty gives this problem its clearest structure.</p>

<p>When people face deterioration in a system, they can leave, protest, or remain attached. Exit means leaving the system. Voice means staying and trying to change it. Loyalty is the attachment that delays exit and gives voice a chance.</p>

<p>Politics matters because it determines the cost of exit, the effectiveness of voice, and the strength of loyalty.</p>

<p>Capital with free movement has exit. Capital under controls does not. A software company can sometimes move servers, headquarters, or intellectual property. A cement plant, mine, port, hospital, telecom network, utility, or airport cannot move so easily. A wealthy family may hold foreign assets, but its domestic real estate, operating business, reputation, social networks, parents, children, and citizenship remain embedded somewhere.</p>

<p>The more fixed the asset, the more political the owner. The more irreversible the investment, the more the investor depends on voice rather than exit.</p>

<p>Rich people often assume they can ignore politics because they have exit. They can invest globally, send children overseas, move money abroad, obtain another residency, and leave if things deteriorate. Sometimes that is true. But exit is not a switch. It is a door that narrows precisely when everyone wants to use it.</p>

<p>Capital controls can appear. Bank withdrawals can be limited. Foreign assets can be frozen. Tax residency can be challenged. Domestic wealth can become illiquid. Factories cannot move. Licenses cannot move. Land cannot move. Family cannot always move. Reputation cannot fully move.</p>

<p>The true cost of exit is often discovered only after the crowd has reached the door.</p>

<h2 id="not-caring-is-still-a-political-position">Not Caring Is Still a Political Position</h2>

<p>If exit is costly, voice becomes essential. This is why indifference to politics is often a luxury belief.</p>

<p>A business that depends on regulation must care about policy. A citizen who cannot easily leave must care about institutions. A depositor must care about banking rules. A founder must care about the moral legitimacy of his industry. An investor must care about state incentives. A family must care about schools, zoning, safety, healthcare, taxation, currency, and infrastructure.</p>

<p>Not caring does not make politics disappear. It simply means your voice is absent when the rules are written.</p>

<p>Loyalty makes this even more complicated. Loyalty is not only patriotism. It is language, family, memory, status, obligation, fear, pride, resentment, gratitude, and belonging. States cultivate loyalty because loyal citizens complain differently and leave later.</p>

<p>This is why political language is always moral language. The state rarely says it wants more control. It says national security, fairness, public interest, farmers, children, consumers, inflation, sovereignty, dignity, or social harmony. These words may describe real problems. They also translate private loss into public necessity.</p>

<h2 id="the-practical-lesson">The Practical Lesson</h2>

<p>For investors, the lesson is to model political permission, not just cash flows.</p>

<p>Before investing, ask whether the state can reinterpret the past, reopen the tax treatment, cancel or reprice a license, subordinate a contract to public interest, ban exports, restrict imports, cap profits, delegitimize the industry, treat the platform as national security infrastructure, trap deposits, freeze reserves, or close the exit door.</p>

<p>The more a business depends on state permission, scarce resources, public legitimacy, regulated prices, crisis profits, or immovable assets, the more politics belongs inside the valuation.</p>

<p>For entrepreneurs, the lesson is similar. Do not ask only whether a market is large. Ask why the profit pool is politically allowed. Every startup is downstream of a political settlement. Some settlements are robust. Others are brittle. Some are simply waiting for their first real confrontation with the state.</p>

<p>For citizens, the lesson is simpler still. Disgust is not independence.</p>

<p>Many intelligent people avoid politics because it feels stupid, corrupt, noisy, tribal, and low status. They are often right about the ugliness. They are wrong about the escape. Withdrawing attention does not withdraw exposure.</p>

<p>Roads are political. Taxes are political. Passports are political. School fees are political. Electricity prices are political. Housing costs are political. City air is political. Bank accounts are political. Investment returns are political. Children’s opportunities are political.</p>

<p>Politics is not everything, but the important things eventually become political.</p>

<p>Money becomes political in a crisis. Food becomes political during inflation. Energy becomes political during war. Education becomes political when parents panic. Healthcare becomes political when costs explode. Housing becomes political when young people cannot buy homes. Platforms become political when they shape speech. AI becomes political when it threatens jobs, truth, and power. Capital becomes political when it tries to leave. Citizenship becomes political when fear redraws the circle of belonging. Law becomes political when the state needs a different answer.</p>

<p>The phrase “I don’t care about politics” is therefore not sophistication. It is usually unpriced exposure.</p>

<p>You may not care about politics, but politics cares about your money, your contracts, your licenses, your bank deposits, your business model, your passport, your speech, your children, your exit options, and your future.</p>

<p>You may own assets. You may even own multiple passports, foreign securities, land, companies, and contracts.</p>

<p>But the state owns the rulebook, and politics is how the rulebook changes.</p>]]></content><author><name>Yash Tambawala</name></author><summary type="html"><![CDATA[Why markets, money, contracts, and even private life are never really outside politics]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://yashtambawala.com/assets/og/you-may-not-care-about-politics-but-politics-cares-about-you.png" /><media:content medium="image" url="https://yashtambawala.com/assets/og/you-may-not-care-about-politics-but-politics-cares-about-you.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">The Country Is Its People</title><link href="https://yashtambawala.com/2026/05/17/country-is-its-people.html" rel="alternate" type="text/html" title="The Country Is Its People" /><published>2026-05-17T00:00:00+05:30</published><updated>2026-05-17T00:00:00+05:30</updated><id>https://yashtambawala.com/2026/05/17/country-is-its-people</id><content type="html" xml:base="https://yashtambawala.com/2026/05/17/country-is-its-people.html"><![CDATA[<p>A country prospers when the people who rule it believe their own future depends on the capability of its people.</p>

<p>This, to me, is the center of the matter.</p>

<p>For a long time, I thought the important questions were about systems. Democracy or authoritarianism? Capitalism or socialism? Stronger markets or a stronger state? Better leaders, better voters, better culture, better institutions?</p>

<p>All of these things matter. But I no longer think they reach the deepest layer.</p>

<p>The deeper question is simpler: does the governing class look at the population and see a vote bank, a labour pool, a market, a burden, a caste equation, or a security problem? Or does it see the population as the country’s most important compounding asset?</p>

<p>When a governing class sees people as an asset, everything changes.</p>

<p>Nutrition starts to look like national security, not welfare. Schools become more than buildings and exams; they become capital formation. Public health becomes productivity policy. Cities stop acting like real estate machines and become engines of human potential.</p>

<p>Governance, then, moves beyond the art of winning elections, managing coalitions, announcing schemes, and surviving the news cycle.</p>

<p>It becomes stewardship. The job of the state becomes morally clear: prevent human damage and compound human potential.</p>

<h2 id="what-unity-really-means">What Unity Really Means</h2>

<p>As a child, I was told that India’s central weakness was disunity. For years, this sounded like a vague moral complaint, something about patriotism, caste, language, religion, or region.</p>

<p>I understand it differently now. India can coordinate. In many ways, it coordinates better than ever. We have national payment rails, national identity systems, national highways, national telecom networks, and increasingly national aspirations.</p>

<p>The question is what we coordinate around.</p>

<p>We have built strong systems to count people, identify them, pay them, sell to them, entertain them, and distribute benefits to them. We have been much less serious about building systems that develop them.</p>

<p>Nutrition is uneven. Schooling is unreliable. Public health is fragile. Women’s work is constrained. Cities drain human energy instead of multiplying it. Manufacturing does not absorb enough workers.</p>

<p>This is the real fragmentation. India has national systems. The trouble is that our strongest systems are not organized around the production of human capability.</p>

<p>We can produce islands of excellence: world-class engineers, doctors, founders, scientists. Brilliant individuals who succeed anywhere. But a thin layer of excellence sitting on top of an underdeveloped base is stratification, not development.</p>

<h2 id="survival-is-not-capability">Survival Is Not Capability</h2>

<p>For a while, I thought the answer might be food. Maybe protein. Maybe South Asia’s cereal-heavy diet: rice, wheat, dal, roti, tea, biscuits, sugar. Maybe we had avoided starvation without ever being fully built.</p>

<p>But that answer was too narrow.</p>

<p>A rich vegetarian Indian child and a poor vegetarian Indian child are not living in the same biological world. One gets milk, curd, paneer, dal, fruit, nuts, vaccines, clean water, doctors, books, conversation, preschool, safety, and parental attention. The other may get rice, wheat, watery dal, tea, repeated infections, anemia, poor sanitation, low stimulation, and weak schooling.</p>

<p>Both are “vegetarian.” But one childhood is developmental. The other is survival.</p>

<p>That distinction changed the frame for me.</p>

<p>South Asia did not fail at producing food. It failed at producing fully developed children. Cereal agriculture was not a mistake; it was a civilizational answer to an ancient problem. How do you keep enormous populations alive under monsoon risk, disease, weak storage, and fragmented land? You grow grains. Storable, countable, divisible, calorie-dense grains. Grain keeps people alive.</p>

<p>But survival and development are different achievements. Confusing the two is where the developmental imagination falls short.</p>

<p>A grain state can be simple and still function. It can procure, store, ration, count, and announce. It can say: no one should sleep hungry. For a poor and densely populated society, that is an achievement.</p>

<p>But a nutrition state has to be far more capable. It has to ensure milk is safe, eggs are fresh, mothers are counselled, children are weighed, anemia is treated, water is clean, school meals contain real nutrition, parents talk to children, and infections are prevented.</p>

<p>That requires a different order of capability.</p>

<p>Calories are low-coordination. Human capital is high-coordination.</p>

<p>This distinction clarifies a lot. India built enough state capacity to prevent famine, distribute grain, run elections, count people, police disorder, and announce schemes. Building human beings asks for something harder: daily reliability, invisible quality, institutions that function instead of merely existing, and trust.</p>

<h2 id="trust-is-productive-infrastructure">Trust Is Productive Infrastructure</h2>

<p>The more I think about development, the more trust looks like productive infrastructure rather than just a moral virtue.</p>

<p>You need trust to buy milk. To eat outside. To take medicine. To send your child to school. To hire strangers. To build factories. To certify quality. To export.</p>

<p>You need trust to move from family businesses to modern firms. From grain to nutrition. From exams to learning. From jugaad to systems.</p>

<p>Without trust, people retreat to what they can see and control. Grain. Gold. Land. Cash. Family. Caste. Community. Political patrons. This retreat is rational in a world where institutions cannot be trusted. Modern prosperity, however, is complexity. A developed country is not only wealthy; it is able to coordinate at scale.</p>

<p>India has islands of this in technology, pharmaceuticals, finance, space, digital infrastructure, and some manufacturing. But these islands sit within a much larger landscape of low trust, uneven nutrition, unreliable schooling, informal labour, unsafe cities, and fragmented coordination.</p>

<p>Services alone cannot complete India’s development. They can create elite prosperity: salaries, foreign exchange, urban consumption, global status. But they do not automatically transform a society where mass capability remains underdeveloped.</p>

<p>Manufacturing matters because it creates jobs, but also because it teaches coordination. Factories require punctuality, quality control, process discipline, supplier reliability, documentation, and export accountability. Manufacturing trains a society to cooperate with strangers at scale.</p>

<p>India still needs more than GDP. It needs the social capacity to coordinate.</p>

<h2 id="why-the-equilibrium-persists">Why the Equilibrium Persists</h2>

<p>What makes this hard is that none of it is hidden.</p>

<p>Everyone can see pieces of it. We know too many children are not learning. We know nutrition is inadequate. We know cities drain people. We know our best people often serve foreign demand. We know Indian businesses pay foreign platforms to reach Indian consumers. Even our attention is mediated from abroad.</p>

<p>A country of 1.4 billion people should feel uneasy about this.</p>

<p>And yet the equilibrium persists.</p>

<p>Not because everyone is malicious, or because someone designed the entire system in a closed room. The equilibrium persists because everyone is optimizing locally.</p>

<p>The engineer takes the global job. The worker goes to Dubai. The founder buys ads on Google and Meta. The politician announces visible schemes. The family sends its child abroad. The business elite diversifies geographically. The bureaucrat manages the file. The poor want survival. The rich want exit. The middle class wants security.</p>

<p>Everyone is making rational choices inside an underbuilt system.</p>

<p>The system remains underbuilt because enough powerful people can succeed without making the average Indian more capable. Too many can win by leaving. Too many businesses can win by importing. Too many politicians can win by distributing. Too many platforms can win by extracting attention. Too many elites can prosper by positioning near foreign capital, foreign demand, or foreign institutions.</p>

<p>When the elite can escape the consequences of low national capability, the country remains underbuilt.</p>

<p>This is why I no longer think the main question is whether India is democratic or authoritarian. The real question is: who has skin in the game?</p>

<h2 id="skin-in-the-game">Skin in the Game</h2>

<p>Do India’s rulers personally rise when the average Indian child becomes healthier, smarter, safer, and more productive? Or do they rise when the population remains dependent, fragmented, and easy to manage?</p>

<p>That question sits beneath all politics.</p>

<p>If rulers depend on the long-term capability of the people, they will invest in the people. If they depend on managing coalitions, distributing benefits, controlling narratives, and arbitraging global systems, they will do that instead.</p>

<p>The form of government matters. The incentive of the governing class matters more.</p>

<p>Any system can waste human capital when its rulers can survive without developing the population. Any system can build human capital when its rulers are forced to depend on the population’s capability. The form matters, but the soul of the system is skin in the game.</p>

<p>A ruler has skin in the game when his own power depends on the capability of the population. A business elite has skin in the game when its wealth depends on building domestic ecosystems, rather than arbitraging cheap labour and protected access. A cultural elite has skin in the game when prestige comes from building here, rather than proximity to elsewhere. A society has skin in the game when every damaged childhood is treated as a national loss.</p>

<p>China is uncomfortable to think about for precisely this reason. Not because India should become China. It should not. The narrower lesson is that China treated mass capability as a state project. Grain security was not the end of development; it was the base layer. From there came industrialization, foreign exchange, imported inputs, better diets, trained workers, supply chains, and coordination at scale.</p>

<p>India built the floor. It still has to build the ladder.</p>

<p>Food security became a welfare achievement. It did not become the first layer of a developmental state. A dense civilization cannot eat its way to development from its own land. It has to coordinate, trade, industrialize, and trust. It has to convert food into health, health into skills, skills into firms, firms into exports, and exports into national power.</p>

<p>We keep breaking that sequence.</p>

<h2 id="the-real-exam-starts-before-birth">The Real Exam Starts Before Birth</h2>

<p>India is obsessed with exams. But by the time a child enters the exam pipeline, much of the inequality has already entered the body and brain.</p>

<p>Was the mother nourished? Was the baby born healthy? Was the child protected from infections? Was there conversation at home? Was there preschool, safety, reading, play, dignity?</p>

<p>India pretends this is a meritocracy problem. It is actually a developmental timing problem.</p>

<p>The most important education policy may happen before school. The most important productivity policy may be preventing anemia in adolescent girls. The most important industrial policy may be maternal nutrition. The most important cognitive intervention may be breastfeeding, complementary feeding, sanitation, and responsive caregiving in the first thousand days.</p>

<p>This sounds strange only because our imagination of development starts too late. We think it begins with college, jobs, startups, highways, factories, exports. But development begins in the womb.</p>

<p>A serious state would understand this as national power, not charity, “women and child welfare,” or NGO language.</p>

<p>Because what is a country? Surely not land alone. Not GDP alone. Not only a flag, an army, or a constitution.</p>

<p>A country is accumulated human capability, organized across generations.</p>

<h2 id="the-next-meaning-of-freedom">The Next Meaning of Freedom</h2>

<p>This changes how I think about freedom.</p>

<p>The independence generation thought freedom meant political sovereignty. The planning generation thought it meant public sector command and grain security. The liberalization generation thought it meant markets, consumption, and global integration.</p>

<p>The next definition has to go deeper.</p>

<p>Freedom means every child gets the developmental inputs required to become fully capable, instead of only being kept alive, enrolled, fed grain, given a vote, or given a subsidy.</p>

<p>That remains the unfinished freedom struggle, and it is much harder than the old one.</p>

<p>Political freedom can be declared. Human capability has to be built every day through nutrition, sanitation, schools, cities, factories, women’s work, honest certification, trustworthy food, safe transport, and institutions that do not lie.</p>

<p>It has to be built by rulers who cannot escape the consequences of neglect.</p>

<p>This is why governance should be judged by stewardship, not rule. Health policy should be judged by capability, not hospitals alone. Education policy should be judged by capability, not enrollment. Food policy should be judged by capability, not grain distribution.</p>

<p>Everything must answer one question: did it make the people more capable?</p>

<h2 id="the-only-real-test">The Only Real Test</h2>

<p>Countries rise when their elites have to treat the people as the source of their own future. Countries stagnate when elites can prosper without developing the people.</p>

<p>This is the challenge India faces.</p>

<p>Too many people can succeed by leaving. Too many businesses can succeed by importing. Too many politicians can succeed by distributing. Too many elites can succeed without the average Indian becoming more capable.</p>

<p>Slogans will not break that equilibrium.</p>

<p>It will break only when capability becomes the country’s central political, economic, and moral project. We will have to become serious about the child before the exam, the mother before the child, nutrition before productivity, trust before scale, manufacturing before slogans, and domestic systems before global status.</p>

<p>Serious about the population not as a burden, a market, a vote bank, or cheap labour, but as the country’s deepest source of strength.</p>

<p>A large population does not make a country prosperous.</p>

<p>A capable population does.</p>

<p>India’s task is larger than growth. It is to build the conditions in which every Indian can become fully capable.</p>

<p>Or, said more plainly: India’s task is to stop wasting Indians.</p>]]></content><author><name>Yash Tambawala</name></author><summary type="html"><![CDATA[Why India's deepest project is not growth, but capability]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://yashtambawala.com/assets/og/country-is-its-people.png" /><media:content medium="image" url="https://yashtambawala.com/assets/og/country-is-its-people.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">The Marginal Price Setter in Indian Real Estate</title><link href="https://yashtambawala.com/2026/04/08/marginal-price-setter.html" rel="alternate" type="text/html" title="The Marginal Price Setter in Indian Real Estate" /><published>2026-04-08T00:00:00+05:30</published><updated>2026-04-08T00:00:00+05:30</updated><id>https://yashtambawala.com/2026/04/08/marginal-price-setter</id><content type="html" xml:base="https://yashtambawala.com/2026/04/08/marginal-price-setter.html"><![CDATA[<p>I increasingly think India’s urban growth model is being shaped by a mechanism that most people do not notice, even though it quietly determines how capital gets allocated across the economy. That mechanism is the marginal price setter in real estate.</p>

<p>At first glance, real estate prices appear to be a story of broad supply and demand. People assume prices rise because cities are growing, land is limited, or India is urbanizing. All of that is true at a surface level, but it misses the more important point. In a supply-constrained market, the relevant question is not what the average household can afford. The relevant question is: who is the buyer at the margin, the one who actually sets the clearing price? Once that buyer transacts, the entire market begins to re-anchor around that price, whether or not the median family can pay it.</p>

<p>That is why I keep coming back to the idea of the marginal price setter. It offers a cleaner lens for understanding not just why real estate becomes expensive, but why an entire economy begins to organize itself around asset inflation rather than production. In India’s case, I increasingly suspect that global income flows enter the country and then get capitalized into urban real estate, where they reshape the incentives of households, developers, banks, and politicians alike.</p>

<h2 id="where-the-idea-comes-from">Where the Idea Comes From</h2>

<p>The idea comes from the marginalist tradition in economics, developed in the late nineteenth century by thinkers like Jevons, Menger, and Walras. Their core insight was simple but profound: prices are not determined by the average participant in a market. They are determined at the margin, where the last unit demanded meets the last unit supplied. In other words, the decisive transaction is not the typical one. It is the transaction that clears the market.</p>

<p>This sounds abstract until you see how universal it is. I understood it especially well because electricity markets were part of my undergraduate studies. In power markets, this logic is explicit. Generators bid power into the market at different prices depending on their costs. The system operator stacks those bids from cheapest to most expensive and dispatches generation until demand is met. The last plant needed to satisfy demand, the one at the margin, sets the market-clearing price for all dispatched power. So even if most of the electricity comes from cheaper plants, the price everyone receives is often set by the cost of the last and most expensive plant required to balance the system.</p>

<p>That is the marginal price setter in action. It is not the average cost of all generators that matters. It is the marginal generator. Once you grasp that, you begin to see the same structure everywhere. In commodity markets, the marginal producer often sets the price. In labor markets, the marginal employer can influence wages in tight segments. In financial markets, the marginal buyer or seller determines valuation. And in urban real estate, especially in thin, supply-constrained markets, the same principle can become incredibly powerful.</p>

<h2 id="why-real-estate-is-so-sensitive-to-marginal-pricing">Why Real Estate Is So Sensitive to Marginal Pricing</h2>

<p>Real estate is one of the purest examples of marginal pricing because it combines two characteristics that make the marginal transaction disproportionately influential. First, supply is slow and constrained. Second, transactions are sparse.</p>

<p>Supply in real estate cannot respond quickly to demand shocks. Zoning restrictions, floor space limits, approval delays, infrastructure constraints, land assembly problems, financing bottlenecks, and construction timelines all mean that when demand rises, the market cannot simply produce new supply at the same speed. In a manufacturing industry, higher prices can sometimes attract new capacity relatively quickly. In urban property, that process is much slower, politically more contested, and physically more constrained.</p>

<p>At the same time, real estate markets are not deep, continuously traded markets like equities. Only a small fraction of properties transact in a given year. That means each transaction carries an outsized role in price discovery. When one apartment in a locality sells at a higher price, that sale quickly becomes the benchmark for brokers, developers, appraisers, banks, lenders, and neighboring sellers. The market does not need thousands of trades to move. A handful of marginal transactions can re-anchor expectations for an entire neighborhood.</p>

<p>This is why real estate feels so strange to people. They look around and think, “No one I know can afford these prices.” But that is exactly the point. The market price does not need to be set by the people they know or by the median household. It only needs to be set by the relatively small group of buyers who are actually transacting at the top of the demand stack. In a thin market, the marginal buyer matters far more than the representative buyer.</p>

<h2 id="who-the-marginal-price-setter-is-in-indian-cities">Who the Marginal Price Setter Is in Indian Cities</h2>

<p>In many Indian cities today, the marginal price setter is not the median salaried household. It is often a narrower class of globally connected urban earners whose incomes are linked, directly or indirectly, to external flows. This includes people benefiting from IT and services exports, remittances, stock market gains, startup exits, venture-funded compensation, global capital inflows, and other forms of internationally linked income.</p>

<p>The key point is not that these groups are numerically dominant. They do not need to be. They only need to be strong enough at the margin to transact at price points above what the median household can support. Once they do that, the market starts taking those transactions as the benchmark.</p>

<p>So global income enters India, but instead of being broadly redirected into industrial capacity, a meaningful chunk of it gets capitalized into urban real estate. That is the structural move that matters. The inflow may originate in software exports, foreign capital, remittance corridors, or bull markets in financial assets, but its local expression often shows up in land and housing prices.</p>

<p>Once this happens repeatedly, urban real estate stops reflecting local productive capacity alone. It begins to reflect the purchasing power of a globally connected buyer at the margin. Once that buyer sets the benchmark, everyone else in the city must orient themselves around a price structure they did not create and often cannot afford.</p>

<h2 id="from-housing-market-to-savings-sink">From Housing Market to Savings Sink</h2>

<p>The consequences go much further than affordability. Once real estate becomes the preferred destination for excess savings, it begins to change how the economy allocates capital.</p>

<p>Households start treating apartments, plots, and second homes not primarily as places to live, but as stores of value. Real estate becomes the default answer to uncertainty. It is seen as safer than entrepreneurship, more understandable than equities, and more culturally legitimate than financial assets. Families begin to save for property, borrow against property, compare status through property, and measure security through property ownership.</p>

<p>This matters because the economy’s savings are finite. Money that is capitalized into existing land and apartment values is money that is not being directed into manufacturing capacity, industrial technology, logistics systems, export capability, or productive enterprise. When enough savings get absorbed into property, the economy gradually becomes asset-driven rather than production-driven.</p>

<p>This is one of the deeper meanings of financialization. It is not merely that people speculate. It is that rising asset prices themselves become the organizing logic of the system. Households pursue asset gains, banks lend against assets, governments depend on asset-linked revenues, and political coalitions form around preserving those valuations.</p>

<p>In that world, real estate is no longer just another sector. It becomes the gravitational center of the economic model.</p>

<h2 id="why-policy-reinforces-the-dynamic">Why Policy Reinforces the Dynamic</h2>

<p>In India, this process has not occurred in a vacuum. Policy has often amplified it.</p>

<p>Real estate enjoys a privileged place in the economic and political system. Tax provisions such as Section 54 have historically made it easier to roll gains from one property into another. Stamp duties make property transactions lucrative for state governments. Approval systems remain cumbersome enough to keep supply constrained. Development politics often revolves less around abundant housing and more around preserving or enhancing land values. In many cases, the state itself becomes a participant in scarcity rather than a neutral referee trying to lower costs.</p>

<p>All of this strengthens the attraction of real estate as a savings vehicle. At the same time, it prevents supply from adjusting smoothly enough to dissipate those price pressures. The result is a market where global and domestic capital can keep pushing into a structurally constrained asset base.</p>

<p>Once that happens, property becomes both a financial asset and a political asset. Rising prices please developers, lenders, asset-owning households, and political actors with direct or indirect links to land. A large coalition emerges that benefits from price appreciation, or at least fears the consequences of falling prices. That is why the system becomes hard to reform. It is not just an economic equilibrium. It is a political one.</p>

<h2 id="the-electricity-market-analogy-properly-understood">The Electricity Market Analogy, Properly Understood</h2>

<p>The analogy with electricity markets is useful here because it clarifies the mechanics with unusual precision.</p>

<p>In power exchanges, the clearing price is determined by the most expensive dispatched unit needed to meet demand. Suppose demand is high in the evening and the grid needs to call upon a gas peaker or a costly imported coal plant to satisfy the last slice of demand. That plant becomes the marginal price setter. The entire market price gets pulled up to that level, even though much of the power may have been supplied by cheaper coal, hydro, or renewables.</p>

<p>Real estate behaves similarly, but with its own distortions. In urban property, the expensive buyer at the margin acts like the costly generator in an electricity market. That buyer is not average, but because the market is supply-constrained and transaction-light, their willingness to pay becomes the benchmark for the entire local price structure. Developers launch projects based on that benchmark. Banks underwrite loans around it. Existing owners refuse to sell below it. Even households far below that income band are forced to navigate a market whose reference point is set by someone else.</p>

<p>The difference, of course, is that electricity markets are explicitly designed clearing systems, while real estate markets are socially and politically messy. But the underlying logic is the same. The marginal transactor, not the average participant, determines the price signal that the rest of the system must respond to.</p>

<p>And that is why the identity of the marginal buyer matters so much. If the marginal price setter is an industrial user who needs affordable land to produce and export, the city evolves differently. If the marginal price setter is a globally linked financial buyer seeking a store of value, the city evolves into a savings sink.</p>

<h2 id="tokyo-and-the-dangers-of-a-marginally-anchored-bubble">Tokyo and the Dangers of a Marginally Anchored Bubble</h2>

<p>A powerful historical example of this dynamic can be seen in Tokyo during the late 1980s. After the Plaza Accord in 1985, the yen appreciated sharply, and Japan responded with easier monetary conditions. Credit expanded, banks lent aggressively against real estate collateral, and property became the center of speculative enthusiasm.</p>

<p>Tokyo’s land market was already supply constrained and transaction-light. That meant it did not take a mass movement of ordinary households to push prices upward. A relatively narrow set of corporate buyers, investors, and financial actors could become the marginal price setters. Once they started transacting at elevated prices, the entire market repriced upward. Rising valuations increased collateral values, which allowed for more lending, which funded even more property purchases. It was a classic reflexive loop.</p>

<p>The famous claim that the land beneath the Imperial Palace was worth more than all the real estate in California captured the absurdity of the peak, even if such comparisons were partly rhetorical. The deeper truth was that national savings and credit creation had become increasingly absorbed by real estate, rather than by productive deployment.</p>

<p>Then conditions changed. When the Bank of Japan tightened policy, the credit impulse weakened. The speculative marginal buyers disappeared. Once the marginal price setter vanished, the benchmark price structure could no longer hold. Property prices fell dramatically over time, banks were damaged, balance sheets deteriorated, and Japan entered the long aftermath that later came to be described as a balance sheet recession.</p>

<p>The lesson is brutal and simple. When a market is anchored by a marginal buyer whose purchasing power is contingent rather than fundamental, the entire valuation structure can prove fragile.</p>

<h2 id="why-this-makes-india-fragile">Why This Makes India Fragile</h2>

<p>India’s situation is not identical to Tokyo’s, but the structural vulnerability is real. If the marginal buyer in Indian urban property depends heavily on externally linked income streams, then the resilience of the entire price structure depends on the durability of those flows.</p>

<p>If IT export demand weakens, remittances slow, foreign capital exits, startup liquidity dries up, or stock-market-linked wealth contracts, then the purchasing power of the marginal buyer starts to erode. Because urban real estate prices were never set by the median Indian household in the first place, the system may be more fragile than it appears. It can look socially unaffordable for years and still remain elevated, so long as the marginal buyer remains funded. But if that buyer weakens, the benchmark itself becomes unstable.</p>

<p>This also has political implications. Many political and economic elites are themselves exposed to real estate, directly or indirectly. Their wealth, financing structures, and local influence are often entangled with land values. So a property slowdown does not remain confined to the housing market. It starts affecting state revenues, development incentives, local patronage networks, bank collateral quality, and the broader political economy.</p>

<p>That is why this is not just a story about expensive apartments. It is a story about how a country’s growth model can become dependent on who is able to set prices at the urban margin.</p>

<h2 id="the-gujarat-counterexample-and-the-production-oriented-city">The Gujarat Counterexample and the Production-Oriented City</h2>

<p>This is also why some places in India feel structurally different.</p>

<p>Where cities cannot rely as heavily on globally driven white-collar income to sustain land values, they are often forced into a more production-oriented equilibrium. Land has to function more as an input into economic activity and less as a pure speculative savings sink. That tends to push policy toward town planning, land pooling, trunk infrastructure, and supply expansion.</p>

<p>This is one reason Gujarat matters as a counterexample. Gujarat’s major cities, whatever their many flaws, often had to be more serious about enabling production, logistics, and industrial land use. They could not depend in the same way on a global-income-fueled urban real estate model to carry the local economy. So the system had stronger incentives to make land usable, tradable, and serviceable for actual economic output.</p>

<p>That does not mean Gujarat somehow escaped land politics. Of course not. But the orientation is meaningfully different. When a region must rely more on manufacturing, trade, and domestic commerce, urban planning becomes less optional. Land cannot remain only a speculative object. It has to work as economic infrastructure.</p>

<p>That is the alternative path India could have built more broadly: cities where land is abundant enough and well planned enough to support production, rather than cities where property functions primarily as the preferred vessel for capitalized global income.</p>

<h2 id="the-deeper-point">The Deeper Point</h2>

<p>The deeper point is not merely that real estate prices are high. It is that the identity of the marginal price setter determines what kind of economy gets built.</p>

<p>If the marginal price setter is financially driven, globally connected, and primarily interested in property as a store of wealth, then savings will keep getting pulled into land and housing. Asset appreciation becomes central. Production becomes secondary. Politics organizes itself around defending valuations.</p>

<p>If the marginal price setter is instead an industrial actor who needs affordable land, reliable infrastructure, and scalable urban systems in order to produce, employ, and export, then the economy starts looking very different. Capital goes into factories, logistics, tooling, warehousing, and productive coordination. Land behaves less like a speculative chip and more like an economic input.</p>

<p>That is why “real estate” is not a side story in development. It is central. The urban land market quietly decides whether savings go toward building the productive base or toward bidding up the price of existing assets.</p>

<p>And that is also why the term marginal price setter is so useful. It forces us to stop looking at averages and start looking at the decisive transaction. Once you do that, a lot of India’s urban political economy begins to make more sense. Global income flows enter the system, but instead of transforming industrial capacity, they often end up hardening an asset-led equilibrium through the urban property market. The price is set at the margin, and then the rest of the economy has to live with the consequences.</p>]]></content><author><name>Yash Tambawala</name></author><summary type="html"><![CDATA[How Global Income Flows Shape India’s Real Estate Economy]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://yashtambawala.com/assets/og/marginal-price-setter.png" /><media:content medium="image" url="https://yashtambawala.com/assets/og/marginal-price-setter.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">The Agency Problem</title><link href="https://yashtambawala.com/2026/03/09/the-agency-problem.html" rel="alternate" type="text/html" title="The Agency Problem" /><published>2026-03-09T00:00:00+05:30</published><updated>2026-03-09T00:00:00+05:30</updated><id>https://yashtambawala.com/2026/03/09/the-agency-problem</id><content type="html" xml:base="https://yashtambawala.com/2026/03/09/the-agency-problem.html"><![CDATA[<p>The world belongs to those who shape it.</p>

<p>There is a line from the poem <em>Invictus</em> that captures the spirit of agency better than almost anything else.</p>

<blockquote>
  <p>I am the master of my fate.<br />
I am the captain of my soul.</p>
</blockquote>

<p>On the surface it sounds like a statement about individual character. A person deciding to take responsibility for their own life.</p>

<p>But the idea behind it is much bigger than the individual. Entire societies display different levels of agency. Some nations behave like authors of the systems they live in. Others behave more like navigators, learning how to move through systems that feel as if they were created somewhere else.</p>

<p>In that sense the real question for any society is very simple: Do citizens feel like they are shaping the system around them, or do they feel like they are merely surviving inside it?</p>

<p>That psychological difference ends up influencing almost everything else.</p>

<h2 id="agency-is-not-just-a-personality-trait">Agency Is Not Just a Personality Trait</h2>

<p>We often talk about agency as if it were purely a personality trait. Some people are disciplined, ambitious, proactive. Others appear passive, reactive, or fatalistic.</p>

<p>But human behavior is far more sensitive to environment than we usually admit.</p>

<p>It helps to separate two distinct types of agency:</p>

<ul>
  <li>Individual agency: the internal capacity for initiative, discipline, and long-term thinking.</li>
  <li>Environmental agency: the degree to which a system actually allows individuals to influence outcomes.</li>
</ul>

<p>A high-agency individual inside a system that rewards initiative becomes a builder. The same individual placed inside a system where effort rarely translates into results may become cynical, opportunistic, or simply extremely skilled at extracting value from complexity.</p>

<p>The person did not fundamentally change. The environment did.</p>

<p>Consider a single individual: a sharp, ambitious engineer working across different contexts in the same week. Inside a well-run startup, she ships quickly, takes ownership, builds for the long term. Negotiating a government permit, she learns to read the right people, move indirectly, trade in favors. Driving abroad, she follows every rule with care. Back home, she treats a red light as optional at midnight.</p>

<p>Same person. Four different behavioral equilibria. What changed was not her character. What changed was what the system rewarded.</p>

<h2 id="discipline-usually-follows-agency">Discipline Usually Follows Agency</h2>

<p>This is why discipline is frequently misunderstood. It is often treated as a moral virtue, something societies either possess or lack as a matter of character.</p>

<p>But discipline usually appears as a consequence of agency rather than its cause.</p>

<p>Think about professions where feedback is immediate and consequences are real. Pilots follow checklists with ritual precision. Surgeons rehearse procedures until every motion becomes second nature. Engineers building spacecraft operate in worlds where tiny mistakes destroy entire missions. These environments reward structure. Discipline emerges because it works.</p>

<p>But when outcomes are loosely tied to effort and systems behave unpredictably, rigid discipline becomes less valuable than flexibility. Improvisation beats procedure. Adaptability beats structure. In India there is even a word that captures this mindset: <em>jugaad</em>.</p>

<p>This does not mean people lack the capacity for discipline. It means disciplined authorship has not always been the rational strategy.</p>

<h2 id="birds-that-look-free-are-not-undisciplined">Birds That Look Free Are Not Undisciplined</h2>

<p>Watch a murmuration of starlings moving across an evening sky and the first impression is of pure freedom. Thousands of birds wheeling and turning with fluid spontaneity, no apparent leader, no visible plan.</p>

<p>But look closer and what you see is extreme discipline operating at speed. Each bird follows precise rules about its distance and alignment with its neighbors. The vast, shifting shape that emerges is not despite the individual discipline. It is because of it.</p>

<p>Or consider migratory birds crossing thousands of kilometers, orienting by magnetic fields, maintaining formation through storms, navigating without instruments. The freedom to cross continents is built entirely on relentless internal structure.</p>

<p>The lesson is counterintuitive: genuine freedom at the collective level requires genuine discipline at the individual level. What looks like spontaneous flight is the output of a very precise system.</p>

<p>High-agency societies work the same way. The freedom to build, innovate, and create does not replace structure. It emerges from it. A civilization that mistakes undisciplined wandering for freedom will eventually discover it has simply drifted.</p>

<h2 id="a-nation-is-an-archipelago-of-trust-environments">A Nation Is an Archipelago of Trust Environments</h2>

<p>One of the simplest ways to observe how environment shapes behavior is through something mundane.</p>

<p>Consider the difference between Indian streets and Indian metro stations. On the street it is not unusual to see people casually litter. Yet the very same individuals often follow rules carefully inside a metro system. They stand in queues. They avoid throwing garbage. They behave with what looks like almost European discipline.</p>

<p>What changed? The individual did not become a different person. The environment changed. Inside the metro the rules are clear, enforcement is visible, and everyone else appears to be complying. Under those conditions disciplined behavior becomes the rational equilibrium.</p>

<p>Once you start noticing this pattern, you see it everywhere. Indians who casually break traffic rules at home drive carefully abroad. People who litter public spaces maintain immaculate homes. Citizens who distrust government institutions operate with remarkable honesty inside well-run companies.</p>

<p>A nation is therefore not a single moral landscape. It is an archipelago of different trust environments, each with its own equilibrium. The real question is never whether high-trust islands exist. They always do. The real question is whether those islands expand into continents.</p>

<h2 id="the-low-agency-equilibrium">The Low Agency Equilibrium</h2>

<p>History plays a role in shaping these patterns. Large parts of South Asia spent centuries under layered systems of authority, imperial courts, caste hierarchies, colonial administration, and later large bureaucratic states. Across many of these systems the recurring lesson was similar: important decisions came from somewhere else.</p>

<p>Over time people internalize this experience. They stop assuming that systems are theirs to shape. Instead they become extremely skilled at surviving within them. Adaptation becomes more valuable than transformation. Social intelligence becomes more important than institutional audacity.</p>

<p>This does not produce unintelligent societies. In fact it often produces extremely clever individuals. But their cleverness becomes tactical rather than civilizational. They navigate history rather than write it.</p>

<h2 id="ownership-and-corruption">Ownership and Corruption</h2>

<p>This framework also provides a deeper explanation for corruption.</p>

<p>Corruption is usually framed in moral terms. People are greedy. Institutions are weak. Enforcement is insufficient. But underneath those explanations lies a simpler psychological principle: people protect systems they feel they own.</p>

<p>When institutions feel distant, extractive, or temporary, the mind quietly shifts: the system is not mine. Once that shift occurs, extracting value from the system begins to feel rational rather than immoral.</p>

<p>This is why corruption so often involves capable individuals rather than incompetent ones. High agency does not disappear in low-ownership environments. It simply changes form. It becomes strategic opportunism.</p>

<h2 id="how-equilibria-actually-change-and-why-it-often-comes-from-outside">How Equilibria Actually Change And Why It Often Comes From Outside</h2>

<p>If environment determines behavior, the critical question is: what actually changes the environment?</p>

<p>The honest answer is uncomfortable. High-agency equilibria are almost never produced by broad organic consensus. They are seeded by a small number of actors who absorb a disproportionate share of what might be called the coordination cost.</p>

<p>Coordination cost is the friction of getting many people to behave differently all at once. In a low-trust environment, acting with institutional integrity before others do is individually irrational. The first movers bear the cost while free-riders capture the benefit. Most people therefore wait. The low-trust equilibrium is self-reinforcing.</p>

<p>A shift happens when someone, or a small founding group, is willing to absorb that cost anyway. To build a system that rewards effort before the surrounding environment confirms it. To hold others to standards that do not yet feel natural. To create conditions under which disciplined behavior becomes rational for everyone else.</p>

<p>But there is a harder implication that follows directly from this: in deeply entrenched low-agency environments, internal actors rarely absorb that cost voluntarily. The equilibrium punishes them for trying. And so, more often than history is comfortable admitting, the initial disruption comes from outside.</p>

<p>Colonial powers, foreign capital, international institutions, multinational corporations, all have functioned as exogenous shocks to local equilibria. The British did not build Indian railways for India’s benefit. But the railways still happened. The infrastructure of the modern Indian state, its legal system, its bureaucratic architecture, its universities, much of it was scaffolded by a foreign power extracting value from the subcontinent. The uncomfortable truth is that external imposition and internal liberation are sometimes the same event, described from different vantage points.</p>

<p>This is not an argument for colonialism. It is an argument for clarity about mechanism. Prussia’s administrative reforms were imposed against resistance by a state that decided to make competence matter. Meiji Japan did not drift toward industrialization, a small elite dismantled feudal structures with deliberate speed. Singapore did not become high-trust organically. Lee Kuan Yew changed what was rational for millions of people by changing what would be consistently enforced.</p>

<p>In each case the population did not become more intelligent or more virtuous overnight. The equilibrium changed. And once effort reliably produced results, discipline and institution-building followed as consequences.</p>

<h2 id="swaraj-and-the-psychology-of-ownership">Swaraj and the Psychology of Ownership</h2>

<p>This perspective sheds direct light on Indian history.</p>

<p>Many smaller kingdoms on the subcontinent struggled not because they lacked courage but because they existed inside fragmented political environments where long-term institutional coordination was difficult. In such environments, even capable rulers operated within systems designed by someone else.</p>

<p>What made the Maratha rise distinctive was not just military success. It was the emergence of Swaraj, and a single concentrated actor in Shivaji who was willing to absorb the coordination cost of that idea before it had broad legitimacy. Shivaji did not wait for consensus. He built the institutions, enforced the standards, and created the conditions under which others could rationally join the new equilibrium.</p>

<p>Swaraj was a psychological claim of ownership. It suggested that the system of rule should belong to those who lived within it. Once people begin to feel that a political order is theirs, behavior changes. They sacrifice, coordinate, and build institutions that endure.</p>

<p>In that sense the deepest struggle in political history is rarely over territory. It is over authorship.</p>

<h2 id="the-illusion-of-escape-freedom">The Illusion of Escape Freedom</h2>

<p>There are two very different ideas of freedom operating beneath all of this.</p>

<p>The first is freedom through escape. This is the desire to reduce dependence on systems that feel broken or arbitrary. People accumulate assets, build financial independence, and gradually withdraw from environments they do not trust. Modern movements around financial independence reflect this idea. The goal is to need the system less.</p>

<p>But there is a paradox concealed inside escape freedom that almost no one examines honestly.</p>

<p>The person who optimizes for escape believes they are maximizing autonomy. What they have actually done is substitute dependence on one system, institutions, employers, the state, for dependence on another system that they have even less influence over. The retiree living off a portfolio is entirely at the mercy of markets, inflation, geopolitical stability, and monetary policy she did not shape and cannot affect.</p>

<p>She replaced a relationship where her effort could matter with one where it cannot. The intermediary became impersonal, so the dependence became invisible. But it did not disappear. It deepened.</p>

<p>Compare this with the founder who built the company generating those returns. She has more genuine autonomy not despite being embedded in a system, but because she is part of the system producing value. Her authorship is a source of resilience, not a liability.</p>

<p>Escape freedom is not actually freedom. It is a more sophisticated form of dependency, one that feels like independence precisely because the chain connecting you to the world has grown longer and more abstract.</p>

<p>The second kind of freedom is freedom through authorship: the freedom to design institutions, shape incentives, and build systems that affect many other lives. The first type produces survivors and portfolio managers. The second produces founders, reformers, and institution builders.</p>

<p>Healthy societies contain both. But civilizations dominated entirely by escape freedom eventually stop building systems. They only learn how to protect themselves from broken ones. And a civilization that only knows how to protect itself has already begun to decline.</p>

<h2 id="becoming-the-captains-of-fate">Becoming the Captains of Fate</h2>

<p>The line from <em>Invictus</em> is not really about self-help. It is about civilizational psychology.</p>

<p>A society becomes the master of its fate when enough capable people stop treating institutions as temporary shelters to exploit and begin treating them as structures they are responsible for building. When freedom stops meaning escape and starts meaning authorship. When discipline stops being a burden and starts being the mechanism of flight.</p>

<p>But that transition does not happen by inspiration alone. It requires someone to go first. Someone willing to act as if the new equilibrium already exists. Someone willing to absorb the coordination cost on behalf of everyone who will later benefit from the system they helped establish. And in deeply entrenched low-agency environments, that someone is often an outsider, which is its own uncomfortable lesson about the relationship between disruption and development.</p>

<p>The world belongs to those who shape it.</p>

<p>Not to those who navigate it most cleverly. Not to those who protect themselves from it most effectively. Not even to those who understand it most clearly.</p>

<p>To those who shape it.</p>

<p>That is the real turning point, not when slogans grow louder, not when a different party wins an election, not even when talent increases. But when a small group of people decide to stop waiting for the equilibrium to change and begin absorbing the cost of changing it themselves.</p>

<p>At that point discipline stops feeling like punishment. Corruption stops feeling clever. Escape stops feeling like freedom.</p>

<p>And a civilization that once survived by navigating systems slowly begins to create them.</p>]]></content><author><name>Yash Tambawala</name></author><summary type="html"><![CDATA[Why Some Societies Build History While Others Learn to Navigate It]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://yashtambawala.com/assets/og/the-agency-problem.png" /><media:content medium="image" url="https://yashtambawala.com/assets/og/the-agency-problem.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry></feed>