How the India higher-education explorer gets its numbers
Aggregate analysis of public government statistics — AISHE, AICTE, NMC, UDISE+ and MoSPI’s PLFS microdata, each downloadable from its own publisher. No Avanti student data and no personally identifiable information: every figure is a count, a weighted population share, or a median.
The explorer is unlisted while its population figures are labelled estimates — reachable by URL, absent from search and from every index on this site. This page is the durable version of the caveats the explorer states inline, and it is deliberately unlisted too, so the method and the tool carry the same status.
One question per source
Section titled “One question per source”Five sources, and the explorer never asks one of them a question it cannot answer.
| Question | Source | Why not the others |
|---|---|---|
| How many graduated, by field and level | AISHE | The only annual out-turn by programme. PLFS cannot count awards. |
| What happens to them — work, pay, job type | PLFS | The only source that observes people rather than institutions. |
| Engineering seats, occupancy, placement | AICTE | Institution-level intake AISHE does not carry. |
| MBBS seats | NMC | The seat matrix, per college. |
| Engineering diploma out-turn | AICTE | AISHE cannot see it — see below. |
| How many finish Class X and XII | MoE board results | UDISE counts enrolment, not results. |
| Grade-by-grade school reach | UDISE+ | Classes 1–12, by management and state. |
The distinction the whole thing rests on: flow versus stock
Section titled “The distinction the whole thing rests on: flow versus stock”AISHE publishes a flow — degrees awarded in a year. PLFS observes a stock — people who hold a qualification right now. They are different quantities and dividing one by the other produces a number that looks like a rate and is not one.
So the explorer keeps them apart by construction:
- Graduating cohort takes size from AISHE and only rates from PLFS, multiplying per bucket and summing. Never one blended rate over a mixed population — a flat rate across a class that is 60% general degrees and 8% engineering is an average of things that behave nothing alike.
- Jobs & wages is PLFS-only. Its counts will not tie to AISHE, and the page says so where the counts are.
The two tabs will report different headline rates. That is the design
Section titled “The two tabs will report different headline rates. That is the design”This is the question the explorer gets asked most, so it is worth being exact.
Both tabs read the same PLFS cells. Pick a single qualification on either and the rate is identical — engineering degrees read 42.2% in formal work at the recent-graduate window whichever tab you are on. What differs is how those cells are averaged into one headline:
| Tab | Weights each qualification by | So it answers |
|---|---|---|
| Graduating cohort | how many AISHE says graduated in it | what this year’s graduating class faces |
| Jobs & wages | how many people PLFS finds holding it | what people who already hold the qualification are doing |
Which way the gap runs follows from the two mixes. Among undergraduate levels, PLFS’s population is roughly three-quarters general degrees, where AISHE’s graduating class is nearer three-fifths — and general degrees carry the lowest regular-wage rate of any of them. So the PLFS-weighted average sits below the AISHE-weighted one: 17.4% against 19.2% earning regular wages at the recent-graduate window, on identical filters.
Neither figure is a correction of the other, and neither is wrong. A flow of graduates and a stock of people holding a qualification are different populations, which is the same distinction as the flow versus stock rule above.
One thing that WAS wrong, and is fixed: until 2026-08-25 the cohort tab’s cells did not pin the level, so a cell for an undergraduate bucket included the PG-diploma holders who sit inside PLFS rungs 10 and 11. That made the same qualification read 13.3% on one tab and 13.0% on the other, for no good reason. The cells now agree; what is left is only the weighting described here.
Three things that make PLFS counts smaller than they look
Section titled “Three things that make PLFS counts smaller than they look”PLFS records only a person’s HIGHEST qualification. Someone with a bachelor’s who went on to a master’s appears once, under postgraduate. AISHE counted both awards. So a PLFS “UG degree” population is not the set of people holding a bachelor’s.
The windows are bands, not cohorts. The recent-graduate window is three age-years per qualification (engineering 22–24, general 21–23, medical 24–26), so its headcount is roughly three cohorts stacked. The fixed band is ages 24–29, six years, divided by six for a per-cohort figure. 24–29 and not 24–30 for two reasons: attainment is complete by 24 (graduate-or-above runs 23.0% at 21, 27.8% at 22, 31.6% at 23, 32.7% at 24, then flat), and age 30 is the worst heap in PLFS’s age data, so excluding it drops a known artefact.
None of this affects a count, because PLFS no longer supplies one. The window matters for the rate it is measured at: the recent-graduate window catches a population still 32–36% in education, which is a real fact about recent graduates rather than an error, but it is not the same question as the fixed band asks.
PLFS supplies rates. It never supplies a population
Section titled “PLFS supplies rates. It never supplies a population”This was not always true, and the change is the single biggest correction the explorer has had.
PLFS weights sum to the survey’s own Census-2011-anchored projection, and it records only a person’s highest qualification. Both push its counts below reality, and by different amounts for different groups. Measured against a ~2.5 crore age cohort, PLFS’s own totals for ages 24–29 land at 0.76.
The explorer used to offer a “population basis” filter with four multipliers so the reader could choose a correction. That was the wrong answer to the right observation: a control that offers three wrong numbers instead of declining to make the claim. It is gone.
What PLFS is genuinely good at is a ratio inside a cell, where the weight cancels and no correction is needed at all. So:
- every percentage on every tab comes from PLFS
- every count comes from AISHE, AICTE, MoE or UDISE
- the Jobs & wages tab shows no population figure at all
If you do want the multiplier for something outside the explorer, it depends on the age band and the
question. For ages 24–29: ×1.17 lands at 0.89 of a 2.5 crore cohort, ×1.32 lands at 1.00. There
is no single right factor, which is why ../plfs/ASSUMPTIONS.md A4 refuses to name one.
Why AISHE and PLFS totals differ by about a third — and why that is not an error
Section titled “Why AISHE and PLFS totals differ by about a third — and why that is not an error”| AISHE qualifications awarded, 2023-24 | 10,953,146 |
| as a share of one ~2.5 crore age cohort | 43.8% |
| share of a cohort PLFS finds holding a post-school qualification | ~30% |
The gap is roughly 1.45 awards per qualified person — what you get when a graduate goes on to a master’s, or adds a B.Ed. AISHE counts awards; PLFS counts people. Neither is wrong and the two must never be divided into each other.
As a check on the whole chain: post-school attainment at ×1.32 comes to 29.9% of a cohort, against India’s tertiary gross enrolment ratio of roughly 28%.
The two axes, and why a rung is not a level
Section titled “The two axes, and why a rung is not a level”PLFS asks two separate questions, and the “education ladder” is their cross-product:
- general education (
gedu_lvl) — below middle → middle → secondary → higher secondary → diploma below graduate → graduate → postgraduate and above - technical education (
tedu_lvl) — none, or one of agriculture, engineering/technology, medicine, crafts, other subjects at each of three tiers: technical degree, diploma below graduate, diploma graduate+
A B.Tech holder is graduate × technical degree in engineering. A PGDM holder is graduate × graduate+ diploma in other subjects — a graduate who also holds a diploma, not a level of its own. The data settles it: 0.69m graduates also hold a below-graduate diploma and 1.10m hold a graduate+ diploma.
Reading them as one ladder is not a stylistic error. An earlier version of this explorer made PG Diploma a sibling of PG Degree and reported it at 2.31× AISHE’s PG-Diploma awards, because it was counting 3.3m people who mostly hold a degree and a diploma against 143k awards a year.
Neither axis records a doctorate. gedu_lvl stops at “postgraduate and above”, so a Ph.D. cannot be
told apart from a master’s, and Ph.D. graduates are marked unrateable rather than given a borrowed
rate.
Mapping an AISHE degree onto a PLFS cell
Section titled “Mapping an AISHE degree onto a PLFS cell”The load-bearing assumption of the Graduating cohort tab, in one sentence: the PLFS rate for a cell applies to the AISHE degrees mapped into that cell.
The mapping comes from the survey instrument’s own category list (InstructionManual_VolI §3.4.10),
not from matching names. That matters, because matching names was measurably wrong. Comparing AISHE’s
annual out-turn against PLFS’s stock per age-year:
| AISHE group → PLFS rung | matched by name | mapped from the instrument |
|---|---|---|
| General | 0.64 | 0.50 |
| Engineering | 0.46 | 0.55 |
| Medicine & Health | 0.15 | 0.47 |
| Other professional + Agri + Crafts | 0.09 | 0.40 |
Those ratios should cluster; a seven-fold spread means the categories describe different people. The two low rows explain it: AISHE’s Medicine group is 69% pharmacy, nursing and physiotherapy, and its Other-professional group is 60% B.Ed — and PLFS has no technical category for any of them, so those graduates report “no technical education”. The old mapping handed 253,581 pharmacy and nursing graduates a year a doctors’ employment rate.
The cost of getting this right: pharmacy and nursing now read the general-degree rate. That is very likely too low for nursing. The alternative was the doctors’ rate, which is certainly too high. Neither is measured — PLFS cannot see those qualifications separately — so the cell each bucket used is named on its row.
Thin cells widen and say so. Agriculture at degree level is 24 respondents nationally; crafts is
10. Below 40 the cell widens to the same general level with every subject in it, and the row reads
“widened — only N in
The diploma lane comes from AICTE, and is carved out rather than added
Section titled “The diploma lane comes from AICTE, and is carved out rather than added”AISHE cannot see the diploma population. Its largest diploma “programme” is literally named
Diploma-Diploma with no subject (503,489 graduates), and there is no engineering diploma anywhere
in its list — while PLFS finds 1.30m people holding one. Polytechnic diplomas are AICTE-regulated and
mostly outside AISHE’s degree-awarding frame.
AICTE’s DIPLOMA level is Engineering and Technology in every year: 309,261 passes in 2021-22,
against PLFS’s 216,667 per age-year — a ratio of 0.70, sensibly above the 0.50 degrees run at, since
a diploma holder who never takes a degree stays on the diploma level.
But AISHE’s universe already includes polytechnics — its own aishe_dim_standalone_institutions
schema says so and names AICTE as their regulator. So AICTE’s figure is 53% of AISHE’s unlabelled
diploma lump, not an addition to it. The engineering-diploma row is therefore carved out of that
lump and capped by it; the residual stays unattributed and visible. Adding it on top, which an
earlier version did, inflated the graduating class by a third of a million.
AICTE’s panel ends at 2022-23, loaded to 3% of intake, so the newest usable year is 2021-22 and the row says which year it borrowed. 2020-21 is a COVID trough (54,072), not a trend.
School leavers: MoE for size, and one year of UDISE is a synthetic cohort
Section titled “School leavers: MoE for size, and one year of UDISE is a synthetic cohort”Everyone who goes to college is on the Graduating cohort tab; the School leavers tab is everyone who does not — around 70% of an age cohort.
Size comes from MoE board results, because UDISE counts enrolment and “highest qualification is Class X” is a statement about who completed a level. Class X in year Y and Class XII in year Y+2 are the same cohort, so the tab only offers pairs exactly two apart; the source has no 2023, which leaves 2022 → 2024 as the only clean pairing.
| passed Class X, 2022 | 15,848,975 |
| passed Class XII, 2024 (same cohort) | 12,929,386 — 81.6% |
| sat Class X and failed | 2,747,689 |
| sat Class XII and failed | 2,015,718 |
UDISE still answers grade-by-grade reach, and shows the largest narrowing at Class 10 → 11 (−20.2%), consistent with the Class X fail figure. But one year of UDISE is a synthetic cohort, not a cohort — Class 1 and Class 12 in the same year are different children eleven years apart, so a grade-to-grade step mixes real attrition with a decade of change in cohort size and coverage. It shows where the ladder narrows, not how many individuals left. A real attrition rate needs several years of UDISE, which is not loaded.
Two independent measurements of nearly the same thing agree, which is worth having: UDISE Class-10 enrolment (17,989,819) and MoE Class X appeared (18,553,676) are within 3%. The Class-12 pair agrees less well, because private and open-school candidates sit boards without being enrolled in a UDISE school.
What the sources cannot do
Section titled “What the sources cannot do”AISHE has no state × discipline cell. State × level, and discipline nationally, are separate published cuts. Selecting a state therefore drops the field breakdown; the page says so instead of showing a filter that quietly returns nothing.
PLFS cannot see a PhD. Its education ladder runs school → diploma-below-graduate → graduate → postgraduate and above, and stops. A doctorate is indistinguishable from a master’s at any level. The class exists on the AISHE-sourced tab, marked unrateable.
PG Diploma comes from a different ladder. It is not a step in the general one, so it is read from PLFS’s technical ladder, which lands only on the graduate and postgraduate steps and carves cleanly out of both. Worth separating: it carries the highest median wage of any level, and folded into the degree classes it was invisible.
AICTE 2022-23 is partially loaded — enrolment about 3% of intake. Years are measured for completeness and labelled; the default is the newest year that clears the bar.
UDISE+ holds only 2024-25 here, so the school rungs cannot be read as a time series.
Medians, and why some say “approx.”
Section titled “Medians, and why some say “approx.””A median is not additive. Each is computed at the cell it belongs to, using 100 quantile buckets. Where the display spans several cells the figure is a weighted average of those cells’ own medians — marked approx. — never a recomputed quantile over pooled data. Medians are also unweighted sample medians while the rates beside them are population-weighted.
The casual daily wage is the exception, and is exact. Blending is not merely approximate on that distribution, it is biased: the blended figure read ₹432 against a pooled ₹400, an 8% overstatement of a number PLFS itself publishes. Counts, unlike quantiles, do compose — so the extract emits the casual daily wage as a histogram of 33 weighted counts, and the page reads the median off those counts merged under whatever filter is showing.
The bucket edges are 16 round rupee values (₹150, ₹200 … ₹800, ₹900, ₹1,000), each its own singleton bucket, with a gap bucket between neighbours. They are pinned, not derived on a rebuild. InstructionManual_VolI §3.6.11(c) requires the daily wage to be entered “in whole number in rupees”, so the distribution is discrete and heaps hard — 899 distinct values, the top 10 carrying 75.5% of all weight. Interpolating across a spike is biased upward rather than merely coarse: uniform buckets return ₹406.7 at width ₹10, ₹431.0 at ₹50 and ₹449.9 at ₹100 against an exact weighted median of ₹400.
Where the cumulative half-weight lands on a singleton the median is that value, exactly — verified on
all 712 state × sex slices of the nine releases carrying the daily block, of which 613 resolve to a
singleton, covering 89.5% of the casual population, and all 613 match the exact weighted median from
microdata (sql/check_casual_median.sql, which must stay at 0 mismatches). Where it lands in a gap
the answer is genuinely an interval between two spikes, and the page prints the interval. A
midpoint would be invented precision.
Both outcomes are visible on the default view: all-ages CY2025 lands on ₹400 exactly, while the age 25–29 slice reads ₹400–₹450 against a true weighted median of ₹408.33.
Days and hours count casual work only. §3.6.9 records hours for every work status (codes 11–72); §3.6.11 records wage earnings only for casual work (41, 42, 51). Both are per activity per day and so pair cell by cell — what is false is that every hours cell has an earnings cell behind it. A usually-casual person’s self-employed or regular days carry hours and a zero wage. So a day counts when it carries casual earnings, and hours come only from cells whose paired earnings cell is positive. Measured on calendar_2025: 41,080 of 385,663 hour-carrying activity-day cells (10.7%) have a zero wage, and none has a blank one, so the test is unambiguous. Counting any day with hours put days a week at 5.414 instead of 5.201 and the exact weighted hourly wage at ₹50.00 instead of ₹53.13. The daily median is ₹400 on either denominator — the heaping makes it robust — which is why this correction did not move the headline figure.
That correction is also what extended the series from two releases to nine. hr11–hr27 exist
only in calendar_2024 and calendar_2025, so a days-from-hours derivation could only ever cover those
two; ern11–ern27 are populated in nine. The old scope was a limit of the method, not the data.
The hourly figure is still two releases, legitimately — it needs the hours block, and the page
blanks that column where it does not exist.
The monthly regular wage gets the same treatment one rung coarser, and always reports a band.
ern_reg heaps too, but on a far longer tail: its top 10 values carry only 45.1% of weight against
casual daily’s 75.5%, and a top-30 union resolves just 68.7–96.1% of the population depending on
release. Singleton spike buckets would therefore cost dozens of measures and still leave most slices
unresolved. So the extract emits 13 counts on band edges instead — ₹5,000, ₹8,000, ₹10,000,
₹12,000, ₹15,000, ₹18,000, ₹20,000, ₹25,000, ₹30,000, ₹40,000, ₹50,000, ₹75,000 — half-open [lo, hi), pinned, and every edge a multiple of ₹1,000 because 80.4% of reported monthly wages are, so an
edge splits almost no value’s mass.
The median is read off those counts merged, which locates it inside a band and no more precisely.
The band is what is shown: “₹15,000–₹18,000”, never a point inside it. sql/check_monthly_wage_bands.sql
asserts the exact weighted median from microdata falls inside the band this method picks, on every
state × sex slice of every release — 741 slices, 0 outside, and the closed bands average 27.7% of
their own lower edge in width. On the default slice the band is ₹15,000–₹18,000 against a blended
point of ₹17,782, at the very top of it.
These counts are weighted, where median_wage is an unweighted sample median. That is a
deliberate change of estimand: the band is the population’s, matching the rates printed beside it.
Their population is ern_reg > 0 with no status filter — the same population median_wage and
wage_earners have always been on. That is about 2% wider than activity code 31 alone (measured on
calendar_2025: 11.77 crore at code 31 plus ~0.24 crore across codes 91, 92, 11, 51, 81, 21, 93), so
the “Regular wage work” row’s count and its pay band are not on identical populations, and the page
says so. Note that sal_b1–sal_b4, the wage bands on the cohort tab, do filter to code 31 —
the two are not interchangeable.
The cohort tab’s median is a MIXTURE, which is the same fix in a different shape. That card asks a different question from the rest of the page: not “what is the median in this slice” but “what would this graduating class earn”, where the class comes from AISHE and the wages from PLFS. It used to be a graduates-weighted average of each bucket’s own median.
Its buckets cannot simply be summed. A bucket whose own (general level × technical category) cell is below the sample floor falls back to the whole general level widened across subjects, so several buckets can share one cell and that cell overlaps the exact cells of its siblings — in AISHE 2023-24, “Postgraduate & above, all subjects” is the fallback for three buckets and overlaps three exact postgraduate cells. Adding their counts would count the same respondents repeatedly.
So each cell is used as a shape, not a count: its 13 band counts are normalised to shares of itself, scaled by that bucket’s graduates, and summed. That is the wage distribution of a synthetic cohort — every graduate assigned the wage distribution of the cell their qualification reads, mixed in AISHE’s proportions. A mixture of distributions is a distribution, so its median is a real median of that mixture, which an average of medians never was. Normalising is what makes the shared and overlapping cells harmless: a cell lends its shape once per bucket pointing at it, weighted by that bucket’s graduates, and never contributes its headcount.
It remains an estimate — the mixing weights are AISHE’s and the shapes are PLFS’s. What it is no longer is a statistic with no definition. For AISHE 2023-24 the card reads ₹12,000–₹15,000; the average of medians it replaces read ₹15,121, outside that band entirely, pulled up by small high-wage buckets (medicine graduates are ₹40,000–₹50,000 on 1.2% of the class) while 63% of the cohort sits in the lowest band shown.
The Jobs tab’s other three went the same way. The formal wage and self-employment income are
bands on the same edge set as the regular monthly wage — deliberately, so the three monthly rows of
the work-status table can be read straight down the column rather than each carrying a private
vocabulary. Measured, both sit inside that range (formal p10 ₹7,000 / p50 ₹15,000 / p90 ₹48,600;
ern_self p10 ₹3,500 / p50 ₹12,500 / p90 ₹30,000). Neither is concentrated enough for singletons —
top-10 weight is 45.2% for ern_reg, 39.7% for ern_self, against casual daily’s 75.5% — so both
always report an interval.
The casual per-hour figure does get singletons, on an eighths lattice rather than round
rupees. §3.6.9 rounds hours to whole numbers and §3.6.11(c) makes the daily wage whole rupees, so the
quotient is a ratio of small integers and heaps on ₹6.25 steps — a round daily wage over an 8-hour
day. ₹50.00 alone carries 27.3% of the weight. It resolves to an exact rate for 61.8% of slices,
covering 71.5% of casual workers, and to an interval for the rest. ₹57.143 (400/7) and ₹66.667 carry
real weight and are deliberately excluded: besides being non-round, they are not exactly representable
in binary floating point, so = against them would silently match nothing.
sql/check_jobs_medians.sql asserts all three — containment for the two band families, and exact
equality for every hourly slice that resolves to a singleton. Current: 729 formal slices and 749
self-employment slices with 0 outside their band, and 88 of 142 hourly slices resolving to a
singleton with 0 mismatches.
What is still a blend, and why. Only the tables reading the occupation cube — by job function and by sector. That cube carries no bucket columns at all, so a weighted average of cell medians is the only thing available there, and it stays marked approx. wherever it appears.
Four measures left the payload with this change. median_wage, median_formal_wage,
median_self_month and median_cas_hourly are all per-cell quantiles that nothing reads any more.
Leaving a blendable median in the payload beside an exact one is a loaded gun — averaging it is the
bug this whole line of work removed — so the surest way that nobody averages it is for it not to be
there. The occupation cube keeps its own median_wage, which it genuinely needs.
What the checks prove, and what they do not. check_casual_median.sql and
check_monthly_wage_bands.sql each re-derive their buckets inline from the microdata; neither reads
an emitted casd_* or ernb_* column. They are regression tests on the edge sets and the
median-picking rules — they fire if an edge moves or a CASE is hand-edited — and they are where the
band-width figures above come from. They are not evidence about the pipeline. Two checks in
check_invariants.py cover that, reading real emitted data: the buckets must sum, per cube row, to
the population they describe; and app.js’s hand-mirrored CASD_EDGES/ERNB_EDGES must derive
exactly the cube’s column names, in order, since that mirror drifting would blank every wage cell on
the page without any other gate noticing.
What the histograms cost. Measured on real rebuilds, gzipped, per measure: casual daily 3.7 KB, casual hourly 7.8 KB — both sparse, since the daily block exists in only two of eleven releases — against 34.4 KB for the formal bands, 54.4 KB for self-employment and 73.8 KB for the regular monthly bands, which are dense because wage earners exist in nearly every cell. Sparsity, not column count, is what sets the price: the 21-column hourly family costs a quarter of what the 13-column self-employment one does.
First paint went 500 KB → 611 KB with the first two families, and 611 KB → 735 KB with the last three. Trimming edges is not a real lever — 13 monthly buckets down to 6 saves 28 KB and roughly doubles every band’s width. Every bucket family rides the roll-up cube only, never the status cube, which is what keeps the cost to that: carrying the casual-daily set on the status cube as well cost +309 KB gzipped, 63% of that family’s whole payload, for columns nothing reads.
Real terms, and why the wage trend defaults to them
Section titled “Real terms, and why the wage trend defaults to them”The cohort trend spans a decade. Plotted in nominal rupees, comparing its ends measures inflation as much as outcomes — so wages on that chart are in constant rupees by default, with a Wage basis control to switch back to nominal for a single year’s figure.
It matters more than a rigour point. On the default view the nominal series runs ₹10k–₹12k → ₹12k–₹15k, which reads as a rise. In constant CY2025 rupees it runs ₹14k–₹17k → ₹13k–₹16k — flat, if anything slightly down. The deflator changes the sign of the headline reading.
The index is cpi_india.csv, refreshed by fetch_cpi.py from the World Bank’s FP.CPI.TOTL
for India (calendar-year averages). The CSV is committed and is the source of truth for the build, so
a network failure cannot silently move a published number. Each calendar year’s average is anchored
at 1 July and interpolated to the release midpoint the build already defines, which reuses one
convention rather than inventing a second: a Jul–Jun release lands midway between two calendar years,
which is their mean.
Its known imprecision is that this is a calendar-year index applied to Jul–Jun releases. A MoSPI CPI-Combined monthly series would remove that, and is the better source for India-specific work; it is not used only because it is not as cleanly fetchable. The interface — one factor per release — would not change.
Deflating a band is exact, not approximate. Every wage figure here is an interval read off pinned
nominal buckets. Deflation is strictly monotone, so a nominal median in [lo, hi) gives a real
median in [lo×f, hi×f). Bucketing stays nominal, the check queries are untouched, and only the
displayed bounds move. The scaled edges stop being round numbers, which does not matter — roundness
mattered at bucketing time so an edge would not split a spike of reported values, and has no meaning
at display time.
Deflation is applied per point, where the release that point was rated against is known. One point can read a different release from its neighbour, so a single page-level factor would be wrong.
Vintage
Section titled “Vintage”Every tab carries a Data built … line naming the UTC timestamp and the source commit for the data
on screen. The extracts are re-run by hand, so that is the vintage of the numbers, not of the page.
The build refuses to publish a payload it cannot date.