Method, and what it cost to get here
Everything settled, why it was settled that way, and what is still open. Written so the analysis
can be picked up cold. The prose piece is README.md; the rebuild commands are in
TECHNICAL.md.
The chain
Section titled “The chain”Four steps, in this order. Each source is used only for what it is good at.
| Step | Source | Gives | Never used for |
|---|---|---|---|
| 1. How many finish | AISHE | counts — degrees actually awarded, per year | rates |
| 2. What happens to them | PLFS | rates — share in formal work, share clearing ₹6L | counts |
| 3. Multiply | national totals | ||
| 4. What ranked colleges claim | NIRF | placements and salaries, per college | anything unranked |
| 5. Subtract | the residual, and whether it is even positive |
Counts never come from PLFS, and rates never come from AISHE. That separation is load-bearing — an earlier version took counts from PLFS, had to gross them up 1.16–1.22× to reach true population, and that adjustment silently drove every derived row.
Why rates need no population gross-up. PLFS weights sum to 1.19bn, roughly 15–20% below India’s true population, because the projections are anchored to Census 2011. Harmless here: both numerator and denominator carry the same undercount, so it cancels. It would not cancel if we took levels from PLFS, which is why we don’t.
The decisions, and the reasoning
Section titled “The decisions, and the reasoning”Employment means FORMAL employment
Section titled “Employment means FORMAL employment”Regular salaried, paid, with a written contract or social security (pas='31' and ern_reg>0
and (job_pas IN ('2','3','4') or ssec_pas NOT IN ('8','9'))).
A NIRF placement is a company hiring you onto a real job. Informal work — self-employment, or
salaried with neither contract nor benefits — is real work but invisible to campus recruitment,
so comparing it against NIRF is a category error. It is carried alongside in
national_totals.csv and reported in its own section, never netted in.
Unpaid family helpers (pas='21') are excluded from employment entirely. PLFS records them as
employed; they earn nothing. Nationally they are 15.6 of the 56.2 points of “self-employment”, so
leaving them in badly overstates work.
The age window matches the moment NIRF measures
Section titled “The age window matches the moment NIRF measures”This was the single most consequential fix. NIRF reports the salary a college’s students were placed at on graduation. Measuring the national comparison at ages 26–29 compared fresher offers against people four to eight years into a career, and Indian engineering wages grow steeply across exactly that span:
| Age | 22 | 23 | 24 | 26 | 29 |
|---|---|---|---|---|---|
| % in formal work | 25.1 | 38.5 | 40.6 | 56.1 | 63.6 |
| % clearing ₹6L | 4.1 | 3.3 | 5.7 | 13.5 | 25.3 |
| Median formal wage/month | ₹27,500 | ₹30,000 | ₹30,000 | ₹32,650 | ₹40,000 |
That gap, not college quality, was what made unranked colleges appear to out-place ranked ones.
The rule adopted: three years from each field’s normal completion age.
| Field | Window | Why |
|---|---|---|
| Engineering | 22–24 | four-year degree entered at 18 |
| Doctors | 24–26 | MBBS is 5.5 years including internship, so licensing is ~23.5 |
| Everything else | 21–23 | three-year degree |
Doctors at 23–24 is too early — 75% are still studying and the sample holds one person above ₹6L.
“Medical” means doctors — MBBS and BDS
Section titled ““Medical” means doctors — MBBS and BDS”It did not, on two of three sides, and the error ran about 3.9× in the denominator:
- NIRF ranks Medical and Dental separately from Pharmacy and Nursing. Correct already.
- AISHE “Medical Science” (330,839) is the whole health cluster — B.Sc Nursing alone is 291,498 graduates. Replaced with M.B.B.S. (65,836) + B.D.S. (20,098) = 85,934.
- PLFS
tedu_lvl='04', “technical degree in medicine”, covers nursing and allied health too. Of 281 people in that class only 45 work as Medical Doctors, against 82,198 weighted nurses, pharmacy technicians and paramedics.
Because no PLFS education code isolates MBBS, doctors are identified by occupation (NCO group 221) with the medical class still studying added to the denominator — at 24–26 a medical-degree holder in full-time study is doing NEET-PG or MDS, which requires the degree. That study group is 63% of the denominator, and omitting it would have claimed near-total employment for doctors by measuring only those who had finished training.
What this misses: doctors who are unemployed, in domestic duties, or working outside clinical occupations. Small for MBBS, but it makes the denominator a lower bound and both rates an upper bound. Cells are small — n=94, n=7 above ₹6L — and are printed with the table.
Nursing and pharmacy graduates sit in the residual, which is also where NIRF puts them.
Count the ranked programme, not the ranked institute
Section titled “Count the ranked programme, not the ranked institute”Filtering institute keys by ranking category and then aggregating all their rows counted every programme the institute runs — a medical college ranked under Overall dragged in its nursing and B.Sc UG programmes:
| Category | UG grads in-category | All UG rows at those institutes | Inflation |
|---|---|---|---|
| Engineering | 225,013 | 551,277 | 2.45× |
| Medical | 9,499 | 121,534 | 12.8× |
| Dental | 4,527 | 97,466 | 21.5× |
Elite medical showed 647 graduates per college against a real MBBS class median of 93. Fixed by filtering rows to the ranking category before aggregating.
Salaries are cut 20%; placement rates are not
Section titled “Salaries are cut 20%; placement rates are not”Cost-to-company bundles what never reaches a payslip: employer PF 5–6% of CTC, gratuity ~2%, unrealised variable 3–5%, annualised joining bonus 1–3%. Eleven to sixteen percent before anyone exaggerates. Three independent checks land in the same place — Avanti’s own alumni started on 0.75× what their colleges filed; DTU’s per-student NAAC list gives a ₹10.50L median against ₹13.25L filed to NIRF; and the CTC arithmetic gets there from statutory rates alone.
No placement haircut. We looked for evidence and could not find it. NIRF’s internal arithmetic is consistent, AICTE’s placement series is unusable, and our own alumni turn out more employed than their colleges claim. An unevidenced haircut is not better than none.
The reconciliation below vindicates that split: placement fits, salary does not.
Elite is defined on the reported median
Section titled “Elite is defined on the reported median”Median salary ≥ ₹6L and placement rate ≥ 50%. Selection uses the reported median, not the post-haircut one, because the elite threshold and the quality-job cut are both ₹6L — discounting the selection too would apply the haircut twice.
Within-college salaries are log-normal, σ = 0.45
Section titled “Within-college salaries are log-normal, σ = 0.45”NIRF publishes one median per college, not a distribution. Treating a college as wholly above or below ₹6L would credit a ₹4L college with zero good jobs and push them into the residual, which then beats the ranked tier — an artefact, not a finding.
σ is not load-bearing (check_pool_plausibility.py). Across σ from 0.20 to 0.80 the ranked
total moves 58k–69k and the over-claim stays 1.67×–1.98×. Removing the log-normal entirely lands
within 1.2% of the same total. σ shifts jobs between elite and ranked-not-elite; it does not
move the conclusion. It rests on DTU alone (which implies 0.471), and that is now acceptable.
The result: placement reconciles, salary does not
Section titled “The result: placement reconciles, salary does not”| Field | Measure | Ranked colleges claim | National (PLFS) | Room left | Ratio |
|---|---|---|---|---|---|
| Engineering | placed | 159,731 | 271,802 | 112,071 | 0.59× |
| Engineering | ≥₹6L | 62,612 | 34,689 | −27,924 | 1.80× |
| Doctors | placed | 7,443 | 22,869 | 15,426 | 0.33× |
| Doctors | ≥₹6L | 3,748 | 8,586 | 4,838 | 0.44× |
| Everything else | placed | 82,767 | 521,569 | 438,802 | 0.16× |
| Everything else | ≥₹6L | 17,404 | 11,833 | −5,571 | 1.47× |
Where the salary ratio exceeds 1.00 the unranked residual goes negative. That is the result, not a defect.
Why the small national pool is believable
Section titled “Why the small national pool is believable”The obvious objection is that ~34,700 entry-level ₹6L engineering jobs is too few, and that PLFS —
a household survey — misses well-paid young engineers in shared urban rentals. Three checks say
otherwise. Run check_pool_plausibility.py and check_plfs_pool_ci.py.
The audited anchor. IITs, NITs and IIITs are the institutions least likely to be misreporting. If PLFS were blind to elite young engineers, they alone would overflow the pool. They don’t:
| Colleges | UG grads | Jobs ≥₹6L | Filed median | |
|---|---|---|---|---|
| IIT | 22 | 13,536 | 9,844 | ₹17.52L |
| NIT | 28 | 19,128 | 10,991 | ₹10.28L |
| IIIT | 5 | 1,122 | 788 | ₹11.00L |
| Total | 55 | 33,786 | 21,623 | |
| Other ranked | 208 | 191,227 | 40,990 | ₹5.00L (placement-weighted) |
The public elite fill 62% of the national pool from 4% of graduates — comfortably inside it, with room for BITS and the top private colleges. The over-claim lives entirely in the other 208. The ranking’s top is credible; its tail is not.
Coverage. 58% of formally employed engineering graduates aged 22–24 are in NIC 62 (computer programming and consultancy) at a median of ₹32,600/month — precisely the TCS/Infosys/Wipro fresher band. 71% urban. PLFS is finding the software workforce; it simply pays ₹3.9L. ₹6L sits at the 90th percentile of employed fresh engineers (p50 ₹28,000, p90 ₹50,000, p99 ₹78,000).
Sampling. Bootstrapped over 916 first-stage units, because treating PLFS persons as independent understates the standard error 3.5–4.7×. The rate is 4.42%, 95% CI [2.19%, 7.19%]; the pool is 34,689 [17,145 … 56,441]. NIRF’s 62,613 is outside the upper bound.
Bugs found, and the lesson each one carries
Section titled “Bugs found, and the lesson each one carries”| Bug | Effect | Lesson |
|---|---|---|
nirf_fact_aggregate sums byte-identical duplicate rows, including median salary |
29–54 institutes inflated 2–3× per year; headline 97% → 74% | Verify a derived table against the primary source (we checked NIRF’s own scorecard PDFs). Filed as external_data_sources#73; use nirf_fact_master with MAX() |
| Institute-key filtering aggregated every programme at a ranked institute | medical grads 12.8×, dental 21.5× | A per-college number is not a per-programme number |
AISHE “Medical Science” and PLFS tedu_lvl='04' both mean health, not doctors |
denominator 3.9× too large, numerator diluted with nursing wages | Check that a category means the same thing on every side before differencing |
| 26–29 window against NIRF’s at-graduation figures | unranked appeared to out-place ranked | Match the measurement moment, not just the population |
Postgraduates included (gedu_lvl IN ('12','13')) |
medical quality rate inflated 82% | |
Diplomas included (tedu_lvl ‘13’/‘14’) |
‘14’ is 29% of PLFS “medical” | |
if x["median"] in Python |
NaN is truthy; the whole unranked median column silently blanked | Test pd.notna, never truthiness, on a float that can be NaN |
gsutil -m cp output/* |
swept a 6,692-row alumni file with names into the public bucket | Never glob into a public bucket; use the grep-excluded form |
APPROX_QUANTILES(..., 2)[OFFSET(1)] for medians in extract_informal_wages.sql |
non-deterministic — two identical runs returned ₹10,800 then ₹10,000 for the same 695 people, and ₹17,000 vs ₹17,500 for the same 1,057 | 2 buckets is too coarse to be stable. Use 100 ([OFFSET(50)]). Caught only because a rebuild diff showed a number moving with no code change — a pipeline that cannot be re-run and diffed hides this class entirely |
What is still open
Section titled “What is still open”Re-verified by rebuilding the whole pipeline from the extracts, not by reading this list — three items that were on it are done, and saying so matters as much as listing what is not.
Closed since this list was written:
README.md and ONE_PAGER.md are stale.Both carry the entry-level window, the MBBS+BDS definition and the post-fix funnels. Verified by grepping every headline figure.Regenerated and compared pixel-by-pixel against the committed files: identical. The byte-level diff a rebuild produces is matplotlib metadata, not content.tier_*.pngneed regenerating.σ / fat-tail sensitivity.check_pool_plausibility.pycovers it: the over-claim holds at 1.67×–1.98× across σ 0.20–0.80, and removing the log-normal entirely moves the total 1.2%.
Still open:
overlay_plfs.pyandextract_plfs.sqlare superseded bybuild_tier_tables.pyandextract_plfs_fields.sql, butnirf_plfs_overlay.pngin the workbook still comes from them. Carve out as its own cleanup.- Nothing has been re-uploaded to GCS this session, so nothing public is currently wrong. Do
not upload until this merges — and never with a bare
output/*glob. - The doctors row rests on small cells (n=94, n=7 above ₹6L). Directionally sound, not precise. Say so wherever it is used.
- The “everything else” unranked residual is negative (−5,571 quality jobs). That is not an estimate of anything; it is the arithmetic saying ranked colleges alone claim more than the national pool holds. It is the finding, not a quantity — never quote it as a count.
- None of this is causal. Elite colleges select hard on entrance rank; their graduates would earn more whoever taught them. This describes who holds the jobs, not what produced them.
SOURCES.mdcharacterisations are legally sensitive (“SRM contradicts itself”). The repo is private and the source bucket returns 403. If ever published, convert to attributed statements and offer right of reply.