Skip to content

Build decisions and folder guide

Companion to the paper, which is this folder’s README.md. This file holds the things a reader of the argument does not need but a rebuilder does: what each file is, the decisions baked into the build, and how to regenerate the page.

Where a 2.5 crore birth cohort ends up: from Class 10 through eighteen final qualifications to the nature of work at ages 25–29, and what each qualification is worth at 25–29, 30–34 and 35–44. PLFS calendar-2025 microdata only — no administrative registers in this pass.

File What it is
pathways.html The page: Sankey, worth table with an age-band switch, dumbbell chart, pay-vs-employment scatter at 35–40, pay / regular-work / people rank table, government-vs-private size table with two bubble-sized scatters, fees table with sources, method notes. Built, never hand-edited.
template.html The page’s HTML/CSS/JS with /*__DATA__*/ and __LOGO__ markers. Edit this.
build.py Reads data/*.csv, derives flows and expected earnings, writes data/pathways.json and pathways.html.
sql/00_bucket_case.sql The endpoint definition (the CASE every extract shares).
sql/01_flow_25_29.sql Endpoint × general level × nature of work, ages 25–29 → data/flow_25_29.csv.
sql/02_worth_by_band.sql Employment, salaried, formal, wage medians and thresholds by band → data/worth_by_band.csv.
sql/03_voc_codes.sql The voc code check behind the “+ vocational” buckets.
sql/04_worth_35_40.sql The 02 query at a single band, ages 35–40, for the pay-vs-employment scatter → data/worth_35_40.csv.
sql/06_worth_35_40_by_sex.sql The 04 query split by sex → data/worth_35_40_by_sex.csv. Men are employed at 96–100% on every pathway, so the pooled ranking on the page is a composition effect; the paper ranks within sex.
sql/07_trainee_school_level.sql School level of people reporting full-time formal vocational training (the ITI-entrant proxy) → data/trainee_school_level.csv.
sql/08_within_quintile_men_30_44.sql Training and diploma premium within household-consumption quintile, men 30–44 (PLFS household join) → data/within_quintile_men_30_44.csv.
sql/09_pseudo_panel_unemployment.sql Unemployment and salaried share for men born 1994–97 and 1998–2001 across six releases → data/pseudo_panel_unemployment.csv.
sql/10_sense_check.sql Five checks behind whitepaper/SENSE_CHECK.md: whole-cohort job status by sex and school level, who makes up the not-salaried and the unemployed, government’s share of salaried work across releases, nominal wage trends for five groups, urban share of holders → data/sense_*.csv.
data/registers.csv Hand-assembled register figures per endpoint: AISHE 2023-24 out-turn (programmes listed per row), AICTE 2021-22 approved intake, NMC 2024-25 MBBS seats, NIRF DCS institute-of-national-importance intake, MoE 2024 board passes. The mapping notes say where the NSS code pools programmes. Loaded into the JSON but no longer drawn — the capacity layer was taken off the page pending the qualitative argument about why seats sit empty.
data/ownership.csv Government vs private size per pathway: AICTE seats by institution type, NMC MBBS seats, AISHE graduating classes apportioned by the ownership mix of the relevant colleges or standalone institutions, UDISE+ school enrolment for the school rows. Each row carries its measure, a confidence flag and the basis text.
sql/05_ownership_inputs.sql The queries behind ownership.csv.
data/fees.csv Typical annual fees per pathway, government vs private, with low/high range, basis text and pipe-separated sources. Snippet-level: the research pass could not open any page (network policy), so figures are search-excerpt readings of the cited pages except where they coincide with the warehouse fee table.
HYPOTHESIS.md Pressure test of the “government funds the top and bottom, leaves the middle to private fees” hypothesis: warehouse evidence, snippet-level literature, claim-by-claim verdict.
CREDENTIALS_AND_NCVT.md What a degree or diploma unlocks by rule (7th CPC ladder), what drives degree demand (literature, ranked by evidence), and what an NCVT certificate is worth — PLFS tables on government jobs, occupations, the waiting curve and ITI-like formality and pay, against the tracer studies and the cheap-labour critique.
HANDOFF_PASS2.md Return from the first web-enabled pass: spending verified (PM-SETU ₹60,000 cr read in full; MERITE ₹4,200 cr), institution-count bias in the ownership estimates, ITI capacity bimodality, source log, next steps.
RESEARCH_BRIEF.md Hand-off brief for a session with web access: verify the excerpt-level fee and capacity figures, fill the gaps, re-test the hypothesis against documents read in full.
README.md The paper, The lost middle (v3). Generated from paper_source.md by render_refs.py — never hand-edit it.
paper_source.md + render_refs.py The paper’s source, with {{key}} reference markers, and the renderer that numbers them and appends the linked reference list.
whitepaper/ The workings behind the paper: literature tables, four critiques, the five-persona debate, the recommendations brief, the verified-source manifest (SOURCES.md) and the open queue (LEFTOVERS.md).
REVERIFICATION.md Line-by-line re-run of the earlier chat analysis this folder grew out of.
india-education-pathways-analysis.md That earlier analysis, as received.
  • Scale. Every flow is lakh per 2.5 crore born. The split at Class 10 uses the reverified CY2025 age-25-29 share below secondary (38.86%); everything to the right is a share of the Class-10-and-above population re-scaled the same way. PLFS weights sit ~16% below the population, so none of this is a PLFS count.
  • Endpoints. Technical qualification first, then general level, with formal vocational training (voc = '1') folded into the general-level bucket it sits on. Engineering PG = technical degree in engineering (03) with general level postgraduate (13), not code 13, which is a graduate-level diploma. gedu_lvl = '11' with no technical code is counted with “Sub-degree diploma, other”.
  • Nature of work. Principal status only. Formal = pas = '31' with a written contract (job_pas IN 2,3,4) or any social security (ssec_pas NOT IN 8,9).
  • Expected monthly earnings = share regular salaried × median regular wage + share self-employed × median self-employment income + share casual × (week’s daily casual earnings × 4.33). Non-earners contribute zero. Cells with fewer than 100 sampled regular earners are marked “~”.
  • Colour. Flows are coloured by qualification family, not by bucket: eighteen hues would fail any colour-vision check. The three series hues were validated with the dataviz skill’s validate_palette.js in both themes (light #A8322A / #E2750F / #1976A8 on #FDFAF1; dark #D64A7E / #D9701C / #3A9AD9 on #1B1A17). The brand orange sits at 2.97:1 against cream, so every node is labelled and a table view exists. School-only flows are a deliberate neutral, outside the categorical set. Brand tokens otherwise follow the deck guide (cream, ink, Lora + Figtree).
Terminal window
python3 analysis/plfs/pathways/run_extracts.py # sql/ -> data/*.csv (needs bq; ~6 GB scanned)
python3 analysis/plfs/pathways/render_refs.py # paper_source.md -> README.md
python3 analysis/plfs/pathways/build.py # data/ -> data/pathways.json + pathways.html

./build_pdf.sh renders the paper to PDF and uploads it to the analysis bucket. The PDF is not committed: it is 1.9 MB and nothing over 1 MB goes in git.

run_extracts.py substitutes the endpoint CASE from 00_bucket_case.sql into every query that marks it, so the definition is written once rather than pasted into fourteen files. It used to be a manual paste, which is how the extracts came to be missing from every checkout.

Three inputs are not machine-generated and are not rebuilt by that first command: registers.csv, ownership.csv and fees.csv are hand-assembled from published registers, with their sources in the basis and sources columns of each row. cpi_annual.csv is keyed in from the Economic Survey. All four are in the bucket with the rest.

  • The paper is this folder’s README.md, per the convention in the convention in the repository’s analysis/README.md, which is an internal index. See whitepaper/LEFTOVERS.md Tier 0 for why, and what publishing it would take.
  • Row-level data. The data/*.csv extracts are gitignored, as the convention requires, and are hosted at gs://avantifellows-bq-assistant/analysis/plfs-pathways/data/. Every [W] reference in the paper links to the hosted copy of the CSV behind it, beside the query that produced it.
  • whitepaper/SOURCES.md is generated, not hand-written, and it is published, so its access: public frontmatter has to be emitted by whatever regenerates it. Adding the frontmatter by hand lasts exactly until the next regeneration and then fails closed and silently, which is how it came off once already.
  • Cited source documents are archived at gs://avantifellows-research-sources/plfs/pathways/whitepaper/, which avantifellows.org accounts can read and nobody else can: the archive is third-party copyrighted material kept for citation checking, and the paper’s references link to it. The first archive, gs://avantifellows-external-data/reference-papers/plfs-pathways/, holds the same 38 files. whitepaper/SOURCES.md is the manifest, with a SHA-256 per file.