Skip to content

What data we hold

These figures are a snapshot, and they will change. Our data is actively being worked on: new sources are still being ingested, cohorts are still graduating, and identity-matching passes are still being run across datasets we already hold. Every number below is therefore a picture of one moment, not a settled total — and the direction of travel is generally upward, particularly for linkage, where much of the gap is passes not yet run rather than data that cannot be joined. Figures generated 2026-08-11. If you are citing anything here, cite it with that date, and re-run the build for a current answer.

Aggregate counts from the analytical warehouse. Nothing on this page identifies an individual student: every figure is a cohort-level count or a share.

Every number here is generated by build.py from a single warehouse table and committed to output/. check.py gates this prose against those extracts, so a figure cannot be edited by hand without failing the check — including the date above. See Reproducing this.

Avanti’s data divides into two halves: what we observe directly about students in our own programmes, and public or partner datasets about the wider system those students sit inside. Most of the interesting questions need both, which makes linking them the hard part — and the part that is least finished.

Our own systems produce five broad kinds of student data:

  • Enrolment and identity over time — which programme, batch, school and centre a student belongs to, and how that changes across academic years. A student is not a fixed row; their enrolment is a journey with transitions.
  • Attendance at live classes — per student, per session. This is the highest-volume operational signal we hold and the one most used to spot disengagement early.
  • Test performance, down to the individual question — results are modelled at overall, chapter and question level, alongside a separate record of who actually sat each test. Question-level granularity is what makes diagnostic analysis possible rather than just ranking.
  • Self-driven practice — practice attempts a student chooses to make outside scheduled tests, kept separate from formal assessments so the two are never conflated.
  • Programme operations — curriculum progress logged by teachers, school-visit records, student documents, and outreach/telecalling contact history.

These are public examination results, government datasets and partner data. We hold them for the entire JNV system, not only for students in Avanti programmes — which is what makes comparison groups possible.

Selection and entrance tests

  • NCST — the common selection test used to admit students into the residential coaching programmes run within the JNV system. This is the entry point for the cohorts we work with, and it carries far more than scores: household income, stream preference, coaching preferences and other background detail collected at application. Held from two sources, NVS and Dakshana.
  • JNVST — a different and earlier test: the entrance examination into JNV schools themselves, sat around class 6. We hold one year (2018), and only students who were selected — so it cannot support any below-the-cutoff comparison.

School examinations

  • Grade 10 and Grade 12 board results — CBSE results for JNV students, including subject-level marks. Grade 10 matters because it is the last common measurement before students specialise; Grade 12 is the terminal school outcome.

Competitive examinations (outcomes)

  • JEE Main and JEE Advanced — the engineering entrance examinations. We hold percentiles, all-India and category ranks, and qualification flags.
  • NEET — the medical entrance examination, with scores, ranks and qualification.

Diagnostics

  • ASSET (Educational Initiatives) — a diagnostic assessment measuring conceptual understanding rather than exam preparedness. Narrow coverage: Mathematics only, one year, a handful of schools.

Reference data for interpreting outcomes

  • Examination cutoffs — official qualifying cutoffs for JEE and NEET by year and category, plus state-level engineering and medical admission cutoffs. Without these, a score is a number with no meaning attached.
  • Seat and admission data — engineering seat allocation ranks, MBBS seat counts by college, approved intake and enrolment by programme and state. These say what a given rank could actually convert into.
  • Institution reference data — college and university registries, rankings and accreditation records.

Wider context

  • National household and labour survey data — used to relate education outcomes to employment and earnings, well beyond our own students.
  • School enrolment statistics and national learning-proficiency surveys — system-level baselines to compare against.

Rows are the year a cohort finishes school (Grade 12). A student’s records sit across many calendar years, so cohort is the only sane axis: a 2024 graduate sat NCST years earlier and JEE at the end.

Counts are distinct students for whom we hold that data. A student who sat an entrance exam in more than one year is counted once (see Counting students, not attempts).

Graduating Students NCST Grade 10 Grade 12 JEE Main JEE Adv NEET of which Avanti
2021 26,342 10,327 17,453 732
2022 45,550 7 59 35,932 6,951 1 17,565 54
2023 38,208 5 85 36,030 8,197 636 11,348 6,946
2024 51,662 4,482 45,519 34,097 13,769 984 9,914 9,566
2025 60,980 2,104 46,009 33,773 11,199 409 7,564 10,842
2026 50,350 5,985 44,693 not yet 1,724 2 33,641
2027 81,502 37,069 44,433 not yet 34,418
2028 48,787 42,931 45,074 not yet 9,389

means we hold nothing. Small counts are shown as the real number rather than rounded away, because “59 Grade 10 records” and “no Grade 10 records” are different claims: 403,381 students in total, across 464,950 rows.

Reading this table:

  • The early cohorts are outcome-heavy and intake-blind. For 2021 and 2022 we hold competitive-exam results and, from 2022, Grade 12 — but essentially nothing from the start of school. We can see where those students landed, not where they began, and that asymmetry is the single biggest limit on longitudinal analysis of older cohorts.
  • Grade 10 coverage effectively begins with the 2024 cohort and is consistent after it. The 59 and 85 records in 2022 and 2023 are stragglers, not coverage.
  • 2026–2028 have not graduated, so Grade 12 and the competitive-exam columns are empty by definition rather than by omission.
  • NCST coverage grows sharply in the recent cohorts, reflecting both wider collection and the programme’s own growth.
  • The last column is the subset in Avanti programmes. Everything else is the surrounding JNV population, which is what gives us a comparison group.

Everything above describes datasets held. Whether they can be joined for the same student is a separate question, and the honest answer is that the crosswalk is incomplete.

The difficulty is structural: every examination issues its own identifier, and none of them is a stable identifier we hold across systems. An NCST roll number, a CBSE roll number and an NTA application number for one student have no shared key. Linkage is therefore probabilistic — matched on name and date of birth, with a recorded method and confidence — rather than looked up.

Share of each cohort linked to two or more datasets:

Graduating Students Linked to 2+ Share What that looks like in practice
2021 26,342 1,438 5.5% JEE and NEET only; nothing school-side to join to
2022 45,550 14,043 30.8% Grade 12 ↔ competitive exams
2023 38,208 16,035 42.0% Grade 12 ↔ competitive exams, a trace of Grade 10
2024 51,662 36,408 70.5% The most complete cohort — NCST, Grade 10, Grade 12 and outcomes
2025 60,980 30,664 50.3% Grade 10 and Grade 12 linked; outcomes still arriving
2026 50,350 5,341 10.6% NCST and Grade 10 both held, mostly not yet joined
2027 81,502 0 0.0% Both datasets present, none joined — see below
2028 48,787 39,218 80.4% NCST ↔ Grade 10 linked for most of the cohort

JEE Advanced is not counted as a separate source here: it is a subset of JEE Main (Advanced requires qualifying in Main), so counting both would report a student as spanning two datasets when they sat two rounds of one exam, from the same examining body, linked by construction rather than by any matching work. In practice the choice barely matters — counting Advanced separately would move 19 students across all eight cohorts and change no figure in this table to one decimal place — so these shares are robust to it either way.

The 2027 row is the clearest illustration of the gap. We hold 37,069 NCST records and 44,433 Grade 10 records for that cohort — and the two sets do not overlap at all, because 37,069 + 44,433 is exactly the 81,502 students in the row. The data exists on both sides; the matching pass has simply not been run for it. Contrast 2028, where the same two datasets are joined for 80.4% of the cohort.

So coverage in the first table and linkage here move independently: a cohort can be rich in data and poor in linkage, and 2027 is precisely that. (The arithmetic behind that claim is recomputed on every build rather than asserted, so the illustration disappears from headline.json the moment a matching pass runs for 2027.)

Of the 464,950 rows in the crosswalk, 119,918 carry a resolved identity; the remainder appear in a single source with nothing to match against.

Method Records
Name + date of birth 91,130
Name + date of birth (strong agreement) 11,206
Ambiguous — multiple plausible candidates 5,228
Our own roster, with no exam record to match against 5,227
Name + date of birth, names transposed 2,465
Direct student identifier 1,823
Existing examination crosswalk 1,688
Fuzzy name + date of birth 513
Grade 10 roll-number crosswalk 397
Name + father’s name 241

Not all 119,918 are cross-source matches. The 5,227 roster rows are enrolled students seeded from our own records who have no exam, board or NCST record in the crosswalk at all — their “match” is that they are their own identifier, so no matching was performed. Genuine links across two independent sources therefore number 114,691. Nor should a roster row be read as “this student had no result”: some of them do have a production exam result that identity resolution could not reach, usually because the exam record carries no name or date of birth.

This table counts rows, not students, which is the right grain for it: it describes how many links were made, not how many people were reached.

Two caveats a reader should carry forward:

  • Only a small minority of links come from a real shared identifier. The great majority are name-and-date-of-birth matches. They are labelled as such, and the ambiguous ones are marked rather than silently resolved — but they are inferences, not lookups.
  • Unlinked does not mean unlinkable. Much of the gap is simply passes not yet run, as the 2027 row shows. The linkage figures are a snapshot of work in progress, not a ceiling on what is possible.

The crosswalk’s grain is one row per (student, entrance-attempt year). A student who sat JEE in 2025 and again in 2026 has two rows, and their Grade 10 and Grade 12 columns repeat across both. Counting rows therefore counts attempts, and silently over-represents exactly the students who tried more than once.

The gap is large in the cohorts that have finished school:

Graduating Rows Distinct students Inflation if you count rows
2021 50,463 26,342 +91.6%
2022 61,649 45,550 +35.3%
2023 51,070 38,208 +33.7%
2024 59,177 51,662 +14.5%
2025 61,952 60,980 +1.6%
2026–2028 unchanged unchanged none

2026–2028 are unaffected because those cohorts have not sat entrance exams yet, so each student holds exactly one row. That is what makes the error easy to miss: it is invisible in the newest cohorts and worst in the oldest.

Every count on this page is therefore COUNT(DISTINCT student_key), and check.py asserts that no cohort reports more data-holders than it has students — the arithmetic symptom that appears the moment rows get counted as people.

Everything comes from one table, the JNV identity crosswalk avantifellows.external_data_sources.jnv_student_outcome_mapping, which stitches a single student across NCST, Grade 10, Grade 12, JEE Main, JEE Advanced and NEET and resolves them to an Avanti student id where one exists.

Terminal window
python3 build.py # query the warehouse, write output/
BQ_REFRESH=1 python3 build.py # ignore the local cache and re-query
python3 check.py # assert this README agrees with output/

build.py writes four files, all committed so a reader can verify the tables above without warehouse access:

File Contents
output/coverage_by_cohort.csv the coverage table
output/linkage_by_cohort.csv the linkage table
output/match_methods.csv the match-method table
output/headline.json the scalars quoted in prose

check.py runs two kinds of assertion — structural invariants on the extracts, and a textual pass that greps this file for every figure it quotes. It does not touch BigQuery, so a reviewer without warehouse access can still run the gate.

Refreshing is a two-step, deliberately. build.py stamps the run date into headline.json, and check.py asserts the date in the note at the top of this page matches it. So after a refresh the gate fails until that date — and any figure that moved — is updated in the prose. That is the intended friction: a page that tells readers its numbers will change must not be able to go stale while still looking current.

Three properties of the source table shape every query, and each one silently corrupts the output if missed:

  1. The grain is per attempt, not per student — hence COUNT(DISTINCT student_key) everywhere, as above.
  2. cohort_year is a derived integer ladder (2021–2028), computed as the Grade 12 exam year, or Grade 10 + 2, or NCST + 2, or the first entrance attempt. It is not academic_year, so the warehouse’s usual current-cohort filter must not be applied here.
  3. has_jee_adv_data is a subset of has_jee_mains_data, which is why Advanced is excluded from the linkage source list.

An earlier version of this page reported the coverage and linkage tables from row counts rather than distinct students, which overstated the 2021 cohort by 91.6% and the 2022–2024 cohorts by 14.5–35.3%, and overstated linkage in every cohort through 2025. The match-method table was unaffected, because rows are the correct grain there. The figures above supersede it. Those numbers were hand-derived from ad-hoc queries; this folder exists so that cannot recur.