What data we hold
These figures are a snapshot, and they will change. Our data is actively being worked on: new sources are still being ingested, cohorts are still graduating, and identity-matching passes are still being run across datasets we already hold. Every number below is therefore a picture of one moment, not a settled total — and the direction of travel is generally upward, particularly for linkage, where much of the gap is passes not yet run rather than data that cannot be joined. Figures generated 2026-08-11. If you are citing anything here, cite it with that date, and re-run the build for a current answer.
Aggregate counts from the analytical warehouse. Nothing on this page identifies an individual student: every figure is a cohort-level count or a share.
Every number here is generated by build.py from a single warehouse table and committed
to output/. check.py gates this prose against those extracts, so a figure cannot be
edited by hand without failing the check — including the date above. See
Reproducing this.
What data we hold
Section titled “What data we hold”Avanti’s data divides into two halves: what we observe directly about students in our own programmes, and public or partner datasets about the wider system those students sit inside. Most of the interesting questions need both, which makes linking them the hard part — and the part that is least finished.
What we collect ourselves
Section titled “What we collect ourselves”Our own systems produce five broad kinds of student data:
- Enrolment and identity over time — which programme, batch, school and centre a student belongs to, and how that changes across academic years. A student is not a fixed row; their enrolment is a journey with transitions.
- Attendance at live classes — per student, per session. This is the highest-volume operational signal we hold and the one most used to spot disengagement early.
- Test performance, down to the individual question — results are modelled at overall, chapter and question level, alongside a separate record of who actually sat each test. Question-level granularity is what makes diagnostic analysis possible rather than just ranking.
- Self-driven practice — practice attempts a student chooses to make outside scheduled tests, kept separate from formal assessments so the two are never conflated.
- Programme operations — curriculum progress logged by teachers, school-visit records, student documents, and outreach/telecalling contact history.
What we bring in from outside
Section titled “What we bring in from outside”These are public examination results, government datasets and partner data. We hold them for the entire JNV system, not only for students in Avanti programmes — which is what makes comparison groups possible.
Selection and entrance tests
- NCST — the common selection test used to admit students into the residential coaching programmes run within the JNV system. This is the entry point for the cohorts we work with, and it carries far more than scores: household income, stream preference, coaching preferences and other background detail collected at application. Held from two sources, NVS and Dakshana.
- JNVST — a different and earlier test: the entrance examination into JNV schools themselves, sat around class 6. We hold one year (2018), and only students who were selected — so it cannot support any below-the-cutoff comparison.
School examinations
- Grade 10 and Grade 12 board results — CBSE results for JNV students, including subject-level marks. Grade 10 matters because it is the last common measurement before students specialise; Grade 12 is the terminal school outcome.
Competitive examinations (outcomes)
- JEE Main and JEE Advanced — the engineering entrance examinations. We hold percentiles, all-India and category ranks, and qualification flags.
- NEET — the medical entrance examination, with scores, ranks and qualification.
Diagnostics
- ASSET (Educational Initiatives) — a diagnostic assessment measuring conceptual understanding rather than exam preparedness. Narrow coverage: Mathematics only, one year, a handful of schools.
Reference data for interpreting outcomes
- Examination cutoffs — official qualifying cutoffs for JEE and NEET by year and category, plus state-level engineering and medical admission cutoffs. Without these, a score is a number with no meaning attached.
- Seat and admission data — engineering seat allocation ranks, MBBS seat counts by college, approved intake and enrolment by programme and state. These say what a given rank could actually convert into.
- Institution reference data — college and university registries, rankings and accreditation records.
Wider context
- National household and labour survey data — used to relate education outcomes to employment and earnings, well beyond our own students.
- School enrolment statistics and national learning-proficiency surveys — system-level baselines to compare against.
Coverage by graduating cohort
Section titled “Coverage by graduating cohort”Rows are the year a cohort finishes school (Grade 12). A student’s records sit across many calendar years, so cohort is the only sane axis: a 2024 graduate sat NCST years earlier and JEE at the end.
Counts are distinct students for whom we hold that data. A student who sat an entrance exam in more than one year is counted once (see Counting students, not attempts).
| Graduating | Students | NCST | Grade 10 | Grade 12 | JEE Main | JEE Adv | NEET | of which Avanti |
|---|---|---|---|---|---|---|---|---|
| 2021 | 26,342 | — | — | — | 10,327 | — | 17,453 | 732 |
| 2022 | 45,550 | 7 | 59 | 35,932 | 6,951 | 1 | 17,565 | 54 |
| 2023 | 38,208 | 5 | 85 | 36,030 | 8,197 | 636 | 11,348 | 6,946 |
| 2024 | 51,662 | 4,482 | 45,519 | 34,097 | 13,769 | 984 | 9,914 | 9,566 |
| 2025 | 60,980 | 2,104 | 46,009 | 33,773 | 11,199 | 409 | 7,564 | 10,842 |
| 2026 | 50,350 | 5,985 | 44,693 | not yet | 1,724 | — | 2 | 33,641 |
| 2027 | 81,502 | 37,069 | 44,433 | not yet | — | — | — | 34,418 |
| 2028 | 48,787 | 42,931 | 45,074 | not yet | — | — | — | 9,389 |
— means we hold nothing. Small counts are shown as the real number rather than rounded away, because “59 Grade 10 records” and “no Grade 10 records” are different claims: 403,381 students in total, across 464,950 rows.
Reading this table:
- The early cohorts are outcome-heavy and intake-blind. For 2021 and 2022 we hold competitive-exam results and, from 2022, Grade 12 — but essentially nothing from the start of school. We can see where those students landed, not where they began, and that asymmetry is the single biggest limit on longitudinal analysis of older cohorts.
- Grade 10 coverage effectively begins with the 2024 cohort and is consistent after it. The 59 and 85 records in 2022 and 2023 are stragglers, not coverage.
- 2026–2028 have not graduated, so Grade 12 and the competitive-exam columns are empty by definition rather than by omission.
- NCST coverage grows sharply in the recent cohorts, reflecting both wider collection and the programme’s own growth.
- The last column is the subset in Avanti programmes. Everything else is the surrounding JNV population, which is what gives us a comparison group.
How much of it is actually linked
Section titled “How much of it is actually linked”Everything above describes datasets held. Whether they can be joined for the same student is a separate question, and the honest answer is that the crosswalk is incomplete.
The difficulty is structural: every examination issues its own identifier, and none of them is a stable identifier we hold across systems. An NCST roll number, a CBSE roll number and an NTA application number for one student have no shared key. Linkage is therefore probabilistic — matched on name and date of birth, with a recorded method and confidence — rather than looked up.
Share of each cohort linked to two or more datasets:
| Graduating | Students | Linked to 2+ | Share | What that looks like in practice |
|---|---|---|---|---|
| 2021 | 26,342 | 1,438 | 5.5% | JEE and NEET only; nothing school-side to join to |
| 2022 | 45,550 | 14,043 | 30.8% | Grade 12 ↔ competitive exams |
| 2023 | 38,208 | 16,035 | 42.0% | Grade 12 ↔ competitive exams, a trace of Grade 10 |
| 2024 | 51,662 | 36,408 | 70.5% | The most complete cohort — NCST, Grade 10, Grade 12 and outcomes |
| 2025 | 60,980 | 30,664 | 50.3% | Grade 10 and Grade 12 linked; outcomes still arriving |
| 2026 | 50,350 | 5,341 | 10.6% | NCST and Grade 10 both held, mostly not yet joined |
| 2027 | 81,502 | 0 | 0.0% | Both datasets present, none joined — see below |
| 2028 | 48,787 | 39,218 | 80.4% | NCST ↔ Grade 10 linked for most of the cohort |
JEE Advanced is not counted as a separate source here: it is a subset of JEE Main (Advanced requires qualifying in Main), so counting both would report a student as spanning two datasets when they sat two rounds of one exam, from the same examining body, linked by construction rather than by any matching work. In practice the choice barely matters — counting Advanced separately would move 19 students across all eight cohorts and change no figure in this table to one decimal place — so these shares are robust to it either way.
The 2027 row is the clearest illustration of the gap. We hold 37,069 NCST records and 44,433 Grade 10 records for that cohort — and the two sets do not overlap at all, because 37,069 + 44,433 is exactly the 81,502 students in the row. The data exists on both sides; the matching pass has simply not been run for it. Contrast 2028, where the same two datasets are joined for 80.4% of the cohort.
So coverage in the first table and linkage here move independently: a cohort can be
rich in data and poor in linkage, and 2027 is precisely that. (The arithmetic behind
that claim is recomputed on every build rather than asserted, so the illustration
disappears from headline.json the moment a matching pass runs for 2027.)
How the matching was done
Section titled “How the matching was done”Of the 464,950 rows in the crosswalk, 119,918 carry a resolved identity; the remainder appear in a single source with nothing to match against.
| Method | Records |
|---|---|
| Name + date of birth | 91,130 |
| Name + date of birth (strong agreement) | 11,206 |
| Ambiguous — multiple plausible candidates | 5,228 |
| Our own roster, with no exam record to match against | 5,227 |
| Name + date of birth, names transposed | 2,465 |
| Direct student identifier | 1,823 |
| Existing examination crosswalk | 1,688 |
| Fuzzy name + date of birth | 513 |
| Grade 10 roll-number crosswalk | 397 |
| Name + father’s name | 241 |
Not all 119,918 are cross-source matches. The 5,227 roster rows are enrolled students seeded from our own records who have no exam, board or NCST record in the crosswalk at all — their “match” is that they are their own identifier, so no matching was performed. Genuine links across two independent sources therefore number 114,691. Nor should a roster row be read as “this student had no result”: some of them do have a production exam result that identity resolution could not reach, usually because the exam record carries no name or date of birth.
This table counts rows, not students, which is the right grain for it: it describes how many links were made, not how many people were reached.
Two caveats a reader should carry forward:
- Only a small minority of links come from a real shared identifier. The great majority are name-and-date-of-birth matches. They are labelled as such, and the ambiguous ones are marked rather than silently resolved — but they are inferences, not lookups.
- Unlinked does not mean unlinkable. Much of the gap is simply passes not yet run, as the 2027 row shows. The linkage figures are a snapshot of work in progress, not a ceiling on what is possible.
Counting students, not attempts
Section titled “Counting students, not attempts”The crosswalk’s grain is one row per (student, entrance-attempt year). A student who sat JEE in 2025 and again in 2026 has two rows, and their Grade 10 and Grade 12 columns repeat across both. Counting rows therefore counts attempts, and silently over-represents exactly the students who tried more than once.
The gap is large in the cohorts that have finished school:
| Graduating | Rows | Distinct students | Inflation if you count rows |
|---|---|---|---|
| 2021 | 50,463 | 26,342 | +91.6% |
| 2022 | 61,649 | 45,550 | +35.3% |
| 2023 | 51,070 | 38,208 | +33.7% |
| 2024 | 59,177 | 51,662 | +14.5% |
| 2025 | 61,952 | 60,980 | +1.6% |
| 2026–2028 | unchanged | unchanged | none |
2026–2028 are unaffected because those cohorts have not sat entrance exams yet, so each student holds exactly one row. That is what makes the error easy to miss: it is invisible in the newest cohorts and worst in the oldest.
Every count on this page is therefore COUNT(DISTINCT student_key), and check.py
asserts that no cohort reports more data-holders than it has students — the arithmetic
symptom that appears the moment rows get counted as people.
Reproducing this
Section titled “Reproducing this”Everything comes from one table, the JNV identity crosswalk
avantifellows.external_data_sources.jnv_student_outcome_mapping, which stitches a
single student across NCST, Grade 10, Grade 12, JEE Main, JEE Advanced and NEET and
resolves them to an Avanti student id where one exists.
python3 build.py # query the warehouse, write output/BQ_REFRESH=1 python3 build.py # ignore the local cache and re-querypython3 check.py # assert this README agrees with output/build.py writes four files, all committed so a reader can verify the tables above
without warehouse access:
| File | Contents |
|---|---|
output/coverage_by_cohort.csv |
the coverage table |
output/linkage_by_cohort.csv |
the linkage table |
output/match_methods.csv |
the match-method table |
output/headline.json |
the scalars quoted in prose |
check.py runs two kinds of assertion — structural invariants on the extracts, and a
textual pass that greps this file for every figure it quotes. It does not touch
BigQuery, so a reviewer without warehouse access can still run the gate.
Refreshing is a two-step, deliberately. build.py stamps the run date into
headline.json, and check.py asserts the date in the note at the top of this page
matches it. So after a refresh the gate fails until that date — and any figure that
moved — is updated in the prose. That is the intended friction: a page that tells readers
its numbers will change must not be able to go stale while still looking current.
Three properties of the source table shape every query, and each one silently corrupts the output if missed:
- The grain is per attempt, not per student — hence
COUNT(DISTINCT student_key)everywhere, as above. cohort_yearis a derived integer ladder (2021–2028), computed as the Grade 12 exam year, or Grade 10 + 2, or NCST + 2, or the first entrance attempt. It is notacademic_year, so the warehouse’s usual current-cohort filter must not be applied here.has_jee_adv_datais a subset ofhas_jee_mains_data, which is why Advanced is excluded from the linkage source list.
What changed
Section titled “What changed”An earlier version of this page reported the coverage and linkage tables from row counts rather than distinct students, which overstated the 2021 cohort by 91.6% and the 2022–2024 cohorts by 14.5–35.3%, and overstated linkage in every cohort through 2025. The match-method table was unaffected, because rows are the correct grain there. The figures above supersede it. Those numbers were hand-derived from ad-hoc queries; this folder exists so that cannot recur.