How the numbers are made
Full details, code, and figures live in the open analysis writeup. This page is the plain-language version.
1. Scores, standardized carefully
We use each school's mean scale score on the CAASPP Smarter Balanced tests (2015–2025) — never "percent proficient", which a school can move just by nudging students across a cut line. Scores are compared within the same grade, subject, and year, in units of student-level standard deviations. We exclude 2020 (no testing) and 2021 (statewide participation collapsed to 24%).
2. Three numbers per school
A precision-weighted model gives each school a level (how its students score vs the state), a growth rate (how much a cohort progresses from one grade to the next vs the state), and a trend (whether results are improving across years). Small schools' estimates are shrunk toward the average — a 40-student grade simply tells us less than a 400-student one — and estimates that stay unreliable after shrinkage are not shown.
3. Comparing schools that serve similar students
Raw score levels correlate −0.76 with the share of economically disadvantaged students — demography, not schooling, dominates raw comparisons. So we also report performance relative to schools serving similar students: what remains after accounting for the tested population's economic disadvantage, race/ethnicity, English-learner and disability shares, parental education, and school size. This residual is not a measure of school quality — it still contains everything our covariates miss.
4. What we refuse to do
- No single composite rating, and no bare integer rankings.
- No causal claims — these are relationships between school averages.
- No school-level subgroup estimates where suppression and small counts make them unreliable.
- No hiding of specification sensitivity: when reasonable modeling choices move a school's result, we say so. Rankings agree between our specifications at Spearman 0.82–0.92 — good, but far from perfect, and that disagreement is the error bar that matters most.
Why we're careful: lessons we inherited
The LA Times' 2010 teacher value-added project published quintile labels whose ratings flipped for half the teachers under an equally-defensible alternative model (Briggs & Domingue, 2011). The Urban Institute's demographically adjusted NAEP work and Stanford's SEDA project showed the responsible alternative: precision weighting, shrinkage, reliability gates, and published sensitivity. We follow the second tradition, and our growth measure passes the falsification test the LA Times model failed (correlation with prior achievement: +0.01 here vs their 0.50).