Digitizing the 1951 Indian Population Census
Overcoming the hurdles in extracting data at scale from noisy historical documents using LLMs.
Overview
At Development Data Lab, I designed a novel large language model based data extraction architecture to digitize historical Indian Population Census microdata. The pipeline produced the first ever digitized version of the 1951 Census, and has since been extended to the 1961 and 1971 Censuses.
Talk: The Fifth Elephant
I presented this work at The Fifth Elephant conference in a talk titled Decoding 1951 census: LLMs extract messy tables.