Portfolio

Digitizing the 1951 Indian Population Census

Overcoming the hurdles in extracting data at scale from noisy historical documents using LLMs.

Last updated
  • Data Science
  • Python
  • Machine Learning

Overview

At Development Data Lab, I designed a novel large language model based data extraction architecture to digitize historical Indian Population Census microdata. The pipeline produced the first ever digitized version of the 1951 Census, and has since been extended to the 1961 and 1971 Censuses.

Talk: The Fifth Elephant

I presented this work at The Fifth Elephant conference in a talk titled Decoding 1951 census: LLMs extract messy tables.