Data Scientist Intern

Ground synthetic Indian populations in real data, and design the tests that tell us when our predictions are right, wrong, or just lucky. Internship, India · remote-friendly.

About the role

Ground synthetic Indian populations in real data, and design the tests that tell us when our predictions are right, wrong, or just lucky.

What you'll do

Build population models from Indian public data — census, NFHS, HCES, PLFS — and state where each number comes from

Replace analyst priors with evidence, one dimension at a time

Design evaluation: ranking accuracy, calibration, baselines, and where the model fails

Work with creators' real outcomes to measure prediction error honestly

You might be a fit if

You can explain a confidence interval to someone who has never heard of one

You have opinions about survey data, and about when not to trust it

You are comfortable with Python or R, SQL and messy public datasets

What you'll learn

Population synthesis and post-stratification

Evaluation design for predictive systems

India's data landscape, in depth

What we ask in the application

Estimate what share of 18–24-year-olds in Tier-2 Tamil Nadu watch personal-finance videos weekly. Walk us through the sources you'd use, the assumptions you'd make, and how wrong you might be.

What evidence would convince you that our predictions beat a creator's own gut feeling? Design the test.

Link to an analysis you did

Tell us about a time you were confidently wrong. How did you find out, and what changed after?

Pick one thing Indian consumers or viewers do that outsiders consistently get wrong. Explain what's really going on.

Simulating how people respond is either very useful or quietly dangerous. Where do you think the line is?

Contact: hello@thommidhi.com