Data Scientist Intern
Ground synthetic Indian populations in real data, and design the tests that tell us when our predictions are right, wrong, or just lucky. Internship, India · remote-friendly.
About the role
Ground synthetic Indian populations in real data, and design the tests that tell us when our predictions are right, wrong, or just lucky.
What you'll do
Build population models from Indian public data — census, NFHS, HCES, PLFS — and state where each number comes from
Replace analyst priors with evidence, one dimension at a time
Design evaluation: ranking accuracy, calibration, baselines, and where the model fails
Work with creators' real outcomes to measure prediction error honestly
You might be a fit if
You can explain a confidence interval to someone who has never heard of one
You have opinions about survey data, and about when not to trust it
You are comfortable with Python or R, SQL and messy public datasets
What you'll learn
Population synthesis and post-stratification
Evaluation design for predictive systems
India's data landscape, in depth
What we ask in the application
Estimate what share of 18–24-year-olds in Tier-2 Tamil Nadu watch personal-finance videos weekly. Walk us through the sources you'd use, the assumptions you'd make, and how wrong you might be.
What evidence would convince you that our predictions beat a creator's own gut feeling? Design the test.
Link to an analysis you did
Tell us about a time you were confidently wrong. How did you find out, and what changed after?
Pick one thing Indian consumers or viewers do that outsiders consistently get wrong. Explain what's really going on.
Simulating how people respond is either very useful or quietly dangerous. Where do you think the line is?
Contact: hello@thommidhi.com