AI Engineer Intern
Build the decision and reasoning layers that turn a sampled population into a prediction — and make every model in the loop replaceable, measurable and honest. Internship, India · remote-friendly.
About the role
Build the decision and reasoning layers that turn a sampled population into a prediction — and make every model in the loop replaceable, measurable and honest.
What you'll do
Integrate decision models (Jev, Claude, open models) behind one contract and benchmark them against each other
Harden the simulation pipeline on Cloudflare Workers, Workflows and D1: retries, validation, provenance, cost
Build guards that keep language models from inventing numbers
Instrument runs so every prediction can later be scored against reality
You might be a fit if
You have shipped something with an LLM and watched it fail in a way you didn't expect
You are comfortable in TypeScript or Python and curious about the other
You read model outputs as data to be validated, not answers to be trusted
What you'll learn
Production LLM systems end to end
Evaluation and calibration in practice
Edge infrastructure at low cost
What we ask in the application
Describe something you built with a language model that broke — or would have broken — in production. What was the failure mode, and how would you catch it automatically?
Our decision layer returns probabilities, not prose. Using only data we would actually have (predictions and later outcomes), how would you tell whether those probabilities are miscalibrated?
Link to code you are proud of
Tell us about a time you were confidently wrong. How did you find out, and what changed after?
Pick one thing Indian consumers or viewers do that outsiders consistently get wrong. Explain what's really going on.
Simulating how people respond is either very useful or quietly dangerous. Where do you think the line is?
Contact: hello@thommidhi.com