AI Engineer Intern

Build the decision and reasoning layers that turn a sampled population into a prediction — and make every model in the loop replaceable, measurable and honest. Internship, India · remote-friendly.

About the role

Build the decision and reasoning layers that turn a sampled population into a prediction — and make every model in the loop replaceable, measurable and honest.

What you'll do

Integrate decision models (Jev, Claude, open models) behind one contract and benchmark them against each other

Harden the simulation pipeline on Cloudflare Workers, Workflows and D1: retries, validation, provenance, cost

Build guards that keep language models from inventing numbers

Instrument runs so every prediction can later be scored against reality

You might be a fit if

You have shipped something with an LLM and watched it fail in a way you didn't expect

You are comfortable in TypeScript or Python and curious about the other

You read model outputs as data to be validated, not answers to be trusted

What you'll learn

Production LLM systems end to end

Evaluation and calibration in practice

Edge infrastructure at low cost

What we ask in the application

Describe something you built with a language model that broke — or would have broken — in production. What was the failure mode, and how would you catch it automatically?

Our decision layer returns probabilities, not prose. Using only data we would actually have (predictions and later outcomes), how would you tell whether those probabilities are miscalibrated?

Link to code you are proud of

Tell us about a time you were confidently wrong. How did you find out, and what changed after?

Pick one thing Indian consumers or viewers do that outsiders consistently get wrong. Explain what's really going on.

Simulating how people respond is either very useful or quietly dangerous. Where do you think the line is?

Contact: hello@thommidhi.com