Application briefing · Google DeepMind
Two open roles, one team, one job: hand-build the realistic, week-long eval tasks that measure what Gemini can actually do as a software engineer. This page is your briefing, your fit case, and your interview prep — in one place.
Same team (Polaris, DeepMind, Mountain View), same core work: design tasks that test Gemini's ability to complete long-horizon software engineering. A task = a prompt that instructs the model + a grader/rubric that scores the attempt. One well-designed task takes about a week, built by hand — no synthetic pipelines.
$147K–210K + 15% bonus + equity
Minimums: software dev experience; experience with AI-assisted coding tools / developer agents / LLM workflows; experience using coding agents for work.
Preferred: competitive programming; rubric/eval-framework/benchmark design; fast learning of new stacks.
$350K–525K base · $550K–1.1M total comp
Minimums: experience using AI coding agents; strong analytical and problem-solving skills. No degree requirements.
Preferred: competitive programming / math competitions.
Extra scope: evaluate weaknesses of latest models; build data and pipelines for improving Gemini.
The play: apply to both. The FTC pays ~3× the full-time base for the same craft; the full-time role is the stable seat. Nothing in either posting rules out applying to both reqs.
The posting asks for someone who uses coding agents and can build the tasks that expose what models can't do yet. That is your daily work, not a line on a resume:
The gap to own, not hide: no competitive-programming background (preferred on both postings). Don't claim it. Your counterweight is stronger: years of designing evaluations for agentic output, which is the actual job.
Proof points to drop in the application and interviews. All live, all checkable:
"I don't just use coding agents — I run a practice where every agentic build passes rubrics I design. Prompt plus grader, iterated until it meets the bar. That is eval-task design on production work, daily."
Polaris interviews will probe one thing above all: can you design a good eval task? Expect to walk through a task end-to-end on a whiteboard. The axes they'll push on:
A worked example in the Polaris shape. Pick one from your own work and rehearse it out loud — this is the closest thing to a take-home for this team:
Prompt: "Port this 400-line Python CLI tool to TypeScript. It must keep the exact CLI surface (flags, exit codes, stdout format), pass the existing 23-test fixture suite, and handle the three documented edge cases in EDGE_CASES.md. You may restructure internals freely."
tsc --strict clean.Your edge in the room: you've lived the failure side. When they ask "where do agents break on this?", you answer from production experience: they skim EDGE_CASES.md, they "fix" tests instead of code, they leave any types that fail --strict. That specificity is what the team hires for.