Application briefing · Google DeepMind

Polaris: the evals team that grades Gemini.

Two open roles, one team, one job: hand-build the realistic, week-long eval tasks that measure what Gemini can actually do as a software engineer. This page is your briefing, your fit case, and your interview prep — in one place.

01The two roles

Same team (Polaris, DeepMind, Mountain View), same core work: design tasks that test Gemini's ability to complete long-horizon software engineering. A task = a prompt that instructs the model + a grader/rubric that scores the attempt. One well-designed task takes about a week, built by hand — no synthetic pipelines.

Full-time

Research Engineer, Polaris

$147K–210K + 15% bonus + equity

Minimums: software dev experience; experience with AI-assisted coding tools / developer agents / LLM workflows; experience using coding agents for work.

Preferred: competitive programming; rubric/eval-framework/benchmark design; fast learning of new stacks.

Req 114077679900074694 →

12-month contract

Evals Engineer, Polaris

$350K–525K base · $550K–1.1M total comp

Minimums: experience using AI coding agents; strong analytical and problem-solving skills. No degree requirements.

Preferred: competitive programming / math competitions.

Extra scope: evaluate weaknesses of latest models; build data and pipelines for improving Gemini.

Req 95145323399127750 →

The play: apply to both. The FTC pays ~3× the full-time base for the same craft; the full-time role is the stable seat. Nothing in either posting rules out applying to both reqs.

02Why you fit

The posting asks for someone who uses coding agents and can build the tasks that expose what models can't do yet. That is your daily work, not a line on a resume:

The gap to own, not hide: no competitive-programming background (preferred on both postings). Don't claim it. Your counterweight is stronger: years of designing evaluations for agentic output, which is the actual job.

03Your ammo

Proof points to drop in the application and interviews. All live, all checkable:

The one-liner

"I don't just use coding agents — I run a practice where every agentic build passes rubrics I design. Prompt plus grader, iterated until it meets the bar. That is eval-task design on production work, daily."

04Interview prep

Polaris interviews will probe one thing above all: can you design a good eval task? Expect to walk through a task end-to-end on a whiteboard. The axes they'll push on:

Questions to ask them

05Eval-task walkthrough (rehearse this)

A worked example in the Polaris shape. Pick one from your own work and rehearse it out loud — this is the closest thing to a take-home for this team:

The task

Prompt: "Port this 400-line Python CLI tool to TypeScript. It must keep the exact CLI surface (flags, exit codes, stdout format), pass the existing 23-test fixture suite, and handle the three documented edge cases in EDGE_CASES.md. You may restructure internals freely."

The grader

Why it's a good task

Your edge in the room: you've lived the failure side. When they ask "where do agents break on this?", you answer from production experience: they skim EDGE_CASES.md, they "fix" tests instead of code, they leave any types that fail --strict. That specificity is what the team hires for.

06Comp & logistics

07Checklist